TY - JOUR ID - ICD2009b T1 - Clinical Data Mining: a Review A1 - Iavindrasana, Jimison A1 - Cohen, Gilles A1 - Depeursinge, Adrien A1 - Müller, Henning A1 - Meyer, Rodolphe A1 - Geissbuhler, Antoine JA - Yearbook of Medical Informatics Y1 - 2009 SP - 121 EP - 133 KW - clinical data analysis KW - data mining N2 - Clinical data mining is the application of data mining techniques using clinical data. We review the literature in order to provide a general overview by identifying the status-of-practice and the challenges ahead. The nine data mining steps proposed by Fayyad in 1996 [4] were used as the main themes of the review. Pubmed was used as primary source and 84 papers were retained based on our inclusion criteria. Clinical data mining has three objectives: understanding the clinical data, assist healthcare professionals, and develop a data analysis methodology suitable for medical data. Classification is the most frequently used data mining function with a predominance of the implementation of Bayesian classifiers, neural networks, and SVMs (Support Vector Machines). A myriad of quantitative performance measures were proposed with a predominance of accuracy, sensitivity, specificity, and ROC curves. The latter are usually associated with qualitative evaluation. Clinical data mining respects its commitment to extracting new and previously unknown knowledge from clinical databases. More efforts are still needed to obtain a wider acceptance from the healthcare professionals and for generalization of the knowledge and reproducibility of its extraction process: better description of variables, systematic report of algorithm parameters including the method to obtain them, use of easy-to-understand models and comparisons of the efficiency of clinical data mining with traditional statistical analyses. More and more data will be available for data miners and they have to develop new methodologies and infrastructures to analyze the increasingly complex medical data. ER -