Missing data imputation using supervised learning methods

Document Type : Original Scientific Paper

Authors

School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran, Tehran, Iran

Abstract

Missing data is a very common problem in all research fields. Case deletion is a simple way to handle incomplete data sets which could mislead to biased statistical results. A more reliable approach to handle missing values is imputation which allows covariate-dependent missing mechanism, as well. This paper aims to prepare guidance for researchers facing missing data problems by comparing various imputation methods including machine learning techniques, to achieve better results in supervised learning tasks. A benchmark dataset has experimented and the results are compared by applying popular classifiers over varying missing mechanisms and rates on this benchmark dataset.

Keywords

Main Subjects


Batista, G.E. and Monard, M.C. (2003). An analysis of four missing data treatment methods for supervised learning. Applied Artificial Intelligence, 17(5-6):519–533.
Breiman, L. (2001). Random forests. Machine Learning, 45(1):5–32.
Graham, J.W. (2009). Missing data analysis: making it work in the real world. Annual Review of Psychology, 60(1):549–576.
Poulos, J. and Valle, R. (2018). Missing data imputation for supervised learning. Applied Artificial Intelligence, 32(2):186–196.
Kohavi, R. (1996). Scaling Up the Accuracy of Naive-Bayes Classifier: a Decision-Tree Hybrid. Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, 96(1):202–207.
Little, R. and Rubin, D. (2014). Statistical Analysis with Missing Data. John Wiley and Sons.
Rubin, D.B. (1976). Inference and Missing Data. Biometrika, 63:581–592.
Schafer, J.L. (1999). Multiple imputation: a primer. Statistical Methods in Medical Research, 8(1):3–15.
Silva-Ramírez, E.L., Pino-Mejías, R., López-Coello, M. and Cubiles-de-la-Vega, M.D. (2011). Missing value imputation on missing completely at random data using multilayer perceptrons. Neural Networks, 24(1):121–129.
Stekhoven, D.J. and Buehlmann, P. (2012). MissForest - nonparametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112–118.
Tsiatis, A. (2007). Semiparametric Theory and Missing Data. Springer Science and Business Media.
Van Buuren, S. and Oudshoorn, K. (1999). Flexible multivariate imputation by MICE (pp. 1–20). Leiden: TNO.