Blind label ratio estimation

Document Type : Original Scientific Paper

Authors

1 Department of Electrical and Computer Engineering‎, ‎Babol Noshirvani University of Technology‎, ‎Tehran‎, ‎Iran

2 Son Corporate Group‎, ‎Tehran‎, ‎Iran

Abstract

Many anomaly detection algorithms require knowledge of the ratio of the two labels to operate‎. ‎In real life‎, ‎however‎, ‎we may not have access to this value‎. ‎As such‎, ‎we often run anomaly detection packages with default values that may differ significantly from the actual value‎. ‎Experiments on multiple datasets show that correctly determination of this ratio or at least obtaining a close estimate can makes a significant difference in the final performance of the anomaly detection algorithm‎. ‎In this paper‎, ‎we address the problem of estimating this ratio using both theoretical and heuristic techniques‎. ‎In the theoretical method‎, ‎we maximize the mutual information between features and labels to find the exact ratio‎. ‎In the heuristic method‎, ‎we sweep the [0,1] range in 0.01 steps to search for the ratio‎. ‎On each iteration‎, ‎we run the anomaly detection algorithm based on the ratio for that iteration and record the correlation coefficient between the features and the label generated by the algorithm‎. ‎After the 100th iteration‎, ‎we declare the ratio that provides the maximum correlation coefficient as our estimate of the label ratio‎. ‎Our experiments on multiple datasets and several anomaly detection algorithms show that maximizing the correlation coefficient leads to the best results.

Keywords

Main Subjects


Abdi, H. (2007). Multiple correlation coefficient. Encyclopedia of Measurement and Statistics, 648(651):19.
Beraha, M., Metelli, A.M., Papini, M., Tirinzoni, A. and Restelli, M. (2019). Feature selection via mutual information: New theoretical insights. In 2019 International Joint Conference on Neural Networks (IJCNN), pp. 1-9. IEEE.
Bootkrajang, J. and Chaijaruwanich, J. (2020). Towards an improved label noise proportion estimation in small data: A Bayesian approach. International Journal of Machine Learning and Cybernetics, 13(4):851–867.
Boyd, S. and Vanderberghe, L. (2004). Convex Optimization. Cambridge University Press.
Brockett, P.L., Derrig, R.A., Golden, L.L., Levine, A. and Alpert, M. (2002). Fraud classification using principal component analysis of RIDITs. Journal of Risk and Insurance, 69(3):341–371.
Chen, L., Zaharia, M. and Zou, J. (2022). Estimating and explaining model performance when both covariates and labels shift. Advances in Neural Information Processing Systems, 35:11467–11479.
Hsu, H.-H. and Hsieh, C.-W. (2010). Feature selection via correlation coefficient clustering. Journal of Software, 5(12):1371–1377.
Iyer, A.S., Nath, J.S. and Sarawagi, S. (2016). Privacy-preserving class ratio estimation. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 925–934.
Kaggle. (n.d.). Medical provider fraud detection. Available at https://www.kaggle.com/code/rohitrox/medical-provider-fraud-detection/input
Nian, K., Zhang, H., Tayal, A., Coleman, T. and Li, Y. (2016). Unsupervised spectral ranking for anomaly and application to auto insurance fraud detection. Journal of Finance and Data Science, 2(1):1–28.
Quadrianto, N., Smola, A.J., Caetano, T.S. and Le, Q.V. (2009). Estimating labels from label proportions. Journal of Machine Learning Research, 10:2349–2374.
Shaeiri, Z. and Kazemitabar, S.J. (2020). Fast unsupervised automobile insurance fraud detection based on spectral ranking of anomalies. International Journal of Engineering, 33(7):1240–1248.
Shannon, C. (1948). A mathematical theory of communication. The Bell System Technical Journal, 27(3):379–423.
Yang, J., Rahardja, S. and Fränti, P. (2019). Outlier detection: How to threshold outlier scores? In Proceedings of the International Conference on Artificial Intelligence, Information Processing and Cloud Computing, 1–6.