Validation Assessments on Resampling Method in Imbalanced Binary Classification for Linear Discriminant Analysis

Authors

  • Ahmad Hakiim Jamaluddin Department of Mathematics, Universiti Putra Malaysia, Malaysia
  • Nor Idayu Mahat Centre for Testing, Measurement and Appraisal, Universiti Utara Malaysia, Malaysia

DOI:

https://doi.org/10.32890/jict2021.20.1.5

Keywords:

Linear discriminant analysis, pre-processing, resampling method, class imbalance, binary classification

Abstract

The curse of class imbalance affects the performance of many conventional classification algorithms including linear discriminant analysis (LDA). The data pre-processing approach through some resampling methods such as random oversampling (ROS) and random undersampling (RUS) is one of the treatments to alleviate such curse. Previous studies have attempted to address the effect of a resampling method on the performance of LDA. However, some studies contradicted with each other based on different performance measures as well as validation strategies. This manuscript attempted to shed more light on the effect of a resampling method (ROS or RUS) on the performance of LDA based on true positive rate and true negative rate through five validation strategies, i.e. leave-one-out cross-validation, k-fold cross-validation, repeated k-fold cross-validation, naive bootstrap, and .632+ bootstrap. 100 two-group bivariate normally distributed simulated and four real data sets with severe class imbalance ratio were utilised. The analysis on the location and dispersion statistics of the performance measures was further enlightened on: (i) the effect of a resampling method on the performance of LDA, and (ii) the enhancement in the learning fairness of LDA on objects regardless of sample size, hence reducing the effect of the curse of class imbalance.

References

Alcalá-Fdez, J., Fernández, A., Luengo, J., Derrac, J., García, S., Sánchez, L., & Herrera, F. (2011). KEEL data-mining software tool: Data set repository, integration of algorithms and experimental analysis framework. J. of Mult.-Valued Logic & Soft Computing, 17, 255–287.

Arlot, S., & Celisse, A. (2010). A survey of cross-validation procedures for model selection. Statistics Surveys, 4, 40–79. https://doi.org/10.1214/09-SS054

Branco, P., Torgo, L., & Ribeiro, P. R. (2016). A survey of predictive modeling on imbalanced domains. ACM Computing Surveys, 49(2), 31:1–31:50. https://doi.org/10.1145/2907070

Burman, P. (1989). A comparative study of ordinary cross-validation, v-fold cross-validation and the repeated learning-testing methods. Biometrika, 76(3), 503–514. https://doi.org/10.2307/2336116

Das, S., Datta, S., & Chaudhuri, B. B. (2018). Handling data irregularities in classification: Foundations, trends, and future challenges. Pattern Recognition, 81, 674–693.

Efron, B., & Tibshirani, R. (1993). An introduction to the bootstrap. New York: Chapman and Hall.

Efron, B., & Tibshirani, R. (1997). Improvements on cross-validation: The.632+ bootstrap method. Journal of the American Statistical Association, 92(438), 548–560.

Geisser, S. (1975). The predictive sample reuse method with applications. Journal of the American Statistical Association, 70(350), 320–328. https://doi.org/10.1080/01621459.1975.10479865

Genz, A., Bretz, F., Miwa, T., Mi, X., Leisch, F., Scheipl, F., & Hothorn, T. (2020). mvtnorm: Multivariate normal and t distributions. R package version 1.1-1.

Hairuddin, N. L., Mi Yusuf, L., & Othman, M. S. (2020). Gender classification on skeletal remains: Efficiency of metaheuristic algorithm method and optimized back propagation neural network. Journal of Information and Communication Technology, 19(2), 251–277.

Jamaluddin, A.H & Mahat, N. I. (2019). The effects of resampling methods on linear discriminant analysis for data set with two imbalanced groups: An empirical evidence. Advances and Applications in Statistics, 59(1), 17–42. https://doi.org/10.17654/AS059010017 Journal of ICT, 20, No. 1 (January) 2021, pp: 83-

Japkowicz, N. (2000). Learning from imbalanced data sets: A comparison of various strategies. In AAAI Workshop on Learning from Imbalanced Data Sets (Vol. 68, pp. 10–15).

Kaur, H., Pannu, H. S., & Malhi, A. K. (2019). A systematic review on imbalanced data challenges in machine learning: Applications and solutions. ACM Comput. Surv., 52(4). https://doi.org/10.1145/3343440

Kim, J.-H. (2009). Estimating classification error rate: Repeated cross-validation, repeated hold-out and bootstrap. Computational Statistics & Data Analysis, 53(11), 3735–3745.

Kohavi, R. (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. Proceedings of the 14th International Joint Conference on Artificial Intelligence - Volume 2, 1137–1143. http://dl.acm.org/citation.cfm?id=1643031.1643047

Kuhn, M. (2008). Building predictive models in R using the caret package. Journal of Statistical Software, 28. https://doi.org/10.18637/jss.v028. i05

Roy, S., Ahmed, M., & Akhand, M. A. H. (2018). Noisy image classification using hybrid deep learning methods. Journal of Information and Communication Technology, 17, 233–269.

Xie, J., & Qiu, Z. (2007). The effect of imbalanced data sets on LDA: A theoretical and empirical analysis. Pattern Recognition, 40(2), 557–562. https://doi.org/10.1016/j.patcog.2006.01.009

Xue, J.-H., & Titterington, D. M. (2008). Do unbalanced data have a negative effect on LDA? Pattern Recognition, 41(5), 1575–1588. https://doi.org/10.1016/j.patcog.2007.11.008

Downloads

Published

04-11-2020

How to Cite

Jamaluddin, A. H., & Mahat, N. I. (2020). Validation Assessments on Resampling Method in Imbalanced Binary Classification for Linear Discriminant Analysis. Journal of Information and Communication Technology, 20(1), 83-102. https://doi.org/10.32890/jict2021.20.1.5

Research impact

Harvested 2026-09-06
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2021.20.1.5

Most read articles by the same author(s)