Validation Assessments on Resampling Method in Imbalanced Binary Classification for Linear Discriminant Analysis
DOI:
https://doi.org/10.32890/jict2021.20.1.5Keywords:
Linear discriminant analysis, pre-processing, resampling method, class imbalance, binary classificationAbstract
The curse of class imbalance affects the performance of many conventional classification algorithms including linear discriminant analysis (LDA). The data pre-processing approach through some resampling methods such as random oversampling (ROS) and random undersampling (RUS) is one of the treatments to alleviate such curse. Previous studies have attempted to address the effect of a resampling method on the performance of LDA. However, some studies contradicted with each other based on different performance measures as well as validation strategies. This manuscript attempted to shed more light on the effect of a resampling method (ROS or RUS) on the performance of LDA based on true positive rate and true negative rate through five validation strategies, i.e. leave-one-out cross-validation, k-fold cross-validation, repeated k-fold cross-validation, naive bootstrap, and .632+ bootstrap. 100 two-group bivariate normally distributed simulated and four real data sets with severe class imbalance ratio were utilised. The analysis on the location and dispersion statistics of the performance measures was further enlightened on: (i) the effect of a resampling method on the performance of LDA, and (ii) the enhancement in the learning fairness of LDA on objects regardless of sample size, hence reducing the effect of the curse of class imbalance.
References
Alcalá-Fdez, J., Fernández, A., Luengo, J., Derrac, J., García, S., Sánchez, L., & Herrera, F. (2011). KEEL data-mining software tool: Data set repository, integration of algorithms and experimental analysis framework. J. of Mult.-Valued Logic & Soft Computing, 17, 255–287.
Arlot, S., & Celisse, A. (2010). A survey of cross-validation procedures for model selection. Statistics Surveys, 4, 40–79. https://doi.org/10.1214/09-SS054
Branco, P., Torgo, L., & Ribeiro, P. R. (2016). A survey of predictive modeling on imbalanced domains. ACM Computing Surveys, 49(2), 31:1–31:50. https://doi.org/10.1145/2907070
Burman, P. (1989). A comparative study of ordinary cross-validation, v-fold cross-validation and the repeated learning-testing methods. Biometrika, 76(3), 503–514. https://doi.org/10.2307/2336116
Das, S., Datta, S., & Chaudhuri, B. B. (2018). Handling data irregularities in classification: Foundations, trends, and future challenges. Pattern Recognition, 81, 674–693.
Efron, B., & Tibshirani, R. (1993). An introduction to the bootstrap. New York: Chapman and Hall.
Efron, B., & Tibshirani, R. (1997). Improvements on cross-validation: The.632+ bootstrap method. Journal of the American Statistical Association, 92(438), 548–560.
Geisser, S. (1975). The predictive sample reuse method with applications. Journal of the American Statistical Association, 70(350), 320–328. https://doi.org/10.1080/01621459.1975.10479865
Genz, A., Bretz, F., Miwa, T., Mi, X., Leisch, F., Scheipl, F., & Hothorn, T. (2020). mvtnorm: Multivariate normal and t distributions. R package version 1.1-1.
Hairuddin, N. L., Mi Yusuf, L., & Othman, M. S. (2020). Gender classification on skeletal remains: Efficiency of metaheuristic algorithm method and optimized back propagation neural network. Journal of Information and Communication Technology, 19(2), 251–277.
Jamaluddin, A.H & Mahat, N. I. (2019). The effects of resampling methods on linear discriminant analysis for data set with two imbalanced groups: An empirical evidence. Advances and Applications in Statistics, 59(1), 17–42. https://doi.org/10.17654/AS059010017 Journal of ICT, 20, No. 1 (January) 2021, pp: 83-
Japkowicz, N. (2000). Learning from imbalanced data sets: A comparison of various strategies. In AAAI Workshop on Learning from Imbalanced Data Sets (Vol. 68, pp. 10–15).
Kaur, H., Pannu, H. S., & Malhi, A. K. (2019). A systematic review on imbalanced data challenges in machine learning: Applications and solutions. ACM Comput. Surv., 52(4). https://doi.org/10.1145/3343440
Kim, J.-H. (2009). Estimating classification error rate: Repeated cross-validation, repeated hold-out and bootstrap. Computational Statistics & Data Analysis, 53(11), 3735–3745.
Kohavi, R. (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. Proceedings of the 14th International Joint Conference on Artificial Intelligence - Volume 2, 1137–1143. http://dl.acm.org/citation.cfm?id=1643031.1643047
Kuhn, M. (2008). Building predictive models in R using the caret package. Journal of Statistical Software, 28. https://doi.org/10.18637/jss.v028. i05
Roy, S., Ahmed, M., & Akhand, M. A. H. (2018). Noisy image classification using hybrid deep learning methods. Journal of Information and Communication Technology, 17, 233–269.
Xie, J., & Qiu, Z. (2007). The effect of imbalanced data sets on LDA: A theoretical and empirical analysis. Pattern Recognition, 40(2), 557–562. https://doi.org/10.1016/j.patcog.2006.01.009
Xue, J.-H., & Titterington, D. M. (2008). Do unbalanced data have a negative effect on LDA? Pattern Recognition, 41(5), 1575–1588. https://doi.org/10.1016/j.patcog.2007.11.008
Published
Issue
Section
How to Cite
Research impact
Harvested 2026-09-06Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
- Crossref 0 View →
- Scopus 0 not indexed View →
- Google Scholar no free count Search →
- Dimensions no free count Search →
2002 - 2020






















