Comparative Performance Evaluation of Efficiency for High Dimensional Classification Methods
DOI:
https://doi.org/10.32890/jict2022.21.3.6Keywords:
Classification, confusion matrix, efficiency, high-dimensional data, robustnessAbstract
This paper aimed to determine the efficiency of classifiers for high-dimensional classification methods. It also investigated whether an extreme minimum misclassification rate translates into robust efficiency. To ensure an acceptable procedure, a benchmark evaluation threshold (BETH) was proposed as a metric to analyze the comparative performance for high-dimensional classification methods. A simplified performance metric was derived to show the efficiency of different classification methods. To achieve the objectives, the existing probability of correct classification (PCC) or classification accuracy reported in five different articles was used to generate the BETH value. Then, a comparative analysis was performed between the application of BETH value and the well-established PCC value ,derived from the confusion matrix. The analysis indicated that the BETH procedure had a minimum misclassification rate, unlike the Optimal method. The results also revealed that as the PCC inclined toward unity value, the misclassification rate between the two methods (BETH and PCC) became extremely irrelevant. The study revealed that the BETH method was invariant to the performance established by the classifiers using the PCC criterion but demonstrated more relevant aspects of robustness and minimum misclassification rate as compared to the PCC method. In addition, the comparative analysis affirmed that the BETH method exhibited more robust efficiency than the Optimal method. The study concluded that a minimum misclassification rate yields robust performance efficiency.
References
Bickel, P. J., & Doksum, K. A. (2015). Mathematical statistics: Basic ideas and selected topics, second edition. In Mathematical Statistics: Basic Ideas and Selected Topics, Second Edition (Vol. 1). https://doi.org/10.1201/b18312
Bielza, C., Li, G., & Larrañaga, P. (2011). Multi-dimensional classification with Bayesian networks. International Journal of Approximate Reasoning, 52(6), 705–727. https://doi.org/10.1016/j.ijar.2011.01.007
Blagus, R., & Lusa, L. (2010). Class prediction for high-dimensional class-imbalanced data. BMC Bioinformatics, 11(1), 1–17. https://doi.org/10.1186/1471-2105-11-523 Journal of ICT, 21, No. 3 (July) 2022, pp: 437–
Blagus, R., & Lusa, L. (2013). SMOTE for high-dimensional classimbalanced data. BMC Bioinformatics, 14(1), 1–13. https://doi.org/10.1186/1471-2105-14-106
Boutell, M. R., Luo, J., Shen, X., & Brown, C. M. (2004). Learning multi-label scene classification. Pattern Recognition, 37(9), 1757–1771. doi:10.1016/j.patcog.2004.03.009
Bradley, A. P. (1997). The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognition, 30(7), 1145-1159. https://doi.org/10.1016/S00313203(96)00142-2
Bühlmann, P., & Geer, S. van de. (2011). Statistics for high-dimensional data: Methods, theory, and applications. In Springer Series in Statistics. Casanova, R., Whitlow, C. T., Wagner, B., Williamson, J., Shumaker, S.
A., Maldjian, J. A., & Espeland, M. A. (2011). High dimensional classification of structural MRI Alzheimer’s disease data based on large-scale regularization. Frontiers in Neuroinformatics, 5, 22. https://doi.org/10.3389/fninf.2011.00022
Chuang, L. Y., Chang, H. W., Tu, C. J., & Yang, C. H. (2008). Improved binary PSO for feature selection using gene expression data. Computational Biology and Chemistry, 32(1), 29–38. https://doi.org/10.1016/j.compbiolchem.2007.09.005
Croux, C., Filzmoser, P., & Joossens, K. (2008). Classification efficiencies for robust linear discriminant analysis. Statistica Sinica, 18(2), 581–599. http://www.jstor.org/stable/24308496
Dua, D., & Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science.
Elisseeff, A., & Weston, J. (2002). A kernel method for multi-labelled classification. In T. Dietterich, S. Becker, Z. Ghahramani (Eds.), Advances in Neural Information Processing Systems, vol. 14, MIT Press, pp. 681–687. ttps://proceedings.neurips.cc/ paper/2001/file/39dcaf7a053dc372fbc391d4e6b5d693-Paper. pdf Fernandes, J. A., Lozano, J. A., Inza, I., Irigoien, X., Pérez, A., &
Rodríguez, J. D. (2013). Supervised pre-processing approaches in multiple class variables classification for fish recruitment forecasting. Environmental Modelling and Software, 40, 245–254. https://doi.org/10.1016/j.envsoft.2012.10.001
Ferraty, F. (2010). High-dimensional data: A fascinating statistical challenge. Journal of Multivariate Analysis, 101(2), 305–306. https://doi.org/10.1016/j.jmva.2009.10.012 Journal of ICT, 21, No. 3 (July) 2022, pp: 437–
Fürnkranz, J., Hüllermeier, E., Loza Mencía, E., & Brinker, K. (2008). Multilabel classification via calibrated label ranking. Machine Learning, 73(2). 133–153. https://doi.org/10.1007/s10994-0085064-8
Ghosh, A., SahaRay, R., Chakrabarty, S., & Bhadra, S. (2021). Robust generalized quadratic discriminant analysis. Pattern Recognition, 117, 107981. https://doi.org/10.1016/j.patcog.2021.107981
Ghosh, J. K. (2012). Statistics for high-dimensional data: Methods, theory, and applications by Peter Bühlmann, Sara van de Geer. International Statistical Review, 80(3), 486-487. https://doi.org/10.1111/j.1751-5823.2012.00196_18.x
Gibaja, E. (2013). A tutorial on multi-label learning. ACM Computing Surveys, 9, 1–38. https://doi.org/10.1145/2716262
Gil-Begue, S., Bielza, C., & Larrañaga, P. (2021). Multi-dimensional Bayesian network classifiers: A survey. Artificial Intelligence Review, 54(1), 519–559. https://doi.org/10.1007/s10462-02009858-x
Guo, B., Damper, R. I., Gunn, S. R., & Nelson, J. D. B. (2008). A fast separability-based feature-selection method for highdimensional remotely sensed image classification. Pattern Recognition, 41(5), 1653–1662. https://doi.org/10.1016/j. patcog.2007.11.007
Hamilton, W. C. (1970). The revolution in crystallography. Science, 169(3941), 133–141. https://doi.org/10.1126/science.169.3941.133 Hu, B., Dai, Y., Su, Y., Moore, P., Zhang, X., Mao, C., Chen, J., & Xu,
L. (2018). Feature selection for optimized high-dimensional biomedical data using an improved shuffled frog leaping algorithm. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 15(6), 1765–1773. https://doi.org/10.1109/
TCBB.2016.2602263
Hüllermeier, E., Fürnkranz, J., Cheng, W., & Brinker, K. (2008). Label ranking by learning pairwise preferences. Artificial Intelligence, 172(16–17), 1897–1916. https://doi.org/10.1016/j. artint.2008.08.002
Jimoh, R. G., Abisoye, O. A., & Uthman, M. M. B. (2022). Ensemble feed-forward neural network and support vector machine for prediction of multiclass malaria infection. Journal of Information and Communication Technology, 21(1), 117–148. https://doi.org/10.32890/jict2022.21.1.6.
Johnson, R. A., & Wichern, D. W. (1992). Applied multivariate statistical analysis (3rd ed.). Prentice-Hall, Inc, Englewood Cliffs. Journal of ICT, 21, No. 3 (July) 2022, pp: 437–
Johnstone, I. M., & Titterington, D. M. (2009). Statistical challenges of high-dimensional data. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering
Sciences, 367(1906), 4237–4253. https://doi.org/10.1098/ rsta.2009.0159
Kim, T. K., & Kittler, J. (2005). Locally linear discriminant analysis for multimodally distributed classes for face recognition with a single model image. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(3), 318–327. https://doi.org/10.1109/
TPAMI.2005.58 Kranenburg, R. F., Peroni, D., Affourtit, S., Westerhuis, J. A., Smilde,
A. K., & van Asten, A. C. (2020). Revealing hidden information in GC–MS spectra from isomeric drugs: Chemometricsbased identification from 15 eV and 70 eV EI mass spectra. Forensic Chemistry, 18, 100225. https://doi.org/10.1016/j. forc.2020.100225
Li, J., & Liu, H. (2004). Kent Ridge Biomedical Data Set Repository. School of Computer Engineering, Nanyang Technological University, Singapore. http://datam.i2r.astar.edu.sg/datasets/ krbd/ index.html.
Lin, W. J., & Chen, J. J. (2013). Class-imbalanced classifiers for highdimensional data. Briefings in Bioinformatics, 14(1), 13–26. https://doi.org/10.1093/bib/bbs006
Luo, L., & Li, L. (2014). Defining and evaluating classification algorithms for high-dimensional data based on latent topics. PLoS ONE, 9(1), e82119. https://doi.org/10.1371/journal. pone.0082119
Michiels, S., Koscielny, S., & Hill, C. (2005). Prediction of cancer outcome with microarrays: A multiple random validation strategy. Lancet, 365(9458), 488–492. https://doi.org/10.1016/ S0140-6736(05)17866-0
Okwonu, F. Z. (2013). Several robust techniques in two-groups unbiased linear classification (Unpublished Doctoral Dissertation). Universiti Sains Malaysia, Malaysia.
Okwonu, F. Z., & Othman, A. R. (2013a). Heteroscedastic variancecovariance matrices for unbiased two groups linear classification methods. Applied Mathematical Sciences, 7(138), 6855–6865. https://doi.org/10.12988/ams.2013.39486
Okwonu, F. Z., & Othman, A. R. (2013b). Robust fisher linear classification technique for two groups. World Applied Sciences Journal, 21(Special Issue 1). https://doi.org/10.5829/idosi. Journal of ICT, 21, No. 3 (July) 2022, pp: 437– wasj.2013.21.mae.99939
Penenberg, D. N. (2016). Mathematical statistics: Basic ideas and selected topics (2nd ed., Vols. I and II). P. J. Bickel and K.
A. Doksum, 2015 Boca Raton, Chapman and Hall-CRC xxii + 548 pp., $99.95 (vol. I); 438 pp., $99.95 (vol. II) ISBN 978-1-498-72380-0. Journal of the Royal Statistical Society: Series A (Statistics in Society), 179(4), 1128–1129. https://doi.org/10.1111/rssa.12217
Provost, F., Fawcett, T., & Kohavi, R. (1998, July). The case against accuracy estimation for comparing induction algorithms. In International Conference on Machine Learning (ICML) (Vol. 98, pp. 445–453).
Saadati, M., & Benner, A. (2014). Statistical challenges of highdimensional methylation data. Statistics in Medicine, 33(30), 5347–5357. https://doi.org/10.1002/sim.6251
Trohidis, K., Tsoumakas, G., Kalliris, G., & Vlahavas, I. P. (2008, September). Multi-label classification of music into emotions. In Proceedings of the Ninth International Society for Music Information Retrieval Conference (ISMIR) (Vol. 8, pp. 325–330).
Vidaurre, D. (2020). The statistical challenge of finding spontaneous changes in functional connectivity in high-dimensional fMRI data. In bioRxiv. https://doi.org/10.1101/2020.12.15.422845
Wang, X. Z., Wang, R., & Xu, C. (2018). Discovering the relationship between generalization and uncertainty by incorporating complexity of classification. IEEE Transactions on Cybernetics, 48(2), 703–715. https://doi.org/10.1109/TCYB.2017.2653223 Yan, C., Chang, X., Luo, M., Zheng, Q., Zhang, X., Li, Z., & Nie, F. (2021). Self-weighted robust LDA for multiclass classification with edge classes. ACM Transactions on Intelligent Systems and Technology, 12(1), 1–19. https://doi.org/10.1145/3418284
Yu, L., & Liu, H. (2003, August). Efficiently handling feature redundancy in high-dimensional data. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 685–690). https://doi.org/10.1145/956750.956840
Zhang, B., & Cao, P. (2019). Classification of high-dimensional biomedical data based on feature selection using redundant removal. PLoS ONE, 14(4), e0214406. https://doi.org/10.1371/ journal.pone.0214406
Zhang, C., Zhou, Y., Guo, J., Wang, G., & Wang, X. (2019). Research on classification method of high-dimensional class-imbalanced datasets based on SVM. International Journal of Machine Journal of ICT, 21, No. 3 (July) 2022, pp: 437– Learning and Cybernetics, 10(7), 1765–1778. https://doi.org/10.1007/s13042-018-0853-2
Zhu, S., Ji, X., Xu, W., & Gong, Y. (2005, August). Multi-labeled classification using the maximum entropy method. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 274–281). https://doi.org/10.1145/1076034.1076082
Published
Issue
Section
License
Copyright (c) 2022 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-05Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















