Comparative Performance Evaluation of Efficiency for High Dimensional Classification Methods

Authors

  • Friday Zinzendoff Okwonu Department of Mathematics, Faculty of Science, Delta State University, Nigeria
  • Nor Aishah Ahad School of Quantitative Sciences, College of Arts and Sciences, Universiti Utara Malaysia, Malaysia
  • Nicholas Oluwole Ogini Department of Computer Science, Delta State University, Nigeria
  • Innocent Ejiro Okoloko Faculty of Computing, Dennis Osadebay University, Nigeria
  • Wan Zakiyatussariroh Wan Husin Faculty of Computer and Mathematical Science, Universiti Teknologi MARA, Kelantan Branch, Malaysia

DOI:

https://doi.org/10.32890/jict2022.21.3.6

Keywords:

Classification, confusion matrix, efficiency, high-dimensional data, robustness

Abstract

This paper aimed to determine the efficiency of classifiers for high-dimensional classification methods. It also investigated whether an extreme minimum misclassification rate translates into robust efficiency. To ensure an acceptable procedure, a benchmark evaluation threshold (BETH) was proposed as a metric to analyze the comparative performance for high-dimensional classification methods. A simplified performance metric was derived to show the efficiency of different classification methods. To achieve the objectives, the existing probability of correct classification (PCC) or classification accuracy reported in five different articles was used to generate the BETH value. Then, a comparative analysis was performed between the application of BETH value and the well-established PCC value ,derived from the confusion matrix. The analysis indicated that the BETH procedure had a minimum misclassification rate, unlike the Optimal method. The results also revealed that as the PCC inclined toward unity value, the misclassification rate between the two methods (BETH and PCC) became extremely irrelevant. The study revealed that the BETH method was invariant to the performance established by the classifiers using the PCC criterion but demonstrated more relevant aspects of robustness and minimum misclassification rate as compared to the PCC method. In addition, the comparative analysis affirmed that the BETH method exhibited more robust efficiency than the Optimal method. The study concluded that a minimum misclassification rate yields robust performance efficiency.

References

Bickel, P. J., & Doksum, K. A. (2015). Mathematical statistics: Basic ideas and selected topics, second edition. In Mathematical Statistics: Basic Ideas and Selected Topics, Second Edition (Vol. 1). https://doi.org/10.1201/b18312

Bielza, C., Li, G., & Larrañaga, P. (2011). Multi-dimensional classification with Bayesian networks. International Journal of Approximate Reasoning, 52(6), 705–727. https://doi.org/10.1016/j.ijar.2011.01.007

Blagus, R., & Lusa, L. (2010). Class prediction for high-dimensional class-imbalanced data. BMC Bioinformatics, 11(1), 1–17. https://doi.org/10.1186/1471-2105-11-523 Journal of ICT, 21, No. 3 (July) 2022, pp: 437–

Blagus, R., & Lusa, L. (2013). SMOTE for high-dimensional classimbalanced data. BMC Bioinformatics, 14(1), 1–13. https://doi.org/10.1186/1471-2105-14-106

Boutell, M. R., Luo, J., Shen, X., & Brown, C. M. (2004). Learning multi-label scene classification. Pattern Recognition, 37(9), 1757–1771. doi:10.1016/j.patcog.2004.03.009

Bradley, A. P. (1997). The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognition, 30(7), 1145-1159. https://doi.org/10.1016/S00313203(96)00142-2

Bühlmann, P., & Geer, S. van de. (2011). Statistics for high-dimensional data: Methods, theory, and applications. In Springer Series in Statistics. Casanova, R., Whitlow, C. T., Wagner, B., Williamson, J., Shumaker, S.

A., Maldjian, J. A., & Espeland, M. A. (2011). High dimensional classification of structural MRI Alzheimer’s disease data based on large-scale regularization. Frontiers in Neuroinformatics, 5, 22. https://doi.org/10.3389/fninf.2011.00022

Chuang, L. Y., Chang, H. W., Tu, C. J., & Yang, C. H. (2008). Improved binary PSO for feature selection using gene expression data. Computational Biology and Chemistry, 32(1), 29–38. https://doi.org/10.1016/j.compbiolchem.2007.09.005

Croux, C., Filzmoser, P., & Joossens, K. (2008). Classification efficiencies for robust linear discriminant analysis. Statistica Sinica, 18(2), 581–599. http://www.jstor.org/stable/24308496

Dua, D., & Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science.

Elisseeff, A., & Weston, J. (2002). A kernel method for multi-labelled classification. In T. Dietterich, S. Becker, Z. Ghahramani (Eds.), Advances in Neural Information Processing Systems, vol. 14, MIT Press, pp. 681–687. ttps://proceedings.neurips.cc/ paper/2001/file/39dcaf7a053dc372fbc391d4e6b5d693-Paper. pdf Fernandes, J. A., Lozano, J. A., Inza, I., Irigoien, X., Pérez, A., &

Rodríguez, J. D. (2013). Supervised pre-processing approaches in multiple class variables classification for fish recruitment forecasting. Environmental Modelling and Software, 40, 245–254. https://doi.org/10.1016/j.envsoft.2012.10.001

Ferraty, F. (2010). High-dimensional data: A fascinating statistical challenge. Journal of Multivariate Analysis, 101(2), 305–306. https://doi.org/10.1016/j.jmva.2009.10.012 Journal of ICT, 21, No. 3 (July) 2022, pp: 437–

Fürnkranz, J., Hüllermeier, E., Loza Mencía, E., & Brinker, K. (2008). Multilabel classification via calibrated label ranking. Machine Learning, 73(2). 133–153. https://doi.org/10.1007/s10994-0085064-8

Ghosh, A., SahaRay, R., Chakrabarty, S., & Bhadra, S. (2021). Robust generalized quadratic discriminant analysis. Pattern Recognition, 117, 107981. https://doi.org/10.1016/j.patcog.2021.107981

Ghosh, J. K. (2012). Statistics for high-dimensional data: Methods, theory, and applications by Peter Bühlmann, Sara van de Geer. International Statistical Review, 80(3), 486-487. https://doi.org/10.1111/j.1751-5823.2012.00196_18.x

Gibaja, E. (2013). A tutorial on multi-label learning. ACM Computing Surveys, 9, 1–38. https://doi.org/10.1145/2716262

Gil-Begue, S., Bielza, C., & Larrañaga, P. (2021). Multi-dimensional Bayesian network classifiers: A survey. Artificial Intelligence Review, 54(1), 519–559. https://doi.org/10.1007/s10462-02009858-x

Guo, B., Damper, R. I., Gunn, S. R., & Nelson, J. D. B. (2008). A fast separability-based feature-selection method for highdimensional remotely sensed image classification. Pattern Recognition, 41(5), 1653–1662. https://doi.org/10.1016/j. patcog.2007.11.007

Hamilton, W. C. (1970). The revolution in crystallography. Science, 169(3941), 133–141. https://doi.org/10.1126/science.169.3941.133 Hu, B., Dai, Y., Su, Y., Moore, P., Zhang, X., Mao, C., Chen, J., & Xu,

L. (2018). Feature selection for optimized high-dimensional biomedical data using an improved shuffled frog leaping algorithm. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 15(6), 1765–1773. https://doi.org/10.1109/

TCBB.2016.2602263

Hüllermeier, E., Fürnkranz, J., Cheng, W., & Brinker, K. (2008). Label ranking by learning pairwise preferences. Artificial Intelligence, 172(16–17), 1897–1916. https://doi.org/10.1016/j. artint.2008.08.002

Jimoh, R. G., Abisoye, O. A., & Uthman, M. M. B. (2022). Ensemble feed-forward neural network and support vector machine for prediction of multiclass malaria infection. Journal of Information and Communication Technology, 21(1), 117–148. https://doi.org/10.32890/jict2022.21.1.6.

Johnson, R. A., & Wichern, D. W. (1992). Applied multivariate statistical analysis (3rd ed.). Prentice-Hall, Inc, Englewood Cliffs. Journal of ICT, 21, No. 3 (July) 2022, pp: 437–

Johnstone, I. M., & Titterington, D. M. (2009). Statistical challenges of high-dimensional data. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering

Sciences, 367(1906), 4237–4253. https://doi.org/10.1098/ rsta.2009.0159

Kim, T. K., & Kittler, J. (2005). Locally linear discriminant analysis for multimodally distributed classes for face recognition with a single model image. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(3), 318–327. https://doi.org/10.1109/

TPAMI.2005.58 Kranenburg, R. F., Peroni, D., Affourtit, S., Westerhuis, J. A., Smilde,

A. K., & van Asten, A. C. (2020). Revealing hidden information in GC–MS spectra from isomeric drugs: Chemometricsbased identification from 15 eV and 70 eV EI mass spectra. Forensic Chemistry, 18, 100225. https://doi.org/10.1016/j. forc.2020.100225

Li, J., & Liu, H. (2004). Kent Ridge Biomedical Data Set Repository. School of Computer Engineering, Nanyang Technological University, Singapore. http://datam.i2r.astar.edu.sg/datasets/ krbd/ index.html.

Lin, W. J., & Chen, J. J. (2013). Class-imbalanced classifiers for highdimensional data. Briefings in Bioinformatics, 14(1), 13–26. https://doi.org/10.1093/bib/bbs006

Luo, L., & Li, L. (2014). Defining and evaluating classification algorithms for high-dimensional data based on latent topics. PLoS ONE, 9(1), e82119. https://doi.org/10.1371/journal. pone.0082119

Michiels, S., Koscielny, S., & Hill, C. (2005). Prediction of cancer outcome with microarrays: A multiple random validation strategy. Lancet, 365(9458), 488–492. https://doi.org/10.1016/ S0140-6736(05)17866-0

Okwonu, F. Z. (2013). Several robust techniques in two-groups unbiased linear classification (Unpublished Doctoral Dissertation). Universiti Sains Malaysia, Malaysia.

Okwonu, F. Z., & Othman, A. R. (2013a). Heteroscedastic variancecovariance matrices for unbiased two groups linear classification methods. Applied Mathematical Sciences, 7(138), 6855–6865. https://doi.org/10.12988/ams.2013.39486

Okwonu, F. Z., & Othman, A. R. (2013b). Robust fisher linear classification technique for two groups. World Applied Sciences Journal, 21(Special Issue 1). https://doi.org/10.5829/idosi. Journal of ICT, 21, No. 3 (July) 2022, pp: 437– wasj.2013.21.mae.99939

Penenberg, D. N. (2016). Mathematical statistics: Basic ideas and selected topics (2nd ed., Vols. I and II). P. J. Bickel and K.

A. Doksum, 2015 Boca Raton, Chapman and Hall-CRC xxii + 548 pp., $99.95 (vol. I); 438 pp., $99.95 (vol. II) ISBN 978-1-498-72380-0. Journal of the Royal Statistical Society: Series A (Statistics in Society), 179(4), 1128–1129. https://doi.org/10.1111/rssa.12217

Provost, F., Fawcett, T., & Kohavi, R. (1998, July). The case against accuracy estimation for comparing induction algorithms. In International Conference on Machine Learning (ICML) (Vol. 98, pp. 445–453).

Saadati, M., & Benner, A. (2014). Statistical challenges of highdimensional methylation data. Statistics in Medicine, 33(30), 5347–5357. https://doi.org/10.1002/sim.6251

Trohidis, K., Tsoumakas, G., Kalliris, G., & Vlahavas, I. P. (2008, September). Multi-label classification of music into emotions. In Proceedings of the Ninth International Society for Music Information Retrieval Conference (ISMIR) (Vol. 8, pp. 325–330).

Vidaurre, D. (2020). The statistical challenge of finding spontaneous changes in functional connectivity in high-dimensional fMRI data. In bioRxiv. https://doi.org/10.1101/2020.12.15.422845

Wang, X. Z., Wang, R., & Xu, C. (2018). Discovering the relationship between generalization and uncertainty by incorporating complexity of classification. IEEE Transactions on Cybernetics, 48(2), 703–715. https://doi.org/10.1109/TCYB.2017.2653223 Yan, C., Chang, X., Luo, M., Zheng, Q., Zhang, X., Li, Z., & Nie, F. (2021). Self-weighted robust LDA for multiclass classification with edge classes. ACM Transactions on Intelligent Systems and Technology, 12(1), 1–19. https://doi.org/10.1145/3418284

Yu, L., & Liu, H. (2003, August). Efficiently handling feature redundancy in high-dimensional data. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 685–690). https://doi.org/10.1145/956750.956840

Zhang, B., & Cao, P. (2019). Classification of high-dimensional biomedical data based on feature selection using redundant removal. PLoS ONE, 14(4), e0214406. https://doi.org/10.1371/ journal.pone.0214406

Zhang, C., Zhou, Y., Guo, J., Wang, G., & Wang, X. (2019). Research on classification method of high-dimensional class-imbalanced datasets based on SVM. International Journal of Machine Journal of ICT, 21, No. 3 (July) 2022, pp: 437– Learning and Cybernetics, 10(7), 1765–1778. https://doi.org/10.1007/s13042-018-0853-2

Zhu, S., Ji, X., Xu, W., & Gong, Y. (2005, August). Multi-labeled classification using the maximum entropy method. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 274–281). https://doi.org/10.1145/1076034.1076082

Downloads

Published

17-07-2022

How to Cite

Friday Zinzendoff Okwonu, Prof. Madya Dr Nor Aishah Ahad, Nicholas Oluwole Ogini, Innocent Ejiro Okoloko, & Wan Zakiyatussariroh Wan Husin. (2022). Comparative Performance Evaluation of Efficiency for High Dimensional Classification Methods. Journal of Information and Communication Technology, 21(3), 437-464. https://doi.org/10.32890/jict2022.21.3.6

Research impact

Harvested 2026-09-05
4 citations, from OpenAlex — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2022.21.3.6 OpenAlex W4285803790 Scopus 85134759869