Multilabel Over-Sampling and Under-Sampling with Class Alignment for Imbalanced Multilabel Text Classification

Authors

  • Adil Yaseen Taha Faculty of Information Science and Technology Universiti Kebangsaan Malaysia, Malaysia
  • Sabrina Tiun Faculty of Information Science and Technology Universiti Kebangsaan Malaysia, Malaysia
  • Abdul Hadi Abd Rahman Faculty of Information Science and Technology Universiti Kebangsaan Malaysia, Malaysia
  • Ali Sabah Faculty of Information Science and Technology Universiti Kebangsaan Malaysia, Malaysia

DOI:

https://doi.org/10.32890/jict2021.20.3.6

Keywords:

Data mining, multilabel text classification, class imbalance problem, resampling method, class alignment

Abstract

Simultaneous multiple labelling of documents, also known as multilabel text classification, will not perform optimally if the class is highly imbalanced. Class imbalanced entails skewness in the fundamental data for distribution that leads to more difficulty in classification. Random over-sampling and under-sampling are common approaches to solve the class imbalanced problem. However, these approaches have several drawbacks; the under-sampling is likely to dispose of useful data, whereas the over-sampling can heighten the probability of overfitting. Therefore, a new method that can avoid discarding useful data and overfitting problems is needed. This study proposes a method to tackle the class imbalanced problem by combining multilabel over-sampling and under-sampling with class alignment (ML-OUSCA). In the proposed ML-OUSCA, instead of using all the training instances, it draws a new training set by over-sampling small size classes and under-sampling big size classes. To evaluate our proposed ML-OUSCA, evaluation metrics of average precision, average recall and average F-measure on three benchmark datasets, namely, Reuters-21578, Bibtex, and Enron datasets, were performed. Experimental results showed that the proposed ML-OUSCA outperformed the chosen baseline random resampling approaches; K-means SMOTE and KNN-US. Thus, based on the results, we can conclude that designing a resampling method based on the class imbalanced together with class alignment will improve multilabel classification even better than just the random resampling method.

References

Abdi, L., & Hashemi, S. (2015). To combat multi-class imbalanced problems by means of over-sampling techniques. IEEE transactions on Knowledge and Data Engineering, 28(1), 238–251. https://doi.org/10.1109/TKDE.2015.2458858

Adel, A., Omar, N., Albared, M., & Al-Shabi, A. (2019). Feature selection method based on statistics of compound words for Arabic text classification. International Arab Journal of Information Technology, 16(2), 178–185.

Kermani, F. Z., Eslami, E., & Sadeghi, F. (2019). Global Filter– Wrapper method based on class-dependent correlation for text Journal of ICT, 20, No. 3 (July) 2021, pp: 423– classification. Engineering Applications of Artificial Intelligence, 85, 619-633. https://doi.org/10.1016/j.engappai.2019.07.003

Ali, H., Salleh, M. N. M., Saedudin, R., Hussain, K., & Mushtaq, M. F. (2019). Imbalance class problems in data mining: A review. Indonesian Journal of Electrical Engineering and Computer Science, 14(10.11591). https://doi.org/10.11591/ijeecs.v14. i3.pp1560-1571

Al-Salemi, B., Ayob, M., & Noah, S. A. M. (2018). Feature ranking for enhancing boosting-based multi-label text categorization. Expert Systems with Applications, 113, 531–543. https://doi.org/10.1016/j.eswa.2018.07.024

Amidan, B. G., Ferryman, T. A., & Cooley, S. K. (2005, March). Data outlier detection using the Chebyshev theorem. In 2005 IEEE Aerospace Conference (pp. 3814–3819). IEEE. https://doi.org/10.1109/AERO. 2005.1559688

Boutell, M. R., Luo, J., Shen, X., & Brown, C. M. (2004). Learning multi-label scene classification. Pattern Recognition, 37(9), 1757–1771. https://doi.org/10.1016/j.patcog.2004.03.009

Cascaro, R. J., Gerardo, B. D., & Medina, R. P. (2019, December). Aggregating filter feature selection methods to enhance multiclass text classification. In Proceedings of the 2019 7th International Conference on Information Technology: IoT and Smart City (pp. 80–84). https://doi.org/10.1145/33 77170.3377209

Charte, F., Rivera,A. J., del Jesus, M. J., & Herrera, F. (2015). MLSMOTE: Approaching imbalanced multilabel learning through synthetic instance generation. Knowledge-Based Systems, 89, 385–397. https://doi.org/10.1016/j.knosys.2015.07.019

Charte, F., Rivera, A. J., del Jesus, M. J., & Herrera, F. (2015). Addressing imbalance in multilabel classification: Measures and random resampling algorithms. Neurocomputing, 163, 3–16. https://doi.org/10.1016/j.neucom.2014.08.091

Charte, F., Rivera,A. J., del Jesus, M. J., & Herrera, F. (2015). MLSMOTE: Approaching imbalanced multilabel learning through synthetic instance generation. Knowledge-Based Systems, 89, 385–397. https://doi.org/10.1016/j.knosys.2015.07.019

Charte, F., Rivera, A., del Jesus, M. J., & Herrera, F. (2013, September). A first approach to deal with imbalance in multi-label datasets. In International Conference on Hybrid Artificial Intelligence Systems, (Vol. 8073, pp. 150–160). Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-40846-5_16 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Charte, F., Rivera, A. J., del Jesus, M. J., & Herrera, F. (2019). Dealing with difficult minority labels in imbalanced mutilabel data sets. Neurocomputing, 326, 39–53. https://doi.org/10.1016/j. neucom.201 6.08.158

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613 /jair.953

Chen, E., Lin, Y., Xiong, H., Luo, Q., & Ma, H. (2011). Exploiting probabilistic topic models to improve text categorization under class imbalance. Information Processing & Management, 47(2), 202–214. https://doi.org/10.1016/j. ipm.2010.07.003

Chen, K., Zhang, Z., Long, J., & Zhang, H. (2016). Turning from TF-IDF to TF-IGM for term weighting in text classification. Expert Systems with Applications, 66, 245–260. https://doi.org/10.1016/j.eswa. 2016.09.009

Cheng, W., & Hüllermeier, E. (2009). Combining instance-based learning and logistic regression for multilabel classification. Machine Learning, 76(2–3), 211–225. https://doi.org/10.1007/s10994-009-5127-5

Daniels, Z., & Metaxas, D. (2017, February). Addressing imbalance in multi-label classification using structured Hellinger forests. In Proceedings of the AAAI Conference on Artificial Intelligence, (Vol. 31, No. 1). https://ojs.aaai.org/index.php/ AAAI/article/view/10908

Díez-Pastor, J. F., Rodríguez, J. J., García-Osorio, C., & Kuncheva, L. I. (2015). Random balance: ensembles of variable priors classifiers for imbalanced data. Knowledge-Based Systems, 85, 96–111. https://doi.org/10.1016/j.knosys.2015.04.022

Dubey, R., Zhou, J., Wang, Y., Thompson, P. M., Ye, J., & Alzheimer’s Disease Neuroimaging Initiative. (2014). Analysis of sampling techniques for imbalanced data: An n=648 ADNI study. NeuroImage, 87, 220–241. https://doi.org/10.1016/j. neuroimage.2013.10.005

Fang, M., Xiao, Y., Wang, C., & Xie, J. (2014, November). Multi-label classification: Dealing with imbalance by combining labels. In 2014 IEEE 26th International Conference on Tools with Artificial Intelligence (pp. 233-237). IEEE. https://doi.org/10.1109/ICTAI.2014.42 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Feng, S., Zhao, C., & Fu, P. (2020). A cluster-based hybrid sampling approach for imbalanced data classification. Review of Scientific Instruments, 91(5), 055101. https://doi.org/10.1063/5.0008935

Fernández, A., del Río, S., Chawla, N. V., & Herrera, F. (2017). An insight into imbalanced big data classification: Outcomes and challenges. Complex & Intelligent Systems, 3(2), 105–120. https://doi.org/10.1007/s40747-017-0037-9

Fürnkranz, J., Hüllermeier, E., Mencía, E. L., & Brinker, K. (2008). Multilabel classification via calibrated label ranking. Machine Learning, 73(2), 133–153. https://doi.org/10.1007/s10994-008-5064-8

Galar, M., Fernández, A., Barrenechea, E., & Herrera, F. (2013). EUSBoost: Enhancing ensembles for highly imbalanced datasets by evolutionary undersampling. Pattern recognition, 46(12), 3460–3471. https://doi.org/10.1016/j.patcog.2013.05.006

García, S., Zhang, Z. L., Altalhi, A., Alshomrani, S., & Herrera, F. (2018). Dynamic ensemble selection for multi-class imbalanced datasets. Information Sciences, 445, 22–37. https://doi.org/10.1016/j.ins. 2018. 03.002

Glazkova, A. (2020). A comparison of synthetic oversampling methods for multi-class text classification. arXiv preprint arXiv:2008.04636.

Haixiang, G., Yijing, L., Shang, J., Mingyun, G., Yuanyue, H., & Bing, G. (2017). Learning from class-imbalanced data: Review of methods and applications. Expert Systems with Applications, 73, 220 –239. https://doi.org/10.1016/j.eswa.2016.12.035

Han, H., Wang, W. Y., & Mao, B. H. (2005, August). Borderline-SMOTE: A new over-sampling method in imbalanced data sets learning. In International conference on intelligent computing 3644, 878–887. Springer, Berlin, Heidelberg. https://doi.org/10.1007/11538059_91

He, H., Bai, Y., Garcia, E. A., & Li, S. (2008, June). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. In 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence) (pp. 1322–1328). IEEE. https://doi.org/10.1109/IJCNN.2008. 4633969

Japkowicz, N., & Stephen, S. (2002). The class imbalance problem: A systematic study. Intelligent Data Analysis, 6(5), 429-449. https://doi.org/10.3233/IDA-2002-6504 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Jian, C., Gao, J., & Ao, Y. (2016). A new sampling method for classifying imbalanced data based on support vector machine ensemble. Neurocomputing, 193, 115–122. https://doi.org/10.1016/j.neucom.2016. 02.006

Johnson, R., & Zhang, T. (2014). Effective use of word order for text categorization with convolutional neural networks. arXiv preprint arXiv:1412.1058.

Kim, Y. G., Kwon, Y., & Paik, M. C. (2019). Valid oversampling schemes to handle imbalance. Pattern Recognition Letters, 125, 661–667. https://doi.org/10.1016/j.patrec.2019.07.006

Koziarski, M., Krawczyk, B., & Woźniak, M. (2019). Radial-based oversampling for noisy imbalanced data classification. Neurocomputing, 343, 19–33. https://doi.org/10.1016/j.neucom.2018.04.089

Koziarski, M., Woźniak, M., & Krawczyk, B. (2020). Combined Cleaning and Resampling Algorithm for Multi-Class Imbalanced Data with Label Noise. Knowledge-Based Systems, 204, 106223. arXiv preprint arXiv:2004.03406.

Last, F., Douzas, G., & Bacao, F. (2017). Oversampling for imbalanced learning based on k-means and SMOTE. arXiv preprint arXiv:1711.00837. https://doi.org/10.1016/j.ins.2018.06.056

Li, F., Miao, D., & Pedrycz, W. (2017). Granular multi-label feature selection based on mutual information. Pattern Recognition, 67, 410–423. https://doi.org/10.1016/j.patcog.2017.02.025

Li, H., Zou, P., Han, W. H., & Xia, R. Z. (2014). Imbalanced Data Classification Based on Clustering. In Applied Mechanics and Materials (Vol. 443, pp. 741–745). Trans Tech Publications Ltd. https://doi.org/10.4028/www.scientific.net/AMM.443.741

Lim, H., Lee, J., & Kim, D. W. (2017). Optimization approach for feature selection in multi-label classification. Pattern Recognition Letters, 89, 25–23. https://doi.org/10.1016/j. patrec.2017.02.004

Lin, W. C., Tsai, C. F., Hu, Y. H., & Jhang, J. S. (2017). Clustering-based undersampling in class-imbalanced data. Information Sciences, 409, 17–26. https://doi.org/10.1016/j.ins.2017.05.008

López, V., Fernández, A., García, S., Palade, V., & Herrera, F. (2013). An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics. Information sciences, 250, 113–141. https://doi.org/10.1016/j.ins.2013.07.007 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Maheshwari, S., Jain, R. C., & Jadon, R. S. (2017). A review on class imbalance problem: Analysis and potential solutions. International Journal of Computer Science Issues (IJCSI), 14(6), 43–51. https://doi.org/10.20943/01201706.4351

Mao, X., Chang, S., Shi, J., Li, F., & Shi, R. (2019). Sentiment-aware word embedding for emotion classification. Applied Sciences, 9(7), 1334. https://doi.org/10.3390/app9071334

Mashaan Abed, A. L. I., Tiun, S., & Albared, M. (2013). Arabic term extraction using combined approach on Islamic document. Journal of Theoretical & Applied Information Technology, 58(3).

Mirończuk, M. M., & Protasiewicz, J. (2018). A recent overview of the state-of-the-art elements of text classification. Expert Systems with Applications, 106, 36–54. https://doi.org/10.1016/j.eswa. 2018.03. 058

Moreo, A., Esuli, A., & Sebastiani, F. (2016, July). Distributional random oversampling for imbalanced text classification. In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval (pp. 805–808). https://doi.org/10.1145/2911451.2914722

Onan, A. (2019). Consensus clustering-based undersampling approach to imbalanced learning. Scientific Programming, 2019. https://doi.org/10.1155/2019/5901087

Pant, P., Sabitha, A. S., Choudhury, T., & Dhingra, P. (2019). Multi-label classification trending challenges and approaches. In Emerging Trends in Expert Applications and Security (pp. 433–444). Springer, Singapore. https://doi.org/10.1007/978-981-13-2285-3_51

Patel, H., Singh Rajput, D., Thippa Reddy, G., Iwendi, C., Kashif Bashir, A., & Jo, O. (2020). A review on classification of imbalanced data for wireless sensor networks. International Journal of Distributed Sensor Networks, 16(4), 1550147720916404. https://doi.org/10.1177/1550147720916404

Pereira, R. M., Costa, Y. M., & Silla Jr, C. N. (2020). MLTL: A multi-label approach for the Tomek Link undersampling algorithm. Neurocomputing, 383, 95–105. https://doi.org/10.1016/j.neucom.2019. 11.076

Qiao, L., Zhang, L., Sun, Z., & Liu, X. (2017). Selecting label-dependent features for multi-label classification. Neurocomputing, 259, 112–118. https://doi.org/10.1016/j.neucom.2016.08.122 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Rao, K. N., & Reddy, C. S. (2020). A novel under sampling strategy for efficient software defect analysis of skewed distributed data. Evolving Systems, 11(1), 119–131. https://doi.org/10.1007/ s12530-018-9261-9

Read, J., Pfahringer, B., Holmes, G., & Frank, E. (2011). Classifier chains for multi-label classification. Machine Learning, 85(3), 333. https://doi.org/10.1007/s10994-011-5256-5

Rivera, W. A. (2017). Noise reduction a priori synthetic over-sampling for class imbalanced data sets. Information Sciences, 408, 146–161. https://doi.org/10.1016/j.ins.2017.04.046

Sadhukhan, P., & Palit, S. (2019). Reverse-nearest neighborhood based oversampling for imbalanced, multi-label datasets. Pattern Recognition Letters, 125, 813–820. https://doi.org/10.1016/j. patrec.201 9.08.009

Sáez, J. A., Krawczyk, B., & Woźniak, M. (2016). Analyzing the oversampling of different classes and types of examples in multi-class imbalanced datasets. Pattern Recognition, 57, 164–178. https://doi.org/10.1016/j.patcog.2016.03.012

Sharef, B. T., Omar, N., & Sharef, Z. T. (2014). An automated Arabic text categorization based on the frequency ratio accumulation. International Arab Journal of Information Technology, 11(2), 213–221.

Shi, H., Gao, Q., Ji, S., & Liu, Y. (2018, July). A hybrid sampling method based on safe screening for imbalanced datasets with sparse structure. In 2018 International Joint Conference on Neural Networks (IJCNN) (pp. 1–8). IEEE. https://doi.org/10.1109/ IJCNN.2018.8489569

Song, J., Huang, X., Qin, S., & Song, Q. (2016, June). A bi-directional sampling based on K-means method for imbalance text classification. In 2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS) (pp. 1–5). https://doi.org/IEEE. 10.1109/ICIS.2016.7550920

Su, C. T., & Hsiao, Y. H. (2007). An evaluation of the robustness of MTS for imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 19(10), 1321–1332. https://doi.org/10.1109/ TKDE. 2007.190623

Sun, K. W., & Lee, C. H. (2017). Addressing class-imbalance in multi-label learning via two-stage multi-label hypernetwork. Neurocomputing, 266, 375–389. https://doi.org/10.1016/j.neucom. 2017.05.04 9 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Sun, K. W., Lee, C. H., & Wang, J. (2016). Multilabel classification via co-evolutionary multilabel hypernetwork. IEEE Transactions on Knowledge and Data Engineering, 28(9), 2438–2451. https://doi.org/10.1109/TKDE.2016.2566621

Taha, A. Y., & Tiun, S. (2016). Binary relevance (BR) method classifier of multi-label classification for Arabic text. Journal of Theoretical & Applied Information Technology, 84(3).

Taha, A. Y., Tiun, S., Abd Rahman, A. H., Ayob, M., & Sabah, A. (2020). A dynamic two-Layers MI and clustering-based ensemble feature selection for multi-labels text classification. Journal of Advanced Computer Science and Applications, 11(7).

Tahir, M. A., Kittler, J., & Yan, F. (2012). Inverse random under sampling for class imbalance problem and its application to multi-label classification. Pattern Recognition, 45(10), 3738–3750. https://doi.org/10.1016/j.patcog.2012.03.014

Tanha, J., Abdi, Y., Samadi, N., Razzaghi, N., & Asadpour, M. (2020). Boosting methods for multi-class imbalanced data classification: An experimental review. Journal of Big Data, 7(1), 1–47. https://doi.org/10.1186/s40537-020-00349-y

Toshniwal, D., & Venkoparao, G. (2017). Distributed sparse class-imbalance learning and its applications. IEEE Transactions on Big Data, 13(9). https://doi.org/10.1109/TBDATA.2017. 2688372

Tsoumakas, G., Katakis, I., & Vlahavas, I. (2006, September). A review of multi-label classification methods. In Proceedings of the 2nd ADBIS workshop on data mining and knowledge discovery (ADMKD 2006) (pp. 99–109).

Tsoumakas, G., Katakis, I., & Vlahavas, I. (2010). Random k-labelsets for multilabel classification. IEEE Transactions on Knowledge and Data Engineering, 23(7), 1079–1089. https://doi.org/10.1109/TKDE.2010.164

Wang, K. J., & Adrian, A. M. (2013). Breast cancer classification using hybrid synthetic minority over-sampling technique and artificial immune recognition system algorithm. International Journal Compute Science Electronics Engineering (IJCSEE), 1(3), 408–412.

Weng, W., Lin, Y., Wu, S., Li, Y., & Kang, Y. (2018). Multi-label learning based on label-specific features and local pairwise label correlation. Neurocomputing, 273, 385–394. https://doi.org/10.1016/j.neuc om.2017.07.044 Journal of ICT, 20, No. 3 (July) 2021, pp: 423–

Xu, Z., Shen, D., Nie, T., & Kou, Y. (2020). A hybrid sampling algorithm combining M-SMOTE and ENN based on random forest for medical imbalanced data. Journal of Biomedical Informatics, 107, 103465. https://doi.org/10.1016/j.jbi.2020.103465

Zhang, L., Zhang, C., Quan, S., Xiao, H., Kuang, G., & Liu, L. (2020). A class imbalance loss for imbalanced object recognition. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13, 2778–2792. https://doi.org/10.1109/ JSTARS.2020.2995703

Zhang, M. L., & Wu, L. (2014). Lift: Multi-label learning with label-specific features. IEEE transactions on pattern analysis and machine intelligence, 37(1), 107–120. https://doi.org/10.1109/ TPAMI.2014.2339815

Zhang, M. L., & Zhou, Z. H. (2007). ML-KNN: A lazy learning approach to multi-label learning. Pattern recognition, 40(7), 2038–2048. https://doi.org/10.1016/j.patcog.2006.12.019

Zhang, M. L., Li, Y. K., Yang, H., & Liu, X. Y. (2020). Towards class-imbalance aware multi-label learning. IEEE Transactions on Cybernetics. https://doi.org/10.1109/TCYB.2020.3027509

Zhang, Y., Liu, G., Luan, W., Yan, C., & Jiang, C. (2018, March). An approach to class imbalance problem based on stacking and inverse random under sampling methods. In 2018 IEEE 15th International Conference on Networking, Sensing and Control (ICNSC) (pp. 1–6). IEEE. https://doi.org/10.1109/ ICNSC.2018.8361344

Zhou, S., Li, X., Dong, Y., & Xu, H. (2020). A decoupling and bidirectional resampling method for multilabel classification of imbalanced data with label concurrence. Scientific Programming, 2020. https://doi.org/10.1155/2020/8829432

Zubiaga, A. (2020). Exploiting class labels to boost performance on embedding-based text classification. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (pp. 3357–3360) arXiv preprint arXiv:2006.02104.

Downloads

Published

11-06-2021

How to Cite

Taha, A. Y., Tiun, S., Abd Rahman, A. H., & Sabah, A. (2021). Multilabel Over-Sampling and Under-Sampling with Class Alignment for Imbalanced Multilabel Text Classification. Journal of Information and Communication Technology, 20(3), 423-456. https://doi.org/10.32890/jict2021.20.3.6

Research impact

Harvested 2026-09-06
33 citations, from Scopus — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2021.20.3.6 OpenAlex W3175414601 Scopus 85109463524