Multi-Class Multi-Level Classification of Mental Health Disorders Based on Textual Data from Social Media
DOI:
https://doi.org/10.32890/jict2024.23.1.4Keywords:
MCML classification, mental health disorders, Reddit, text mining, transfer learningAbstract
Mental health disorders pose a significant global public health challenge. Social media data provides insights into these conditions.
Analysing text can help identify indications of mental health disorders through text-based analysis. However, despite the large
number of studies on the analysis of mental health disorders, the predominant algorithm in the existing literature is the Multi-Class
Single-Level (MCSL) classification algorithm, which is often used for simple classification tasks involving a limited number of classes.
Typically, these classes are binary, representing either an unhealthy or a healthy mental state. This paper uses English text data from Reddit to classify mental health disorders. The Multi-Class Multi-Level (MCML) classification algorithm was applied to perform detailed
classification and address the limitations of the research scope using several approaches, including machine learning, deep learning, and transfer learning approaches. Two different pre-processing scenarios were proposed to handle unstructured text data, one of the most challenging aspects of classifying text from social media. The results of the experiments show that the MCML classification algorithm successfully performs detailed classification and produces promising results for each classification level. The proposed pre-processing scenario influences the performance of each classifier and improves classification accuracy. The best accuracy results were obtained for the Robustly Optimised BERT Pre-training Approach (RoBERTa) classifier at level 1 and level 2 classifications, namely 0.98 and 0.85, respectively. Overall, the MCML classification algorithm is proven to be used as a benchmark for early detection of text-based mental health disorders.
References
Al Hamoud, A., Hoenig, A., & Roy, K. (2022). Sentence subjectivity analysis of a political and ideological debate dataset using LSTM and BiLSTM with attention and GRU models. Journal of King Saud University - Computer and Information Sciences, 34(10), 7974–7987. https://doi.org/10.1016/j.jksuci.2022.07.014
Ameer, I., Arif, M., Sidorov, G., Gòmez-Adorno, H., & Gelbukh, A. (2022). Mental illness classification on social media texts using deep learning and transfer learning. http://arxiv.org/ abs/2207.01012
Amurark. (2021). Reddit Dataset. Github. https://github.com/ amurark/mental-health-classification
Bae, Y. J., Shim, M., & Lee, W. H. (2021). Schizophrenia detection using machine learning approach from social media content. Sensors, 21(17), 5924. https://doi.org/10.3390/s21175924
Casola, S., Lauriola, I., & Lavelli, A. (2022). Pre-trained transformers: An empirical comparison. Machine Learning with Applications, 9, 100334. https://doi.org/10.1016/j.mlwa.2022.100334
Chang, M.-Y., & Tseng, C.-Y. (2020). Detecting social anxiety with online social network data. 2020 21st IEEE International Conference on Mobile Data Management (MDM), 2020-June, 333–336. https://doi.org/10.1109/MDM48529.2020.00073
Coppersmith, G., Harman, C., & Dredze, M. (2014). Measuring post traumatic stress disorder in Twitter. Proceedings of the International AAAI Conference on Web and Social Media, 8(1), 579–582. https://doi.org/10.1609/icwsm.v8i1.14574
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. http://arxiv.org/abs/1810.04805
Duong, H. T., & Nguyen-Thi, T. A. (2021). A review: Pre-processing techniques and data augmentation for sentiment analysis. Computational Social Networks, 8(1). https://doi.org/10.1186/ s40649-020-00080-x
Eichstaedt, J. C., Smith, R. J., Merchant, R. M., Ungar, L. H., Crutchley, P., Preoţiuc-Pietro, D., Asch, D. A., & Schwartz, H. A. (2018). Facebook language predicts depression in medical records. Proceedings of the National Academy of Sciences, 115(44), 11203–11208. https://doi.org/10.1073/pnas.1802331115
Gaye, B., Zhang, D., & Wulamu, A. (2021). Sentiment classification for employees reviews using regression vector-stochastic gradient descent classifier (RV-SGDC). PeerJ Computer Science, 7, e712. https://doi.org/10.7717/peerj-cs.712 Journal of ICT, 23, No. 1 (January) 2024, pp: 77-
HaCohen-Kerner, Y., Miller, D., & Yigal, Y. (2020). The influence of pre-processing on text classification using a bag-of-words representation. PLOS ONE, 15(5), e0232525. https://doi.org/10.1371/journal.pone.0232525
Hameed, N., Shabut, A. M., Ghosh, M. K., & Hossain, M. A. (2020). Multi-class multi-level classification algorithm for skin lesions classification using machine learning techniques. Expert Systems with Applications, 141, 112961. https://doi.org/10.1016/j.eswa.2019.112961
Hosseini, P., Khoshsirat, S., Jalayer, M., Das, S., & Zhou, H. (2022). Application of text mining techniques to identify actual wrong-way driving (WWD) crashes in police reports. International Journal of Transportation Science and Technology. https://doi.org/10.1016/j.ijtst.2022.12.002
Hsu, B.-M. (2020). Comparison of supervised classification models on textual data. Mathematics, 8(5), 851. https://doi.org/10.3390/ math8050851
Iyortsuun, N. K., Kim, S.-H., Jhon, M., Yang, H.-J., & Pant, S. (2023). A review of machine learning and deep learning approaches on mental health diagnosis. Healthcare, 11(3), 285. https://doi.org/10.3390/healthcare11030285
Khattak, F. K., Jeblee, S., Pou-Prom, C., Abdalla, M., Meaney, C., & Rudzicz, F. (2019). A survey of word embeddings for clinical text. In Journal of Biomedical Informatics: X (Vol. 4). Academic Press Inc. https://doi.org/10.1016/j.yjbinx.2019.100057
Kim, J., Lee, D., & Park, E. (2021). Machine learning for mental health in social media: Bibliometric study. Journal of Medical Internet Research, 23(3), e24870. https://doi.org/10.2196/24870
Kim, J., Lee, J., Park, E., & Han, J. (2020). A deep learning model for detecting mental illness from user content on social media. Scientific Reports, 10(1), 11846. https://doi.org/10.1038/ s41598-020-68764-y
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2019). ALBERT: A lite BERT for self-supervised learning of language representations. http://arxiv.org/abs/1909.11942
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimised BERT pre-training approach. http://arxiv. org/abs/1907.11692
Luo, X. (2021). Efficient English text classification using selected machine learning techniques. Alexandria Engineering Journal, 60(3), 3401–3409. https://doi.org/10.1016/j.aej.2021.02.009 Journal of ICT, 23, No. 1 (January) 2024, pp: 77-
Maryamah, M., Arifin, A. Z., Sarno, R., Indraswari, R., & Sholikah, R. W. (2021). Pseudo-relevance feedback combining statistical and semantic term extraction for searching Arabic documents. International Journal of Intelligent Engineering and Systems, 14(5), 238–246. https://doi.org/10.22266/ijies2021.1031.22
Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C., & Joulin, A. (2017). Advances in pre-training distributed word representations. http://arxiv.org/abs/1712.09405
Nandwani, P., & Verma, R. (2021). A review on sentiment analysis and emotion detection from text. Social Network Analysis and Mining, 11(1), 81. https://doi.org/10.1007/s13278-021-00776-6
Noori, B. (2021). Classification of customer reviews using machine learning algorithms. Applied Artificial Intelligence, 35(8), 567–588. https://doi.org/10.1080/08839514.2021.1922843
Rai, N., Kumar, D., Kaushik, N., Raj, C., & Ali, A. (2022). Fake news classification using transformer based enhanced LSTM and BERT. International Journal of Cognitive Computing in Engineering, 3, 98–105. https://doi.org/10.1016/j. ijcce.2022.03.003
Ramírez-Cifuentes, D., Freire, A., Baeza-Yates, R., Puntí, J., Medina-Bravo, P., Velazquez, D. A., Gonfaus, J. M., & Gonzàlez, J. (2020). Detection of suicidal ideation on social media: Multimodal, relational, and behavioural analysis. Journal of Medical Internet Research, 22(7), e17758. https://doi.org/10.2196/17758
Rehm, J., & Shield, K. D. (2019). Global burden of disease and the impact of mental and addictive disorders. Current Psychiatry Reports, 21(2), 10. https://doi.org/10.1007/s11920-019-0997-0
Ríssola, E. A., Aliannejadi, M., & Crestani, F. (2022). Mental disorders on online social media through the lens of language and behaviour: Analysis and visualisation. Information Processing & Management, 59(3), 102890. https://doi.org/10.1016/j. ipm.2022.102890
Ríssola, E. A., Bahrainian, S. A., & Crestani, F. (2019). Anticipating depression based on online social media behaviour. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics): Vol. 11529 LNAI (pp. 278–290). Springer Verlag. https://doi.org/10.1007/978-3-030-27629-4_26
Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. http://arxiv.org/abs/1910.01108 Journal of ICT, 23, No. 1 (January) 2024, pp: 77-
Sarker, I. H. (2021). Deep learning: A comprehensive overview on techniques, taxonomy, applications and research directions. SN Computer Science, 2(6), 420. https://doi.org/10.1007/s42979-021-00815-1
Stein, D. J., Palk, A. C., & Kendler, K. S. (2021). What is a mental disorder? An exemplar-focused approach. Psychological Medicine, 51(6), 894–901. https://doi.org/10.1017/ S0033291721001185
Su, C., Xu, Z., Pathak, J., & Wang, F. (2020). Deep learning in mental health outcome research: A scoping review. Translational Psychiatry, 10(1), 116. https://doi.org/10.1038/s41398-020-0780-3
Sun, Y., Li, Y., Zeng, Q., & Bian, Y. (2020). Application research of text classification based on random forest algorithm. 2020 3rd International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE), 370–374. https://doi.org/10.1109/AEMCSE50948.2020.00086
Tadesse, M. M., Lin, H., Xu, B., & Yang, L. (2019). Detection of depression-related posts in Reddit social media forum. IEEE Access, 7, 44883–44893. https://doi.org/10.1109/ ACCESS.2019.2909180
Taha, A. Y., Tiun, S., Abd Rahman, A. H., & Sabah, A. (2021). Multilabel over-sampling and under-sampling with class alignment for imbalanced multilabel text classification. Journal of Information and Communication Technology, 20(3), 423–456. https://doi.org/10.32890/jict2021.20.3.6
Uban, A. S., Chulvi, B., & Rosso, P. (2021). An emotion and cognitive based analysis of mental health disorders from social media data. Future Generation Computer Systems, 124, 480–494. https://doi.org/10.1016/j.future.2021.05.032
Vinod Prakash, & Dharmender Kumar. (2023). A modified gated recurrent unit approach for epileptic electroencephalography classification. Journal of Information and Communication Technology, 22(4), 587–617. https://doi.org/10.32890/ jict2023.22.4.3
Zhang, T., Schoene, A. M., Ji, S., & Ananiadou, S. (2022). Natural language processing applied to mental illness detection: A narrative review. Npj Digital Medicine, 5(1), 46. https://doi.org/10.1038/s41746-022-00589-7
Published
Issue
Section
License
Copyright (c) 2024 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-04Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
- Scopus 21 View →
- Semantic Scholar 20 View →
- OpenAlex 17 View →
- OpenCitations 16 View →
- Crossref 16 View →
- Google Scholar no free count Search →
- Dimensions no free count Search →
2002 - 2020






















