Time-Distributed Attention-Layered Convolution Neural Network with Ensemble Learning using Random Forest Classifier for Speech Emotion Recognition
DOI:
https://doi.org/10.32890/jict2023.22.1.3Keywords:
Ensemble classifiers, Random Forest, Speech Emotion Recognition, Human Computer Interaction, time-distributed layers, spatiotemporal featuresAbstract
Speech Emotion Detection (SER) is a field of identifying human emotions from human speech utterances. Human speech utterances
are a combination of linguistic and non-linguistic information. Nonlinguistic SER provides a generalized solution in human–computer
interaction applications as it overcomes the language barrier. Machine learning and deep learning techniques were previously proposed for classifying emotions using handpicked features. To achieve effective and generalized SER, feature extraction can be performed using deep neural networks and ensemble learning for classification. The proposed model employed a time-distributed attention-layered convolution neural network (TDACNN) for extracting spatiotemporal features at the first stage and a random forest (RF) classifier, which is an ensemble classifier for efficient and generalized classification of emotions, at the second stage. The proposed model was implemented on the RAVDESS and IEMOCAP data corpora and compared with the CNN-SVM and CNN-RF models for SER. The TDACNN-RF model exhibited test classification accuracies of 92.19 percent and 90.27 percent on the RAVDESS and IEMOCAP data corpora, respectively. The experimental results proved that the proposed model is efficient in extracting spatiotemporal features from time-series speech signals and can classify emotions with good accuracy. The class confusion among the emotions was reduced for both data corpora, proving that the model achieved generalization.
References
Agajanian, S., Oluyemi, O., & Verkhivker, G. M. (2019). Integration of Random Forest Classifiers and Deep Convolutional Neural Networks for classification and biomolecular modeling of cancer driver mutations. Frontiers in molecular biosciences. https://doi.org/10.3389/fmolb.2019.00044 Journal of ICT, 22, No. 1 (January) 2023, pp: 49–
Albornoz, E. M., & Milone, D. H. (2015). Emotion recognition in never-seen languages using a novel ensemble method with emotion profiles. IEEE Transactions on Affective Computing, 8(1), 43-53. https://doi.org/10.1109/TAFFC.2015.2503757
Alonso, C. J., Rodrı, J. J., Kuncheva, L. I., & Alonso, C. J. (2006). Rotation forest: A new classifier ensemble method. IEEE Transactions on Pattern Analysis and Machine Intelligence, March 2014. https://doi.org/10.1109/TPAMI.2006.211
Amiriparian, S., Gerczuk, M., Ottl, S., Cummins, N., Freitag, M., Pugachevskiy, S., Baird, A., & Schuller, B. (2017). Snore sound classification using image-based deep spectrum features. Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 3512–3516. https://doi.org/10.21437/Interspeech.2017-434
Atila, O., & Şengür, A. (2021). Attention guided 3D CNN-LSTM model for accurate speech based emotion recognition. Applied Acoustics, 182, 108260. https://doi.org/10.1016/J. APACOUST.2021.108260
Bhavan, A., Chauhan, P., Hitkul, & Shah, R. R. (2019). Bagged support vector machines for emotion recognition from speech. Knowledge-Based Systems, 184, 104886. https://doi.org/10.1016/j.knosys.2019.104886
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32.
Busso, C., Bulut, M., Lee, C.-C., Kazemzadeh, A., Mower, E., Kim, S., Chang, J. N., Lee, S., & Narayanan, S. S. (2007). IEMOCAP: Interactive emotional dyadic motion capture database. Lang Resources & Evaluation, 42, 335–359 (2008). https://doi.org/10.1007/s10579-008-9076-6
Chen, L., Su, W., Feng, Y., Wu, M., She, J., & Hirota, K. (2020a). Two-layer fuzzy multiple random forest for speech emotion recognition in human-robot interaction. Information Sciences, 509, 150–163. https://doi.org/10.1016/j.ins.2019.09.005
Chen, M., He, X., Yang, J., & Zhang, H. (2018). 3-D Convolutional recurrent neural networks with attention model for speech emotion recognition. IEEE Signal Processing Letters, 25(10), 1440–1444. https://doi.org/10.1109/LSP.2018.2860246
Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., & Bengio, Y. (2015). Attention-based models for speech recognition. Advances in neural information processing systems, 28. https://doi.org/10.48550/arXiv.1506.07503 Journal of ICT, 22, No. 1 (January) 2023, pp: 49–
Cummins, N., Amiriparian, S., Hagerer, G., Batliner, A., Steidl, S., & Schuller, B. W. (2017). An image-based deep spectrum feature representation for the recognition of emotional speech. MM 2017 - Proceedings of the 2017, ACM Multimedia Conference, 478–484. https://doi.org/10.1145/3123266.3123371
Fayek, H. M., Lech, M., & Cavedon, L. (2017). Evaluating deep learning architectures for Speech Emotion Recognition. Neural Networks, 92, 60–68. https://doi.org/10.1016/j. neunet.2017.02.013
Fayek, H. M., Lech, M., & Cavedon, L., (2017). Evaluating deep learning architectures for Speech Emotion Recognition. Neural Networks, 92(December), 60–68. https://doi.org/10.1016/j. neunet.2017.02.013
Ganaie, M. A., Hu, M., Tanveer, M., & Suganthan, P. N. (2021). Ensemble deep learning: A review. arXiv. https://doi.org/10.1016/j.engappai.2022.105151
Gudmalwar, A. P., Rama Rao, C. V., & Dutta, A. (2019). Improving the performance of the speaker emotion recognition based on low dimension prosody features vector. International Journal of Speech Technology, 22(3), 521–531. https://doi.org/10.1007/ s10772-018-09576-4
Issa, D., Fatih Demirci, M., & Yazici, A. (2020). Speech emotion recognition with deep convolutional neural networks. Biomedical Signal Processing and Control, 59. https://doi.org/10.1016/j.bspc.2020.101894
Jiang, W., Wang, Z., Jin, J. S., Han, X., & Li, C. (2019). Speech emotion recognition with heterogeneous feature unification of deep neural network. Sensors (Switzerland), 19 (12), 1–15. https://doi.org/10.3390/s19122730
Kondo, K., & Taira, K. (2018). Estimation of binaural speech intelligibility using machine learning. Applied Acoustics, 129, 408–416. https://doi.org/10.1016/j.apacoust.2017.09.001
Kong, Y., & Yu, T. (2018). A Deep Neural Network model using Random Forest to extract feature representation for gene expression data classification. Scientific Reports, 8, 16477. https://doi.org/10.1038/s41598-018-34833-6
Kuchibhotla, S., Yalamanchili, B. S., Vankayalapati, H. D., & Anne, K. R. (2014). Speech emotion recognition using regularized discriminant analysis. In Advances in Intelligent Systems and Computing, 247, 363-369. https://doi.org/10.1007/978-3-319-02931-3_41 Journal of ICT, 22, No. 1 (January) 2023, pp: 49–
Kumar, S., Ratnoo, S., & Vashishtha, J. (2021). Hyper-heuristic evolutionary approach for constructing decision tree classifiers. Journal of Information and Communication Technology. 20, 2, 249–276. https://doi.org/10.32890/jict2021.20.2.5.
Lalitha, S., Tripathi, S., & Gupta, D. (2019). Enhanced speech emotion detection using deep neural networks. International Journal of Speech Technology, 22(3), 497–510. https://doi.org/10.1007/ s10772-018-09572-8
Lech, M., Stolar, M., Best, C., & Bolia, R. (2020). Real-time speech emotion recognition using a pre-trained image classification network: Effects of bandwidth reduction and companding. Frontiers in Computer Science, 2(May), 1–14. https://doi.org/10.3389/fcomp.2020.00014
Lee, T. H., Ullah, A., & Wang, R. (2020). Bootstrap aggregating and random forest. In Macroeconomic forecasting in the era of big data. Advanced Studies in Theoretical and Applied Econometrics (pp. 389-429). Springer.
Li, P., Song, Y., Mcloughlin, I., Guo, W., & Dai, L. (2018). An attention pooling based representation learning method for speech emotion recognition. Interspeech, September, 3087–3091.
Lieskovská, E., Jakubec, M., Jarina, R., & Chmulík, M. (2021). A review on speech emotion recognition using deep learning and attention mechanism. Electronics (Switzerland), 10 (10), 1163. https://doi.org/10.3390/electronics10101163
Livingstone, S. R., & Russo, F. A. (2018). The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north American english. PLoS ONE, 13 (5), e0196391. https://doi.org/10.1371/journal.pone.0196391
Lucky, H., & Suhartono, D. (2022). Botnet detection in IoT devices using random forest classifier with independent component analysis. Journal of Information and Communication Technology Technology, 1(1), 71–94. https://doi.org/10.32890/ jict2022.21.2.3
Mao, Q., Dong, M., Huang, Z., & Zhan, Y. (2014). Learning salient features for speech emotion recognition using convolutional neural networks. IEEE Transactions on Multimedia 16(8), 2203–2213. https://doi.org/10.1109/TMM.2014.2360798
Mirsamadi, S., & Barsoum, E. (2017). Automatic speech emotion recognition using recurrent neural networks with local attention. IEEE International Conference on Acoustics, Speech and Journal of ICT, 22, No. 1 (January) 2023, pp: 49– Signal Processing (ICASSP) October. https://doi.org/10.1109/ ICASSP.2017.7952552
Mustaqeem, & Kwon, S. (2020). CLSTM: Deep feature-based speech emotion recognition using the hierarchical convlstm network. Mathematics, 8(12), 1–19. https://doi.org/10.3390/ math8122133
Mustaqeem, Sajjad, M., & Kwon, S. (2020). Clustering-based speech emotion recognition by incorporating learned features and deep BiLSTM. IEEE Access, 8, 79861–79875. https://doi.org/10.1109/ACCESS.2020.2990405
Neumann, M., & Vu, N. T. (2017). Attentive convolutional neural network based speech emotion recognition: A study on the impact of input Features, signal length, and acted speech. Interspeech.
Noroozi, F., Marjanovic, M., Njegus, A., Escalera, S., & Member, S. (2018). A study of language and classifier-independent feature analysis for vocal emotion recognition. November, 1–24. arXiv preprint arXiv:1811.08935.
Noroozi, F., Sapiński, T., & Kamińska, D. (2017). Vocal ‑ based emotion recognition using random forests and decision tree. International Journal of Speech Technology, 20(2), 239–246. https://doi.org/10.1007/s10772-017-9396-2
Patrick, M. K., Adekoya, A. F., Mighty, A. A., & Edward, B. Y. (2022). Capsule networks–a survey. Journal of King Saud University-computer and information sciences, 34(1), 1295-1310. https://doi.org/10.1016/j.jksuci.2019.09.014
Pawar, M. D., & Kokate, R. D. (2021). Convolution neural network based automatic speech emotion recognition using Mel-frequency Cepstrum coefficients. Multimedia tools and applications. https://doi.org/10.1007/s11042-020-10329-2
Qiao, H., Wang, T., Wang, P., Qiao, S., & Zhang, L. (2018). A time-distributed spatiotemporal feature learning method for machine health monitoring with multi-sensor time series. Sensors, 18(9), 2932. https://doi.org/10.3390/s18092932
Ren, Z., Pandit, V., Qian, K., & Yang, Z. (2017). Deep sequential image features for acoustic scene classification. Detection and Classification of Acoustic Scenes and Events 2017, 113-117.
Sainath, T. N., Vinyals, O., Senior, A., & Sak, H. (2015). Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Journal of ICT, 22, No. 1 (January) 2023, pp: 49– Proceedings, 2015-August,4580–4584. https://doi.org/10.1109/ ICASSP.2015.7178838
Sun, H., Zheng, X., Lu, X., & Wu, S. (2020). Spectral-spatial attention network for hyperspectral image classification. IEEE Transactions on Geoscience and Remote Sensing, 58(5), 3232–3245. https://doi.org/10.1109/TGRS.2019.2951160
Sun, L., Zou, B., Fu, S., Chen, J., & Wang, F. (2019). Speech emotion recognition based on DNN-decision tree SVM model. Speech Communication,115, 29–37. https://doi.org/10.1016/J. SPECOM.2019.10.004
Tuncel, K. S., & Baydogan, M. G. (2018). Autoregressive forests for multivariate time series modeling. Pattern Recognition, 73, 202–215. https://doi.org/10.1016/j.patcog.2017.08.016
Tzirakis, P., Trigeorgis, G., Nicolaou, M. A., Schuller, B., & Zafeiriou, S. (2017). End-to-end multimodal emotion recognition using Deep Neural Networks. IEEE Journal of Selected Topics in Signal Processing, 11(8), 1301-1309. https://doi.org/10.1109/ JSTSP.2017.2764438.
Wang, H., Zhang, Q., Wu, J., Pan, S., & Chen, Y. (2019). Time series feature learning with labeled and unlabeled data. Pattern Recognition, Pattern Recognition, 89, 55–66. https://doi.org/10.1016/J.PATCOG.2018.12.026
Wei, C., Chen, L. lan, Song, Z. zhen, Lou, X. guang, & Li, D. dong. (2020). EEG-based emotion recognition using simple recurrent units network and ensemble learning. Biomedical Signal Processing and Control, 58, 101756. https://doi.org/10.1016/j. bspc.2019.101756
Yao, Z., Wang, Z., Liu, W., Liu, Y., & Pan, J. (2020). Speech emotion recognition using fusion of three multi-task learning-based classifiers: HSF-DNN, MS-CNN and LLD-RNN. Speech Communication, 120(March), 11–19. https://doi.org/10.1016/j. specom.2020.03.005
Zehra, W., Javed, A. R., Jalil, Z., Khan, H. U., & Gadekallu, T. R. (2021). Cross corpus multilingual speech emotion recognition using ensemble learning. Complex & Intelligent Systems, 7(4), 1845–1854. https://doi.org/10.1007/s40747-020-00250-4
Zhao, Z., Bao, Z., Zhao, Y., Zhang, Z., Cummins, N., Ren, Z., & Schuller, B. (2019). Exploring Deep Spectrum representations via attention-based recurrent and convolutional neural networks for speech emotion recognition. IEEE Access, 7, 97515–97525. https://doi.org/10.1109/ACCESS.2019.2928625 Journal of ICT, 22, No. 1 (January) 2023, pp: 49–
Zheng, C., Wang, C., & Jia, N. (2020). An ensemble model for multi-level speech emotion recognition. Applied Sciences (Switzerland), 10 (1), 205. https://doi.org/10.3390/app10010205
Zhou, Z. H., Wu, J., & Tang, W. (2002). Ensembling neural networks: Many could be better than all. Artificial Intelligence, 137(1–2), 239–263. https://doi.org/10.1016/S0004-3702(02)00190-X
Zvarevashe, K., & Olugbara, O. (2020a). Ensemble learning of hybrid acoustic features for speech emotion recognition. Algorithms, 13(3), 70. https://doi.org/10.3390/a13030070
Zvarevashe, K., & Olugbara, O. (2020b). Recognition of cross-language acoustic emotional valence using stacked ensemble learning. Algorithms, 13(10), 246. https://doi.org/10.3390/ a13100246.
Published
Issue
Section
License
Copyright (c) 2023 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-05Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















