Improved Speaker-Independent Emotion Recognition from Speech Using Two-Stage Feature Reduction

Authors

  • Hasrul Mohd Nazid School of Mechatronic Engineering, Universiti Malaysia Perlis, Malaysia
  • Hariharan Muthusamy School of Mechatronic Engineering, Universiti Malaysia Perlis, Malaysia
  • Vikneswaran Vijean School of Mechatronic Engineering, Universiti Malaysia Perlis, Malaysia
  • Sazali Yaacob Kulim Hi-Tech Park, Malaysia

DOI:

https://doi.org/10.32890/jict2015.14.4

Keywords:

Emotional speech, cepstral features, feature reduction, emotion recognition

Abstract

In the recent years, researchers are focusing to improve the accuracy of speech emotion recognition. Generally, high emotion recognition accuracies were obtained for two-class emotion recognition, but multi-class emotion recognition is still a challenging task . The main aim of this work is to propose a two-stage feature reduction using Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA) for improving the accuracy of the speech emotion recognition (ER) system. Short-term speech features were extracted from the emotional speech signals. Experiments were carried out using four different supervised classifi ers with two different emotional speech databases. From the experimental results, it can be inferred that the proposed method provides better accuracies of 87.48% for speaker dependent (SD) and gender dependent (GD) ER experiment, 85.15% for speaker independent (SI) ER experiment, and 87.09% for gender independent (GI) experiment.

 

References

Bondugula, R., Duzlevski, O., & Xu, D. (2005). Profiles and fuzzy K-nearest neighbor algorithm for protein secondary structure prediction. In Y. P. P. Chen and L. Wong (Ed.), Proceedings of the 3rd Asia Bioinformatics Conference (pp. 237-288). London: Imperial College Press.

Bozkurt, E., Erzin, E., Erdem, C. E., & Erdem, A. T. (2010). Use of line spectral frequencies for emotion recognition from speech. Paper presented at the 2010 20th International Conference on Pattern Recognition, Istanbul, Turkey.

Burkhardt, F., Paeschke, A., Rolfes, M., Sendlmeier, W. F., & Weiss, B. (2005). A database of German emotional speech. Paper presented at the Interspeech 2005, Lisbon, Portugal.

Canu, S., Grandvalet, Y., Guigue, V., & Rakotomamonjy, A. (2005). SVM and kernel methods MATLAB toolbox. Perception Systmes et Information, INSA de Rouen, Rouen, France, 2, 21. Journal of ICT, 14, 2015, pp: 57–

Chee, L. S., Ai, O. C., Hariharan, M., & Yaacob, S. (2009). MFCC based recognition of repetitions and prolongations in stuttered speech using k-NN and LDA. Paper presented at the 2009 IEEE Student Conference on Research and Development (SCOReD), Serdang, Malaysia.

Chia Ai, O., Hariharan, M., Yaacob, S., & Sin Chee, L. (2012). Classification of speech dysfluencies with MFCC and LPCC features. Expert Systems with Applications, 39(2), 2157-2165.

Deng, H.-B., Jin, L.-W., Zhen, L.-X., & Huang, J.-C. (2005). A new facial expression recognition method based on local gabor filter bank and PCA plus LDA. International Journal of Information Technology, 11(11), 86-96.

Dey, S., Rajan, R., Padmanabhan, R., & Murthy, H. A. (2011). Feature diversity for emotion, language and speaker verification. Paper presented at the 2011 National Conference on Communications (NCC), Bangalore, India.

Duda, R. O., Hart, P. E., & Stork, D. G. (2001). Pattern classification (2nd ed.). New York: John Wiley & Sons.

El Ayadi, M., Kamel, M. S., & Karray, F. (2011). Survey on speech emotion recognition: Features, classification schemes, and databases. Pattern Recognition, 44(3), 572-587.

Giannoulis, P., & Potamianos, G. (2012). A hierarchical approach with feature selection for emotion recognition from speech. Paper presented at the Language Resources and Evaluation (LREC), Istanbul, Turkey.

Haq, S., & Jackson, P. (2009). Speaker-dependent audio-visual emotion recognition. Paper presented at the International Conference on Audio-Visual Speech Processing, Norwich, UK.

Haq, S., Jackson, P. J., & Edge, J. (2008). Audio-visual feature selection and reduction for emotion classification. Paper presented at the Proceedings of the International Conference on Auditory-Visual Speech Processing (AVSP’08), Tangalooma, Australia.

Hariharan, M., Chee, L. S., Ai, O. C., & Yaacob, S. (2012). Classification of speech dysfluencies using LPC based parameterization techniques. Journal of Medical Systems, 36(3), 1821-1830.

Huang, G.-B., Zhou, H., Ding, X., & Zhang, R. (2012). Extreme learning machine for regression and multiclass classification. IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics, 42(2), 513-529.

Huang, G.-B., Zhu, Q.-Y., & Siew, C.-K. (2006). Extreme learning machine: theory and applications. Neurocomputing, 70(1), 489-501.

Huang, J., Yang, W., & Zhou, D. (2012). Variance-Based Gaussian Kernel Fuzzy Vector Quantization for Emotion Recognition with Short Speech. Paper presented at the 2012 IEEE 12th International Conference on Computer and Information Technology (CIT), Chengdu, China. Journal of ICT, 14, 2015, pp: 57–

Iliou, T., & Anagnostopoulos, C.-N. (2010). SVM-MLP-PNN classifiers on speech emotion recognition field-A comparative study. Paper presented at the 2010 5th International Conference on Digital Telecommunications (ICDT), Athens, Greece.

Jang, R. (2011). Audio signal processing and recognition. Retrieved from http://neural.cs.nthu.edu.tw/jang/books/audiosignal processing/

Keller, J. M., Gray, M. R., & Givens, J. A. (1985). A fuzzy k-nearest neighbor algorithm. IEEE Transactions on Systems, Man and Cybernetics,(4), 580-585.

Kim, Y. K., & Han, J. H. (1995). Fuzzy K-NN algorithm using modified K-selection. Paper presented at the International Joint Conference of the 4th IEEE International Conference on Fuzzy Systems and The Second International Fuzzy Engineering Symposium, Yokohama, Japan.

Koolagudi, S. G., & Rao, K. S. (2012). Emotion recognition from speech: a review. International Journal of Speech Technology, 15(2), 99-117.

Kotti, M., & Paternò, F. (2012). Speaker-independent emotion recognition exploiting a psychologically-inspired binary cascade classification schema. International Journal of Speech Technology, 15(2), 131-150.

Liu, C.-L., Lee, C.-H., & Lin, P.-M. (2010). A fall detection system using k-nearest neighbor classifier. Expert Systems with Applications, 37(10), 7174-7181.

Pan, Y., Shen, P., & Shen, L. (2005). Feature Extraction and Selection in Speech Emotion Recognition. Paper presented at the IEEE Conference on Advanced Video and Signal Based Surveillance (AVSS 2005), Como, Italy.

Pan, Y., Shen, P., & Shen, L. (2012). Speech Emotion Recognition Using Support Vector Machine. International Journal of Smart Home, 6(2).

Rabiner, L. R., & Juang, B.-H. (1993). Fundamentals of Speech Recognition (1st ed.): New Jersey: Prentice Hall Englewood Cliffs.

Sedaaghi, M. (2008). Documentation of the Sahand emotional speech database (SES). Technical Report, Department of Electrical Engineering, Sahand University of Technology, Iran,.

Sezgin, M. C., Gunsel, B., & Kurt, G. K. (2012). Perceptual audio features for emotion detection. EURASIP Journal on Audio, Speech, and Music Processing, 2012(1), 1-21.

Shahzadi, A., Ahmadyfard, A., Harimi, A., & Yaghmaie, K. (2013). Speech emotion recognition using non-linear dynamics features. Turkish Journal of Electrical Engineering & Computer Sciences. doi: 10.3906/ elk-1302-90

Shen, P., Changjun, Z., & Chen, X. (2011). Automatic speech emotion recognition using support vector machine. Paper presented at the 2011 International Conference on Electronic and Mechanical Engineering and Information Technology (EMEIT), Heilongjiang, China. Journal of ICT, 14, 2015, pp: 57–

Shlens, J. (2005). A tutorial on principal component analysis: University of California at San Diego.

Ververidis, D., & Kotropoulos, C. (2006). Emotional speech recognition: Resources, features, and methods. Speech Communication, 48(9), 1162-1181.

Wang, Y., & Guan, L. (2004). An investigation of speech-based human emotion recognition. Paper presented at the 2004 IEEE 6th Workshop on Multimedia Signal Processing, Siena, Italy.

Yusuf, S. A. M., Mahat, N. I., Siraj, F., & Yaacob, S. (2012). Noise robustness of first formant bandwidth (f1bw) features in malay vowel recognition. Journal of Information and Communication Technology, 11, 147-162.

Downloads

Published

28-04-2015

How to Cite

Mohd Nazid, H., Muthusamy, H., Vijean, V., & Yaacob, S. (2015). Improved Speaker-Independent Emotion Recognition from Speech Using Two-Stage Feature Reduction. Journal of Information and Communication Technology, 14, 57-76. https://doi.org/10.32890/jict2015.14.4

Research impact

Harvested 2026-09-06
13 citations, from OpenAlex — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2015.14.4 OpenAlex W4255005081 Scopus 85182138125

Most read articles by the same author(s)