An Iterated Two-Step Sinusoidal Pitch Contour Formulation for Expressive Speech Synthesis
DOI:
https://doi.org/10.32890/jict2021.20.4.2Keywords:
Pitch contour formulation, prosody modification, speech synthesis, storytellingAbstract
Intonation generation in expressive speech such as storytelling is essential to produce high quality Malay language expressive speech
synthesizer. Intonation generation, for instance explicit control, has shown good performance in terms of intelligibility with reasonably
natural speech; thus, it was selected in this research. This approach modifies the prosodic features, such as pitch contour, intensity, and duration, to generate the intonation. However, modification of pitch contour remains a problem because the desired pitch contour is not achieved. This paper formulated an improved pitch contour algorithm to develop a modified pitch contour resembling the natural pitch contour. In this work, the syllable pitch contours of nine storytellers were extracted from their storytelling speeches to create an expressive speech syllable dataset called STORY_DATA. All the shapes of pitch contours from STORY_DATA were analyzed and clustered into the standard six main pitch contour clusters for storytelling. The clustering was performed using one minus the Pearson product moment correlation. Then, an improved iterative two-step sinusoidal pitch contour formulation was introduced to modify the pitch contours of a neutral speech into an expressive pitch contour of natural speeches. Overall, the improved pitch contour formulation was able to achieve 93 percent high correlated matches, indicating the high resemblance as compared to the previous pitch contour formulation at 15 percent. Therefore, the improved formula can be used in a text-to-speech (TTS) synthesizer to produce a more natural expressive speech. The paper also discovered unique expressive pitch contours in the Malay language that need further investigations in the future.
References
AAparicio, R. M., Salle, L., Ramon, U., Tècnica, E., & Electrònica,
E. (2016). Prosodic and voice quality cross-language analysis of storytelling expressive categories oriented to text-to-speech synthesis. (Doctoral dissertation, Universitat Ramon Llull).
Birkholz, P. (2013). Modeling consonant-vowel coarticulation for articulatory speech synthesis. PLoS ONE, 8(4), 1–17. https://doi.org/10.1371/journal.pone.0060603 Journal of ICT, 20, No. 4 (October) 2021, pp: 489–
Boersma, P., David, W., & Heuven, V. (2015). Praat doing phonetics by computer (v. 5.3.39). http://www.praat.org/
Chang, R. R., Yu, X. Q., Yuan, Y. Y., & Wan, W. G. (2014). Emotional analysis and synthesis of human voice based on STRAIGHT. Applied Mechanics and Materials, 536–537, 105–110. https://doi.org/10.4028/www.scientific.net/AMM.536-537.105
Chowdhury, S. (2006). Concatenative text-to-speech synthesis: A study on standard colloquial Bengali. (Doctoral dissertation, Indian Statistical Institute).
Gelin, R., D’Alessandro, C., & Le, Q. (2010, November). Towards a storytelling humanoid robot. In AAAI Fall Symposium Series on Dialog with Robots (pp. 137–138).
Gu, H., & Jiang, K. (2015, July). A pitch-contour generation method combining ANN, global variance, and real-contour selection. In 2015 International Conference on Machine Learning and Cybernetics (ICMLC) (pp. 396–402). https://doi.org/10.1109/
ICMLC.2015.7340954
Hamzah, R. (2016). Discriminative classification model of filled pause and elongation for Malay language spontaneous speech. (Doctoral dissertation, Universiti Teknologi MARA).
Hamzah, R., & Jamil, N. (2019). Investigation of speech disfluencies classification on different threshold selection techniques using energy feature. Malaysian Journal of Computing, 4(1), 178–192.
Hargus, S. (2005). Athabaskan phonetics and phonology. Language and Linguistics Compass, 4(10), 1019–1040.
Ikkunointi. (2016). Windowing. http://www.cs.tut.fi/kurssit/ SGN-4010/ikkunointi_en.pdf
Jamil, N., Apandi, F., & Hamzah, R. (2017, July). Influences of age in emotion recognition of spontaneous speech: A case of an under-resourced language. In 2017 International Conference on Speech Technology and Human-Computer Dialogue (SpeD) (pp. 1–6). IEEE.
Joaquim, L. (1992, July). Speaking styles in speech research. In Workshop on Integrating Speech and Natural Language (pp. 15–17).
Kato, S., Yasuda, Y., Wang, X., Cooper, E., Takaki, S., & Yamagishi, J. (2020). Modeling of Rakugo speech and its limitations: Toward speech synthesis that entertains audiences. IEEE Access, 8, 138149–138161. Journal of ICT, 20, No. 4 (October) 2021, pp: 489–
Keith, G. (2005). Correlation coefficients. https://www.andrews. edu/~calkins/math/edrm611/ edrm05.htm
Klabbers, E., & Santen, J. (2004, June). Clustering of foot-based pitch contours in expressive speech. In Fifth ISCA Workshop on Speech Synthesis (pp. 73–78).
Lunce, B. C. (2007). Digital storytelling as an educational tool. Indiana Libraries, 30(1), 77–80.
Lutfi, S. L. (2007). Adding emotions to synthesized Malay speech using diphone-based templates. (Master’s thesis, University of Malaya).
Md Saad, M., Jamil, N., & Hamzah, R. (2018). Evaluation of support vector machine and decision tree for emotion recognition of Malay folklores. Bulletin of Electrical Engineering and Informatics, 7(3), 479–486. https://doi.org/10.11591/eei. v7i3.1279
Montaño, R., Alías, F., & Ferrer, J. (2013, September). Prosodic analysis of storytelling discourse modes and narrative situations oriented to text-to-speech synthesis. In 8th ISCA Speech Synthesis Workshop (pp. 171–176).
Paliwal, K. K., James, L., & Kamil, W. (2010, December). Preference for 20-40 ms window duration in speech analysis.. In 2010 4th International Conference on Signal Processing and Communication Systems (ICSPCS) (pp. 1–4). IEEE.
Parker, V. (2014). 200 kisah teladan haiwan. Edukid Publication. Plaisant, C., Druin, A., Lathan, C., Dakhane, K., Edwards, K., Vice,
J. M., & Montemayor, J. (2000, November). A storytelling robot for pediatric rehabilitation. In Proceedings of the Fourth International ACM Conference on Assistive Technologies Assets ’00 (pp. 50–55). https://doi.org/10.1145/ 354324.354338
Podder, P., Khan, Zaman, T., & Haque Khan, M. (2014). Comparative performance analysis of hamming, hanning and blackman window. International Journal of Computer Applications, 96(18), 1–7.
Ramli, I., Jamil, N., Seman, N., & Ardi, N. (2016, November). An improved pitch contour formulation for Malay language storytelling text-to-speech (TTS). In IEEE Industrial Electronics and Applications Conference (IEACon) (pp. 250–255). IEEE.
Raúl, M., & Francesc, A. (2017). The role of prosody and voice quality in indirect storytelling speech: A cross-narrator perspective in four European languages. Speech Communication, 88, 1–16. Journal of ICT, 20, No. 4 (October) 2021, pp: 489–
Roekhaut, S., Goldman, J., & Simon, A. C. (2010, May). A model for varying speaking style in TTS systems. In Fifth International Conference on Speech Prosody (pp. 4–7). Sarkar, P., Haque, A., Dutta, A. K., Gurunath Reddy, M., Harikrishna, D. M., Dhara, P., Verma, R., Narendra, N. P., Sunil Kr, S. B.,
Yadav, J., & Rao, K. S. (2014, August). Designing prosody ruleset for converting neutral TTS speech to storytelling style speech for Indian languages: Bengali, Hindi and Telugu. In 2014 7th International Conference on Contemporary Computing (IC3) (pp. 473–477). https://doi.org/10.1109/IC3.2014.6897219
Schröder, M. (2009). Expressive speech synthesis: Past, present, and possible futures. Affective Information Processing, 2, 111–126.
Sebastian, D. (2014). Unit selection text to speech system for Polish. [Unpublished thesis] Faculty of Mechanical Engineering and Robotics, AGH University of Science and Technology.
Theune, M., Meijs, K., Heylen, D., & Ordelman, R. (2006). Generating expressive speech for storytelling applications. IEEE Transactions on Audio, Speech, and Language Processing, 14(4), 1099–1108.
Um, S.-Y., Oh, S., Byun, K., Jang, I., Ahn, C., & Kang H.-G. (2020, May). Emotional speech synthesis with rich and granularized control. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 7254–7258). IEEE.
Verma, R., Sarkar, P., & Rao, K. S. (2015, January). Conversion of neutral speech to storytelling style speech. In 8th International Conference on Advances in Pattern Recognition (pp. 1–6). https://doi.org/10.1109/ICAPR.2015.7050705
Published
Issue
Section
License
Copyright (c) 2022 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-05Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
- Scopus 2 View →
- OpenAlex 1 View →
- Semantic Scholar 1 View →
- OpenCitations 0 View →
- Crossref 0 View →
- Google Scholar no free count Search →
- Dimensions no free count Search →
2002 - 2020






















