Normalization of Noisy Texts in Malaysian Online Reviews

Authors

  • Norlela Samsudin Faculty of Computer and Mathematical Science, Universiti Teknologi MARA Terengganu, Dungun, 23000, Terengganu, Malaysia
  • Mazidah Puteh Faculty of Computer and Mathematical Science, Universiti Teknologi MARA Terengganu, Dungun, 23000, Terengganu, Malaysia
  • Abdul Razak Hamdan Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia 43600, Bangi, Selangor, Malaysia
  • Mohd Zakree Ahmad Nazri Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia 43600, Bangi, Selangor, Malaysia

DOI:

https://doi.org/10.32890/jict2013.12.8

Keywords:

Noisy texts, normalization of noisy texts, artificial abbreviation

Abstract

The process of gathering useful information from online messages has increased as more and more people use the Internet and other online applications such as Facebook and Twitter to communicate with each other. One of the problems in processing online messages is the high number of noisy texts that exist in these messages. Few studies have shown that the noisy texts decreased the result of text mining activities. On the other hand, very few works have investigated on the patterns of noisy texts that are created by Malaysians. In this study, a common noisy terms list and an artificial abbreviations list were created using specific rules and were utilized to select candidates of correct words for a noisy term. Later, the correct term was selected based on a bi-gram words index. The experiments used online messages that were created by the Malaysians. The result shows that normalization of noisy texts using artificial abbreviations list compliments the use of common noisy texts list.

 

References

Acharyya, S., Negi, S., Subramaniam, L. V., & Roy, S. (2008). Unsupervised learning of multilingual short message service (SMS) dialect from noisy examples. Paper presented at the Second Workshop on Analytics for Noisy Unstructured Text Data, Singapore.

Aw, A., Zhang, M., Xiao, J., & Su, J. (2006). A phrase-based statistical model for SMS text normalization. Paper at the COLING/ACL on Main Conference Poster Sessions, Sydney, Australia.

Choudhury, M., Saraf, R., Jain, V., Mukherjee, A., Sarkar, S., & Basu, A. (2007). Investigation and modeling of the structure of texting language. International Journal on Document Analysis and Recognition, 10(3), 157-174. doi: 10.1007/s10032-007-0054-0

Clark, A. (2003). Pre-processing very noisy text. Paper presented the Workshop Shallow Processing at Large, Corpora, Lancester. Journal of ICT, 12, 2013, pp: 147–

Cook, P., & Stevenson, S. (2009). An unsupervised model for text message normalization. Paper presented at the Workshop on Computational Approaches to Linguistic Creativity, Boulder, Colorado.

Dey, L., & Haque, S. K. M. (2008). Opinion mining from noisy text data. Paper presented at the Second Workshop on Analytics for Noisy Unstructured http://jict.uum.edu.my/ Text Data. Singapore.

Dey, L., & Haque, S. K. M. (2009). Studying the effects of noisy text on text mining applications. Paper presented at the Third Workshop on Analytics for Noisy Unstructured Text Data, Barcelona, Spain.

Foster, J., Wagner, J., & Genabith, J. V. (2008). Adapting a WSJ-trained parser to grammatically noisy text. Paper presented at the 46th Annual Meeting of the Association for Computational Linguistics on Human Language Technologies: Short Papers, Columbus, Ohio.

Hussin, S. (2009). Bahasa SMS. Retrieved from http://supyanhussin. wordpress.com/2009/07/11/bahasa-sms/

Jing, H., Lopresti, D., & Shih, C. (2003). Summarization of noisy documents: a pilot study. Paper presented at the HLT-NAACL 03 on Text summarization workshop - Volume 5.

Kobus, C., Yvon, F., & Damnati, G. (2008). Normalizing SMS: Are two methaphors better than one? Paper presented at the 22nd International Conference on Computational Linguistics, Manchester.

Kothari, G., Negi, S., Faruquie, T. A., Chakaravarthy, V. T., & Subramaniam, L. V. (2009). SMS based interface for FAQ retrieval. Paper presented at the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2 - Volume 2, Suntec, Singapore.

Dewan Bahasa dan Pustaka, (2008). Panduan Singkatan Khidmat Pesanan Ringkas. Retrieved from http://www.dbp.gov.my/khidmatsms.pdf

Samsudin, N., Puteh, M., & Hamdan, A. R. (2011). Bess or xbest: Mining the Malaysian online reviews. Paper presented at the 3rd Conference on Data Mining and Optimization (DMO). Journal of ICT, 12, 2013, pp: 147–

Tang, J., Li, H., Cao, Y., & Tang, Z. (2005). Email data cleaning. Paper at the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining, Chicago, Illinois, USA.

Toutanova, K., & Moore, R. C. (2002). Pronounciation modeling for improved spelling correction. Paper presented at the 40th Annual Meeting on http://jict.uum.edu.my/ Association for Computational Linguistics, Philadelphia, Pennsylvania.

Vinciarelli, A. (2004, 23-26 Aug). Noisy text categorization. Paper presented at the 17th International Conference on Pattern Recognition, ICPR.

Wong, W., Leu, W., & Bennamoun, M. (2006). Integrated scoring for spelling error correction, abbreviation expansion and case restoration in dirty text. Paper presented at the Australasian Data Mining Conference, Sydney.

Wong, W., Liu, W., & Bennamoun, M. (2006). Integrated scoring for spelling error correction, abbreviation expansion and case restoration in dirty text. Paper presented at the Fifth Australasian Conference on Data Mining and Analystics - Volume 61, Sydney, Australia.

Downloads

Published

23-04-2013

How to Cite

Samsudin, N., Puteh, M., Hamdan, A. R., & Ahmad Nazri, M. Z. (2013). Normalization of Noisy Texts in Malaysian Online Reviews. Journal of Information and Communication Technology, 12, 147-159. https://doi.org/10.32890/jict2013.12.8

Research impact

Harvested 2026-09-06
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2013.12.8

Most read articles by the same author(s)