NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization

Authors

  • Juhaida Abu Bakar Universiti Utara Malaysia, Malaysia
  • Khairuddin Omar Universiti Kebangsaan Malaysia, Malaysia
  • Mohammad Faidzul Nasrudin Universiti Kebangsaan Malaysia, Malaysia
  • Mohd Zamri Murah Universiti Kebangsaan Malaysia, Malaysia

DOI:

https://doi.org/10.32890/jict2016.15.1.5

Keywords:

Malay corpora, word tokenization, regular expression, Buckwalter character code

Abstract

This paper describes the design and creation of a monolingual parallel corpus for the Malay language written in Jawi. This paper proposes a new corpus called the National University of Malaysia Word Tokenization (NUWT) corpora To the best of our knowledge, currently, there is no sufficiently comprehensive, well-designed standard corpus that is annotated and made available for the public for the Jawi script corpora. This corpus contains the Jawi-specific Buckwalter character code and can be used to evaluate the performance of word tokenization tasks, as well as further language processing. The objective of this work is to conform and standardize the corpora between similar characters in Jawi. It consists of three subcorporas with documents from different genres. The gathering and processing steps, as well as the definition of several evaluation tasks regarding the use of these corpora, are included in this paper. One of the important roles and fundamental tasks of the corpus, which is the tokenization, is also presented in this paper. The development of the Malay language tokenizer is based on the syntactic data compatibility of Malay words written in Jawi. A series of experiments were performed to validate the corpus and to fulfill the requirement of the Jawi script tokenizer with an average error rate of 0.020255. Based on this promising result, the token will be used for the disambiguation and unknown word resolution, such as out-of-vocabulary (OOV) problem in the tagging process.

 

References

Abdul Ghani, R., Zakaria, M. S., & Omar, K. (2009). Jawi-Malay Transliteration. In International Conference on Electrical Engineering and Informatics 2009 (ICEEI’09) (Volume:01) (pp. 154–157). Selangor: http://jict.uum.edu.my IEEE. doi:10.1109/ICEEI.2009.5254799

Abdul Rahman, H. (1999). Panduan menulis dan mengeja Jawi. Kuala Lumpur: Dewan Bahasa dan Pustaka.

Abdullah, I.-H., Hashim, R. S., & Mohamed Husin, N. (2011). Lexical associations of Malayness in Hikayat Abdullah: A collocational analysis. Research Journal of Applied Sciences, 5(6), 429–433.

Abu Bakar, J. (2008). Transliterasi Jawi Lama-Jawi Baru berasaskan Grafem (Kajian Kes Hikayat Merong Mahawangsa) (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.

Ahmad, C. W. S. B. C. W., Omar, K., Nasrudin, M. F., Murah, M. Z., & Azmi, S. M. (2013). Machine transliteration for old Malay manuscript. In 2nd International Conference on Machine Learning and Computer Science (IMLCS’2013) (pp. 23–26). Kuala Lumpur.

Ahmad, C. W. S. C. W. (2007). Penterjemah Jawi lama kepada Jawi baru (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.

Atwell, E. (2008). Development of tag sets for part-of-speech tagging. In A. Ludeling & M. Kyto (Eds.), Corpus linguistics: An international handbook Volume 1 (pp. 501–526). Walter de Gruyter. Retrieved from http://eprints.whiterose.ac.uk/81781/

Azmi, M. S. (2013). Fitur baharu dari kombinasi geometri segitiga dan pengezonan utk paleografi Jawi digital.

Bakar, J. A., Omar, K., Nasrudin, M. F., Murah, M. Z., & Ahmad, C. W. S. C. W. (2013). Implementation of Buckwalter transliteration to Malay corpora. In 13th International Conference on Intelligent System Design and Applications (ISDA’13) (pp. 214–218). Serdang, Selangor: Mirlabs. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–

Balossi, G. (2014). A corpus linguistic approach to literary language and characterization: Virginia Woolf’s the waves. John Benjamins Publishing Company.

Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with python (1st ed.). USA: O’Reilly Media, Inc.

Buckwalter, T. (2002). Buckwalter Arabic morphological analyzer version 1.0. linguistic data consortium. Retrieved from https://catalog.ldc. http://jict.uum.edu.my upenn.edu/LDC2002L49

DBP. (2008). Daftar kata Bahasa Melayu Rumi-Sebutan-Jawi. Dlm. Dahaman & M. Ahmad (Eds.) (Kedua.). Kuala Lumpur: Dawama.

DBP. (2015). Pusat Rujukan Persuratan Melayu. Retrieved from http://prpm. dbp.gov.my/

Diab, M., Hacioglu, K., & Jurafsky, D. (2004). Automatic tagging of Arabic text:From raw text to base phrase chunks. In HLT-NAACL-Short ’04 Proceedings of HLT-NAACL 2004: Short Papers (pp. 149–152). Stroudsburg, PA, USA: Association for Computational Linguistics.

Diah, N. M., Ismail, M., Ahmad, S., & Abdullah, S. A. S. S. (2010). Jawi on mobile devices with Jawi WordSearch game application. In 2010 International Conference on Science and Social Research (CSSR 2010) (pp. 326–329). Kuala Lumpur, Malaysia: IEEE. doi:10.1109/ CSSR.2010.5773793

Diah, N. M., Ismail, M., Hami, P. M. A., & Ahmad, S. (2011). Assisted Jawi-Writing (AJaW) Software for children. In 2011 IEEE Conference on Open Systems (ICOS2011) (pp. 322–326). Langkawi: IEEE. doi:10.1109/ICOS.2011.6079260

Dukes, K., & Habash, N. (2010). Morphological annotation of Quranic Arabic. In Language Resources and Evaluation Conference (LREC) (pp. 2530–2536). Valletta, Malta: ELRA.

Habash, N., & Metsky, H. (2008). Automatic learning of morphological variations for handling out-of-vocabulary terms in Urdu-English machine translation. In Proceedings of the Association for Machine Translation in the Americas (AMTA-08). Waikiki, Hawai’i. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–

Habash, N., Soudi, A., & Buckwalter, T. (2007). On Arabic transliteration. In A. Soudi, A. van den Bosch, & G. Neumann (Eds.), Arabic computational morphology: Knowledge-based and Empirical Methods. Springer.

Heryanto, A., Nasrudin, M. F., & Omar, K. (2008). Offline Jawi Handwritten recognizer using hybrid artificial neural networks and dynamic programming. In International Symposium on Information Technology, 2008 (ITSim 2008) (Volume:2) (pp. 1–6). Kuala Lumpur: IEEE. doi:10.1109/ITSIM.2008.4631722 http://jict.uum.edu.my

Hock, O. Y. (2009). Kamus dwibahasa. Petaling Jaya: Pearson Longman.

Irvine, A., Weese, J., & Callison-Burch, C. (2012). Processing informal, romanized Pakistani text messages. In Proceedings of the 2012 Workshop on Language in Social Media (LSM 2012) (pp. 75–78). Association for Computational Linguistics.

Ismail, K., Yusof, R. J. R., & Jomhari, N. (2010). A case study of Jawi Editor in the XO-laptop simulated environment. In 2010 International Conference on User Science and Engineering (i-USEr) (pp. 21–25). Shah Alam: IEEE. doi:10.1109/IUSER.2010.5716716

Knowles, G., & Don, Z. M. (2003). Tagging a corpus of Malay texts, and coping with “syntactic drift.” In Proceedings of the corpus linguistics (pp. 422–428). Retrieved from http://eprints.lancs.ac.uk/8620/

Leech, G. (2005). Adding linguistic annotation. In M. Wynne (Ed.), Developing linguistic corpora: A Guide to good practice (pp. 17–29). Oxford: Oxbow Books for the Arts and Humanities Data Service. Retrieved from http://www.ahds.ac.uk/creating/guides/linguistic-corpora/chapter2.htm

Mohamed, H., Omar, N., & Ab Aziz, M. J. (2011). Statistical Malay Part-of-Speech (POS) tagger using hidden Markov approach. In 2011 International Conference on Semantic Technology and Information Retrieval (pp. 231–236). IEEE.

NLTK Project. (2015). NLTK corpora. Retrieved from http://www.nltk.org/ nltk_data

Noor, N. K. M., Noah, S. A., Aziz, M. J. A., & Hamzah, M. P. (2010). Anaphora resolution of Malay text: Issues and proposed solution model. 2010 International Conference on Asian Language Processing, 174–177. doi:10.1109/IALP.2010.80 Journal of ICT, 15, No. 1 (June) 2016, pp: 107–

Outahajala, M., Zenkouar, L., Benajiba, Y., Rosso, P., & Elirf. (2013). The development of a fine grained class set for amazigh POS tagging. In ACS International Conference on. IEEE Computer Systems and Applications (AICCSA) (pp. 1–8). IEEE.

Pasha, A., Al-Badrashiny, M., Diab, M., Kholy, A. El, Eskander, R., Habash, N., … Roth, R. M. (2014). MADAMIRA : A fast, comprehensive tool for morphological analysis and disambiguation of Arabic. In Proceedings of the Language Resources and Evaluation Conference (LREC) (pp. 1094–1101). Reykjavik, Iceland. http://jict.uum.edu.my

Rahman, S. A., & Omar, N. (2013). Transforming noun phrase structure form into rules to detect compound nouns in Malay sentences. Journal of ICT, 12, 161–173.

Rahman, S. A., Omar, N., & Aziz, M. J. A. (2011). A fundamental study on detecting head modifier noun phrases in Malay sentence. In 2011 International Conference on Semantic Technology and Information Retrieval (pp. 255–259). Putrajaya: IEEE. doi:10.1109/ STAIR.2011.5995798

Rahman, S. A., Omar, N. B., Mohamed, H., Juzaidin, M., & Aziz, A. (2011). A synonym contextual-based process for handling word similarity in Malay sentence.

Redika, R., Omar, K., & Nasrudin, M. F. (2008). Handwritten Jawi words recognition using hidden Markov models. In International Symposium on Information Technology, 2008 (ITSim 2008) (Volume:2) (pp. 1–5). Kuala Lumpur: IEEE. doi:10.1109/ITSIM.2008.4631723

Saad, N. H. M., Bakar, J. A., Karim, R. A., Tukiman, N., & Nor, K. M. (2012). Pembangunan korpus cerpen bertag Bahasa Melayu: Analisis linguistik korpora. In Research, Invention, Innovation & Design (RIID 2012). Universiti Teknologi MARA Kampus Melaka.

SEAlang Projects. (2011). Retrieved from http://sealang.net/malay/dictionary. htm

Shaalan, K., & Raza, H. (2007). Person name entity recognition for Arabic. In Proceedings of the 2007 Workshop on Computational Approaches to Semitic Languages: Common Issues and Resources (pp. 17–24). Stroudsburg, PA, USA. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–

Sharum, M. Y., Abdullah, M. T., Sulaiman, M. N., Murad, M. A. A., & Hamzah, Z. A. Z. (2011). Name extraction for unstructured Malay text. 2011 IEEE Symposium on Computers & Informatics, 787–791. doi:10.1109/ISCI.2011.5959017

Sulaiman, S. (2013). Pencantas perkataan Melayu untuk aksara Jawi berasaskan petua. Bangi: Universiti Kebangsaan Malaysia.

Sulaiman, S., Omar, K., Omar, N., Murah, M. Z., & Abdul Rahman, H. (2014). The effectiveness of a Jawi stemmer for retrieving relevant Malay http://jict.uum.edu.my documents in Jawi characters. ACM Transactions on Asian Language Information Processing, 13(2), 6.

Sulaiman, S., Omar, K., Omar, N., Murah, M. Z., & Rahman, H. A. (2011). A Malay stemmer for Jawi characters. In D. Wang & M. Reynolds (Eds.), AI 2011: Advances in Artificial Intelligence (pp. 668–676). Perth, Australia: Springer Berlin / Heidelberg.

Tmshkina, J. (2006). Development of a multilingual parallel corpus and a part-of-speech tagger for Afrikaans. IFIP International Federation for Information Processing, 228, 453–462.

Unicode. (2014). Unicode. Retrieved from http://unicode.org

Yonhendri. (2008). Enjin transliterasi rumi Jawi (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.

Downloads

Published

25-05-2016

How to Cite

Abu Bakar, J., Omar, K., Nasrudin, M. F., & Murah, M. Z. (2016). NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization. Journal of Information and Communication Technology, 15(1), 107-131. https://doi.org/10.32890/jict2016.15.1.5

Research impact

Harvested 2026-09-06
8 citations, from OpenAlex — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2016.15.1.5 OpenAlex W2592820646

Most read articles by the same author(s)