NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
DOI:
https://doi.org/10.32890/jict2016.15.1.5Keywords:
Malay corpora, word tokenization, regular expression, Buckwalter character codeAbstract
References
Abdul Ghani, R., Zakaria, M. S., & Omar, K. (2009). Jawi-Malay Transliteration. In International Conference on Electrical Engineering and Informatics 2009 (ICEEI’09) (Volume:01) (pp. 154–157). Selangor: http://jict.uum.edu.my IEEE. doi:10.1109/ICEEI.2009.5254799
Abdul Rahman, H. (1999). Panduan menulis dan mengeja Jawi. Kuala Lumpur: Dewan Bahasa dan Pustaka.
Abdullah, I.-H., Hashim, R. S., & Mohamed Husin, N. (2011). Lexical associations of Malayness in Hikayat Abdullah: A collocational analysis. Research Journal of Applied Sciences, 5(6), 429–433.
Abu Bakar, J. (2008). Transliterasi Jawi Lama-Jawi Baru berasaskan Grafem (Kajian Kes Hikayat Merong Mahawangsa) (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.
Ahmad, C. W. S. B. C. W., Omar, K., Nasrudin, M. F., Murah, M. Z., & Azmi, S. M. (2013). Machine transliteration for old Malay manuscript. In 2nd International Conference on Machine Learning and Computer Science (IMLCS’2013) (pp. 23–26). Kuala Lumpur.
Ahmad, C. W. S. C. W. (2007). Penterjemah Jawi lama kepada Jawi baru (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.
Atwell, E. (2008). Development of tag sets for part-of-speech tagging. In A. Ludeling & M. Kyto (Eds.), Corpus linguistics: An international handbook Volume 1 (pp. 501–526). Walter de Gruyter. Retrieved from http://eprints.whiterose.ac.uk/81781/
Azmi, M. S. (2013). Fitur baharu dari kombinasi geometri segitiga dan pengezonan utk paleografi Jawi digital.
Bakar, J. A., Omar, K., Nasrudin, M. F., Murah, M. Z., & Ahmad, C. W. S. C. W. (2013). Implementation of Buckwalter transliteration to Malay corpora. In 13th International Conference on Intelligent System Design and Applications (ISDA’13) (pp. 214–218). Serdang, Selangor: Mirlabs. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–
Balossi, G. (2014). A corpus linguistic approach to literary language and characterization: Virginia Woolf’s the waves. John Benjamins Publishing Company.
Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with python (1st ed.). USA: O’Reilly Media, Inc.
Buckwalter, T. (2002). Buckwalter Arabic morphological analyzer version 1.0. linguistic data consortium. Retrieved from https://catalog.ldc. http://jict.uum.edu.my upenn.edu/LDC2002L49
DBP. (2008). Daftar kata Bahasa Melayu Rumi-Sebutan-Jawi. Dlm. Dahaman & M. Ahmad (Eds.) (Kedua.). Kuala Lumpur: Dawama.
DBP. (2015). Pusat Rujukan Persuratan Melayu. Retrieved from http://prpm. dbp.gov.my/
Diab, M., Hacioglu, K., & Jurafsky, D. (2004). Automatic tagging of Arabic text:From raw text to base phrase chunks. In HLT-NAACL-Short ’04 Proceedings of HLT-NAACL 2004: Short Papers (pp. 149–152). Stroudsburg, PA, USA: Association for Computational Linguistics.
Diah, N. M., Ismail, M., Ahmad, S., & Abdullah, S. A. S. S. (2010). Jawi on mobile devices with Jawi WordSearch game application. In 2010 International Conference on Science and Social Research (CSSR 2010) (pp. 326–329). Kuala Lumpur, Malaysia: IEEE. doi:10.1109/ CSSR.2010.5773793
Diah, N. M., Ismail, M., Hami, P. M. A., & Ahmad, S. (2011). Assisted Jawi-Writing (AJaW) Software for children. In 2011 IEEE Conference on Open Systems (ICOS2011) (pp. 322–326). Langkawi: IEEE. doi:10.1109/ICOS.2011.6079260
Dukes, K., & Habash, N. (2010). Morphological annotation of Quranic Arabic. In Language Resources and Evaluation Conference (LREC) (pp. 2530–2536). Valletta, Malta: ELRA.
Habash, N., & Metsky, H. (2008). Automatic learning of morphological variations for handling out-of-vocabulary terms in Urdu-English machine translation. In Proceedings of the Association for Machine Translation in the Americas (AMTA-08). Waikiki, Hawai’i. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–
Habash, N., Soudi, A., & Buckwalter, T. (2007). On Arabic transliteration. In A. Soudi, A. van den Bosch, & G. Neumann (Eds.), Arabic computational morphology: Knowledge-based and Empirical Methods. Springer.
Heryanto, A., Nasrudin, M. F., & Omar, K. (2008). Offline Jawi Handwritten recognizer using hybrid artificial neural networks and dynamic programming. In International Symposium on Information Technology, 2008 (ITSim 2008) (Volume:2) (pp. 1–6). Kuala Lumpur: IEEE. doi:10.1109/ITSIM.2008.4631722 http://jict.uum.edu.my
Hock, O. Y. (2009). Kamus dwibahasa. Petaling Jaya: Pearson Longman.
Irvine, A., Weese, J., & Callison-Burch, C. (2012). Processing informal, romanized Pakistani text messages. In Proceedings of the 2012 Workshop on Language in Social Media (LSM 2012) (pp. 75–78). Association for Computational Linguistics.
Ismail, K., Yusof, R. J. R., & Jomhari, N. (2010). A case study of Jawi Editor in the XO-laptop simulated environment. In 2010 International Conference on User Science and Engineering (i-USEr) (pp. 21–25). Shah Alam: IEEE. doi:10.1109/IUSER.2010.5716716
Knowles, G., & Don, Z. M. (2003). Tagging a corpus of Malay texts, and coping with “syntactic drift.” In Proceedings of the corpus linguistics (pp. 422–428). Retrieved from http://eprints.lancs.ac.uk/8620/
Leech, G. (2005). Adding linguistic annotation. In M. Wynne (Ed.), Developing linguistic corpora: A Guide to good practice (pp. 17–29). Oxford: Oxbow Books for the Arts and Humanities Data Service. Retrieved from http://www.ahds.ac.uk/creating/guides/linguistic-corpora/chapter2.htm
Mohamed, H., Omar, N., & Ab Aziz, M. J. (2011). Statistical Malay Part-of-Speech (POS) tagger using hidden Markov approach. In 2011 International Conference on Semantic Technology and Information Retrieval (pp. 231–236). IEEE.
NLTK Project. (2015). NLTK corpora. Retrieved from http://www.nltk.org/ nltk_data
Noor, N. K. M., Noah, S. A., Aziz, M. J. A., & Hamzah, M. P. (2010). Anaphora resolution of Malay text: Issues and proposed solution model. 2010 International Conference on Asian Language Processing, 174–177. doi:10.1109/IALP.2010.80 Journal of ICT, 15, No. 1 (June) 2016, pp: 107–
Outahajala, M., Zenkouar, L., Benajiba, Y., Rosso, P., & Elirf. (2013). The development of a fine grained class set for amazigh POS tagging. In ACS International Conference on. IEEE Computer Systems and Applications (AICCSA) (pp. 1–8). IEEE.
Pasha, A., Al-Badrashiny, M., Diab, M., Kholy, A. El, Eskander, R., Habash, N., … Roth, R. M. (2014). MADAMIRA : A fast, comprehensive tool for morphological analysis and disambiguation of Arabic. In Proceedings of the Language Resources and Evaluation Conference (LREC) (pp. 1094–1101). Reykjavik, Iceland. http://jict.uum.edu.my
Rahman, S. A., & Omar, N. (2013). Transforming noun phrase structure form into rules to detect compound nouns in Malay sentences. Journal of ICT, 12, 161–173.
Rahman, S. A., Omar, N., & Aziz, M. J. A. (2011). A fundamental study on detecting head modifier noun phrases in Malay sentence. In 2011 International Conference on Semantic Technology and Information Retrieval (pp. 255–259). Putrajaya: IEEE. doi:10.1109/ STAIR.2011.5995798
Rahman, S. A., Omar, N. B., Mohamed, H., Juzaidin, M., & Aziz, A. (2011). A synonym contextual-based process for handling word similarity in Malay sentence.
Redika, R., Omar, K., & Nasrudin, M. F. (2008). Handwritten Jawi words recognition using hidden Markov models. In International Symposium on Information Technology, 2008 (ITSim 2008) (Volume:2) (pp. 1–5). Kuala Lumpur: IEEE. doi:10.1109/ITSIM.2008.4631723
Saad, N. H. M., Bakar, J. A., Karim, R. A., Tukiman, N., & Nor, K. M. (2012). Pembangunan korpus cerpen bertag Bahasa Melayu: Analisis linguistik korpora. In Research, Invention, Innovation & Design (RIID 2012). Universiti Teknologi MARA Kampus Melaka.
SEAlang Projects. (2011). Retrieved from http://sealang.net/malay/dictionary. htm
Shaalan, K., & Raza, H. (2007). Person name entity recognition for Arabic. In Proceedings of the 2007 Workshop on Computational Approaches to Semitic Languages: Common Issues and Resources (pp. 17–24). Stroudsburg, PA, USA. Journal of ICT, 15, No. 1 (June) 2016, pp: 107–
Sharum, M. Y., Abdullah, M. T., Sulaiman, M. N., Murad, M. A. A., & Hamzah, Z. A. Z. (2011). Name extraction for unstructured Malay text. 2011 IEEE Symposium on Computers & Informatics, 787–791. doi:10.1109/ISCI.2011.5959017
Sulaiman, S. (2013). Pencantas perkataan Melayu untuk aksara Jawi berasaskan petua. Bangi: Universiti Kebangsaan Malaysia.
Sulaiman, S., Omar, K., Omar, N., Murah, M. Z., & Abdul Rahman, H. (2014). The effectiveness of a Jawi stemmer for retrieving relevant Malay http://jict.uum.edu.my documents in Jawi characters. ACM Transactions on Asian Language Information Processing, 13(2), 6.
Sulaiman, S., Omar, K., Omar, N., Murah, M. Z., & Rahman, H. A. (2011). A Malay stemmer for Jawi characters. In D. Wang & M. Reynolds (Eds.), AI 2011: Advances in Artificial Intelligence (pp. 668–676). Perth, Australia: Springer Berlin / Heidelberg.
Tmshkina, J. (2006). Development of a multilingual parallel corpus and a part-of-speech tagger for Afrikaans. IFIP International Federation for Information Processing, 228, 453–462.
Unicode. (2014). Unicode. Retrieved from http://unicode.org
Yonhendri. (2008). Enjin transliterasi rumi Jawi (Unpublished master’s thesis). Universiti Kebangsaan Malaysia.
Published
Issue
Section
How to Cite
Research impact
Harvested 2026-09-06Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















