Malay Named Entity Recognition System Using Machine Learning for Tourism in Malaysia

Authors

  • Juhaida Abu Bakar Universiti Utara Malaysia, Malaysia
  • Muhammad Asyraf Ariffin Universiti Utara Malaysia, Malaysia
  • Nor Hazlyna Harun Universiti Utara Malaysia, Malaysia
  • Ruziana Mohamad Rasli Universiti Utara Malaysia, Malaysia
  • Nurul Huda Mohamad Saad Universiti Teknologi Mara, Malaysia
  • Lisnawita Universitas Lancang Kuning, Indonesia

DOI:

https://doi.org/10.32890/jdsd2024.2.2.4

Keywords:

Named Entity Recognition, Malay Text, Machine Learning, BERT, ALXLNET

Abstract

Analysing unstructured textual data has become increasingly common due to its rich informational value across various fields. Named Entity Recognition (NER) is crucial for identifying entities in open-domain text documents. Current NER techniques often rely on manually labelled documents, which are timeconsuming and prone to inaccuracies. While methods such as Spacy and Polyglot have been used, more research is needed on applying machine learning techniques to this problem. This work addressed this gap by developing a Malay language NER system using machine learning. The system used available Malay corpus resources to identify, learn, tag, and store entities from Malay texts. It was designed to handle structured and unstructured data, extracting names of people, places, organisations, and other entities. The Malay NER System using Machine Learning was developed as a web-based application. It employed advanced machine learning models, specifically BERT and ALXLNET, to process and analyse data. This study shows good agreement among the respondents regarding the usability, perception, and feedback on the specific pages, with the lowest mean score being 76.67%. Regarding system functionalities, there is room for refinement to ensure more accurate and reliable output. The system featured a web interface allowing users to input Malay text and receive recognised entities as output. Performance was assessed using standard evaluation metrics. This work advanced natural language processing capabilities in Malay by creating a user-friendly, efficient tool for NER in Malay. 

References

Honnibal, M., & Montani, I. (2017). spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. GitHub.

Salleh, M. S., Asmai, S. A., Basiron, H., & Ahmad, S. (2017). A Malay-named entity recognition using conditional random fields. In 2017 5th International Conference on information and Communication Technology (ICOIC7) (pp. 1-6). IEEE.

Mbouopda, M. F., & Yonta, P. M. (2019). A Word Representation to Improve Named Entity Recognition in Low-Resource Languages. In 2019 Sixth International Conference on Social Networks Analysis, Management and Security (SNAMS) (pp. 333-337). IEEE.

Salleh, M. S., Asmai, S. A., Basiron, H., & Ahmad, S. (2018). Named entity recognition using fuzzy c-means clustering method for Malay textual data analysis. Journal of Telecommunication, Electronic and Computer Engineering (JTEC), 10(2-7), 121-126.

Rosmayati, M., Nazratul Naziah, M. M., Noor Maizura, M. N., & Zulaiha, A. O. (2020, December 15). A Review of Named Entity Recognition and Classification on Unstructured Malay Data.

Rami, A.-R. (2015). Named Entity Extraction — polyglot 16.07.04 documentation. https://polyglot.readthedocs.io/en/latest/NamedEntityRecognition.htm l

Ma, R., Peng, M., Zhang, Q., & Huang, X. (2019). Simplify the usage of the lexicon in Chinese NER— arXiv preprint arXiv:1908.05969.

Alfred, R., Leong, L. C., On, C. K., & Anthony, P. (2014). Malay named entity recognition based on a rule-based approach.

Sazali, S. S., Rahman, N. A., & Bakar, Z. A. (2016). Information extraction: Evaluating named entity recognition from classical Malay documents. In 2016, the third international conference on Information Retrieval and knowledge management (CAMP) (pp. 48-53). IEEE.

Asmai, S. A., Salleh, M. S., Basiron, H., & Ahmad, S. (2018). An enhanced Malay named entity recognition using a combination approach for crime textual data analysis. International Journal of Advanced Computer Science and Applications, 9(9), 474-483.

Chang, S. S., Bakar, J. A., & Katuk, N. (2021). Malay Roman Corpus Annotation System. Multidisciplinary Applied Research and Innovation, 2(3), 001–004.

Ulanganathan, T., Ebrahim, A., Xian, B. C. M., Bouzekri, K., Mahmud, R., & Hoe, O. H. (2017). Benchmarking Mi-NER: Malay entity recognition engine. In 9th International Conference on information, process, and knowledge management (pp. 52-58).

Zadgaonkar, A., & Agrawal, A. J. (2024). An Approach for analysing unstructured text data using topic modelling techniques for efficient information extraction. New Generation Computing, 42(1), 109-134.

Zennaki, O., Semmar, N., & Besacier, L. (2015). Utilisation des réseaux de neurones récurrents pour la projection interlingue d'étiquettes morpho-syntaxiques à partir d'un corpus parallèle. In TALN 2015.

Zolkepli, H. (2018). Malaya, a natural-language-toolkit library for Bahasa Malaysia, powered by Pytorch. https://github.com/huseinzol05/malaya.

Downloads

Published

20-10-2024

Issue

Section

Articles

How to Cite

Abu Bakar, J., Ariffin, M. A., Harun , N. H., Mohamad Rasli , R., Mohamad Saad, N. H., & Lisnawita. (2024). Malay Named Entity Recognition System Using Machine Learning for Tourism in Malaysia. Journal of Digital System Development, 2(2), 46-63. https://doi.org/10.32890/jdsd2024.2.2.4

Most read articles by the same author(s)