Projecting Named Entity Tags from a Resource Rich Language to a Resource Poor Language

Authors

  • Norshuhani Zamin Faculty of Science and Information Technology, Universiti Teknologi PETRONAS Bandar Seri Iskandar, 31750 Tronoh, Perak, Malaysia
  • Alan Oxley Faculty of Science and Information Technology, Universiti Teknologi PETRONAS Bandar Seri Iskandar, 31750 Tronoh, Perak, Malaysia
  • Zainab Abu Bakar Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA 40450 Shah Alam, Selangor, Malaysia

DOI:

https://doi.org/10.32890/jict2013.12.7

Keywords:

Named entity recognition, information projection, bitext alignment, resource poor language, unsupervised learning, Malay terrorism corpus

Abstract

Named Entities (NE) are the prominent entities appearing in textual documents. Automatic classification of NE in a textual corpus is a vital process in Information Extraction and Information Retrieval research. Named Entity Recognition (NER) is the identification of words in text that correspond to a pre-defined taxonomy such as person, organization, location, date, time, etc. This article focuses on the person (PER), organization (ORG) and location (LOC) entities for a Malay journalistic corpus of terrorism. A projection algorithm, using the Dice Coefficient function and bigram scoring method with domain-specific rules, is suggested to map the NE information from the English corpus to the Malay corpus of terrorism. The English corpus is the translated version of the Malay corpus. Hence, these two corpora are treated as parallel corpora. The method computes the string similarity between the English words and the list of available lexemes in a pre-built lexicon that approximates the best NE mapping. The algorithm has been effectively evaluated using our own terrorism tagged corpus; it achieved satisfactory results in terms of precision, recall, and F-measure. An evaluation of the selected open source NER tool for English is also presented.

 

References

Lucarelli, G., Vasilakos, V., & Androutsopoulos, I. (2007). Named entity recognition in Greek texts.

McCallum, A. & Li, W. (2003). Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons. Proceedings of the Natural Language Learning, 188-191.

Alfonseca, E., & Manandhar, S. (2002). An unsupervised method for general named entity recognition and automated concept discovery. Proceedings of the International Conference on General WordNet, 34-43.

Alias-I. (2008). LingPipe 4.1.0. Retrieved from http://alias-i.com/lingpipe. Association of Computational Linguistic, 31, 531-574. doi: 10.1162/ 089120105775299177.

Benajiba, Y., Zitouni, I., Diab, M., & Rosso, P. (2010). Arabic named entity recognition: using features extracted from noisy data. Proceedings of the Association for Computational Linguistics Conference Short Papers, 281-285.

Bender, O., Och, F. J., & Ney, H. (2003). Maximum entropy models for named entity recognition. Proceedings of the Natural Language Learning, 148-151.

Bick, E. (2004). A named entity recognizer for Danish. Proceedings of the Language Resources and Evaluation, 305-308.

Boutsis, S., Demitros, L., Giouli, V., Liakata, M., Papageorgiou, H., Piperidis, S. (2000). A system for recognition of named entities in Greek. Proceedings of the Natural Language Processing, 424-436

Budi, I., & Bressan, S. (2007). Application of association rules mining to named entity recognition and co-reference resolution for the Indonesian language. International Journal of BI and DM, 2, 426-446. Journal of ICT, 12, 2013, pp: 121–

Budi, I., Bressan, S., Wahyudi, G., Hasibuan, Z., & Nazief, B. (2005). Named entity recognition for the Indonesian language: combining contextual, morphological and part-of-speech features into a knowledge engineering approach. Discovery Science, 57-69. Springer Berlin/Heidelberg.

Chanlekha, H., & Kawtrakul, A. (2004). Thai named entity extraction http://jict.uum.edu.my/ by incorporating maximum entropy model with simple heuristic information. Proceedings of the International Joint Conference in Natural Language Processing.

Chen, Z., & Ji, H. (2009). Can one language bootstrap the other: A case study on event extraction. Proceedings of the NAACL HLT 2009 Workshop on Semi-Supervised Learning for Natural Language Processing, 66-74.

Ciaramita, M., Gangemi, A., Ratsch, E., Šarić, J., & Rojas, I. (2008). Unsupervised learning of semantic relations for molecular biology ontologies. Proceeding of the Ontology Learning and Population: Bridging the Gap between Text and Knowledge, 91-104.

Cucchiarelli, A., & Velardi, P. (2001). Unsupervised named entity recognition using syntactic and semantic contextual evidence. Computational Linguistics, 27, 123-131.

Da Silva, J. F., Kozareva, Z., & Lopes, G.P. (2004). Cluster analysis and classification of named entities. Proceedings of the Language Resources and Evaluation, 321-324. doi: 10.1.1.99.4830.

Dice, L. R. (1945). Measures of the amount of ecologic association between species. Ecology, 26, 297-302.

Dien, D. I. N. H. (2001). Building an English-Vietnamese bilingual corpus (Unpublished master’s thesis in comparative linguistics). University of Social Sciences and Humanity of Ho Chi Minh City, Vietnam.

Dien, D. I. N. H. (2005). Building an annotated English-Vietnamese parallel corpus. MKS: A Journal of Southeast Asian Linguistics and Languages, 35, 21-36.

Doddington, G, Mitchell, A., Przybocki, M., Ramshaw, L., Strassel, S. & Weischedel, R. (2004). Automatic Content Extraction (ACE) program - task definitions and performance measures. Proceedings of the Language Resources and Evaluation. Journal of ICT, 12, 2013, pp: 121–

Ekbal, A., & Bandyopadhyay, S. (2008,). Bengali named entity recognition using Support Vector Machine. Proceedings of the Workshop on NER for South and South East Asian Languages and International Joint Conference on Natural Language Processing, 51-58.

Ekbal, A., & Bandyopadhyay, S. (2009). A Conditional Random Field http://jict.uum.edu.my/ approach for named entity recognition in Bengali and Hindi. Linguistic Issues in Language Technology, 2, 1-44.

El-Imam, Y. A. & Don Z. M. (2005). Rules and algorithms for phonetic transcription of standard Malay. IEICE Transaction of Information Systems, E88-D, 2354-2372.

Etzioni, O., Cafarella, M., Downey, D., Popescu, A., Shaked, T., Soderland, S., Weld, D., & Yates, A. (2005). Unsupervised named-entity extraction from the web: An experimental study. Artificial Intelligence, 165, 91–134.

Evans, R. (2003). A framework for named entity recognition in the open domain. Proceedings of the Recent Advances in Natural Language Processing, 137-144, doi: 10.1.1.105.7150.

Farmakiotou, D., Karkaletsis, V., Koutsias, J., Sigletos, G., Spyropoulos, C. D., & Stamatopoulos, P. (2000). Rule-based named entity recognition for Greek financial texts. Proceedings of the Workshop on Computational lexicography and Multimedia Dictionaries, 75-78.

Federico, M., Nicola, B., & Vanessa, S. (2002). Bootstrapping entity recognition for Italian broadcast news. Proceedings of the Empirical Methods in Natural Language Processing, 296-303. doi: 10.1.1.11.3083.

Feldman, R., & Sanger, J. (2007). Text mining handbook. Cambridge: Cambridge University Press.

Finkel, J.R., Grenager T., & Manning, C. (2005). Incorporating non-local information into information extraction systems by Gibbs sampling. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 363-370.

Gao, J., Lee, M. (2005). Chinese word segmentation and named entity recognition: A pragmatic approach. Computational Linguistics, 31, 531-574. Journal of ICT, 12, 2013, pp: 121–

Georgi, R., Xia, F., & Lewis, W. (2012). Measuring the divergence of dependency structures cross-linguistically to improve syntactic projection algorithms. Proceedings of the Language Resources and Evaluation, 771-778.

Georgiev, G., Nakov, P., & Ganchev, K. (2009). Feature-rich named http://jict.uum.edu.my/ entity recognition for Bulgarian using Conditional Random Fields. Proceedings of the Recent Advances in Natural Language Processing, 113-117.

Ghahramani, Z. (2004). Unsupervised learning. Advanced Lectures on Machine Learning, 72-112.

Giouli, V., Konstandinidis, A., Desypri, E., & Papageorgiou, Harris. (2006). Multi-domain multi-lingual named entity recognition: Revisiting and rounding the resources issue. Proceedings of the Associations for Computational Linguistics, 59-64.

Gruhl, D., Nagarajan, M., Pieper, J., Robson, C., & Sheth., A. (2009). Context and domain knowledge enhanced entity spotting in informal text. Proceedings of the Semantic Web Conference, 260-276. doi: 10.1007/978-3-642-04930-9_17.

Hamza, O., Bontcheva, K., Maynard, D., Tablan, V., & Cunningham, H. (2003). Named entity recognition in Romanian. Technical Report, Department of Computer Science. University of Sheffield.

Han, X., & Ruonan, R. (2011). The method of medical named entity recognition based on semantic model and improved SVM-KNN algorithm. Proceedings of the Semantic Knowledge and Grid, 21-27.

Heng, J., & Grishman, R. (2006). Data selection for semi supervised learning for named tagging. Proceedings of the Workshop on Information Extraction Beyond The Document, 48-55. Association for Computational Linguistics.

Hofmann, T. (2001). Unsupervised learning by probabilistic latent semantic analysis. Machine Learning, 42, 177-196.

Huang, F. (2005). Multilingual named entity extraction and translation from texts and speech (Unpublished doctoral dissertation). Pittsburgh: Carnegie Mellon University. Journal of ICT, 12, 2013, pp: 121–

Isozaki, H. (2001). Japanese named entity recognition based on a simple rule generator and decision tree learning. Proceedings of the Annual Meeting on Association of Computational Linguistics, 314-321.

Jansche, M., & Abney, S. (2002). Information extraction from voicemail transcripts. Proceedings of the Empirical Methods in Natural Language http://jict.uum.edu.my/ Processing, 10, 320-327. doi: 10.3115/1118693.1118734.

Johannessen, J. B., Hagen, K., Haaland, A., Jonsdottir, A.B., Noklestad, A., Kokkinakis, D., & Meurer, P. (2005). Named entity recognition for the mainland Scandinavian languages. Literary and Linguistics Computing, 20, 91-102.

Karaa, W.B.A. (2011). Named entity recognition using web document corpus. Managing Information Technology, 3. doi: 10.5121/ijmit.2011.3104.

Kim, J.H., & Woodland, P.C. (2000). Rule-based named entity recognition. Technical Report CUED/F-INFENG/TR.385. Cambridge University.

Kim, M., & Compton, P. (2012). Improving the performance of a named entity recognition system with knowledge acquisition. Knowledge Engineering and Knowledge Management, 97-113.

Kucuk, D., & Yazici, A. (2012). A hybrid named entity recognizer for Turkish. Expert Systems with Applications, 39, 2733–2742.

Lim, N. R., New, J. C., Ngo, M. A., Sy, M., & Lim, N. R. (2007). A named-entity recognizer for Filipino texts. Proceedings of the National Natural Language Processing Research Symposium, 20-25.

Lucarelli, G., Vasilakos, X., & Androutsopoulos, I. (2007). Named entity recognition in Greek with an ensemble of SVMs and active learning, Artificial Intelligence Tools Journal, 16, 1015-1045.

Ma, X. (2010). Toward a named entity aligned bilingual corpus. Proceedings of the Language Resources and Evaluation.

Mayfield, J., Lawrie, D., McNamee, P., & Oard, D. (2011). Building a cross-language entity linking collection in twenty-one languages. Multilingual and Multimodal Information Access Evaluation, 6941, 3-13. Springer-Verlag. Journal of ICT, 12, 2013, pp: 121–

Maynard, D., Tablan, V., & Cunningham, H. (2003). NE recognition without training data on a language you don’t speak. Proceedings of the ACL Workshop on Multilingual and Mixed-Language Named Entity Recognition: Combining Statistical and Symbolic Models, 33-40. Association for Computational Linguistics. http://jict.uum.edu.my/

Maynard, D., Tablan, V., Ursu, C., Cunningham, H., & Wilks, Y. (2001). Named entity recognition from diverse text types. Proceedings of the Recent Advances in Natural Languages Processing, 257-274. doi: 10.1.1.18.7395.

McCallum, A., & Li, W. (2003). Early results for named entity recognition with Conditional Random Felds, feature induction and web-enhanced lexicons. Proceedings of the Computational Natural Language Learning, 4, 188-191. doi: 10.3115/1119176.1119206.

Metin, K.S., Kisla, T., & Bahar, K. (2012). Named entity recognition in Turkish using association measures. Advanced Computing: An International Journal, 3, 43-49. doi: 10.5121/aclj.2012.3406.

Minkov, E., Wang, R., & Cohen, W. (2005). Extracting personal names from emails: applying named entity recognition to informal text. Proceedings of the Human Language Technology and Conference on Empirical Methods in Natural Language Processing, 443-450. doi:10.3115/1220575.1220631.

Nadeau, D. (2007). Semi-supervised named entity recognition: Learning to recognize 100 entity types with little supervision (Unpublished doctoral dissertation). University of Ottawa.

Nadeau, D., & Satoshi, S. (2007). A survey of named entity recognition and classification. Linguisticae Investigaciones, 30, 3-26. doi:10.1075/ li.30.1.03nad.

Nadeau, D., Turney, P., & Matwin, S. (2006). Unsupervised named-entity recognition: generating gazetteers and resolving ambiguity. Advances in Artificial Intelligence: Conference of the Canadian Society for Computational Studies of Intelligence, 19, 266-277. Springer.

Nguyen, D., Hoang, S., Pham, S., & Nguyen, T. (2010). Named entity recognition for Vietnamese. Proceedings of the Conference on Intelligent Information and Database Systems: Part II, 205-214. Springer-Verlag. Journal of ICT, 12, 2013, pp: 121–

Pasca, M., Lin, D., Bigham, J., Lifchits, A., & Jain, A. (2006). Organizing and searching the world wide web of facts – step one: the one million of fact extraction challenge. Proceedings of the National Conference on Artificial Intelligence, 1400-1405.

Pinnis, M. (2012). Latvian and Lithuanian named entity recognition with http://jict.uum.edu.my/ TildeNER. Proceedings of the Language Resources and Evaluation, 1258-1265.

Poibeau, T. (2003). The multilingual named entity recognition framework. Proceedings of the Conference on European Chapter of the Associations for Computational Linguistics, 2, 155-158. doi: 10.3115/ 1067737.1067772.

Poibeau, T., & Kosseim, L. (2001). Proper name extraction from non-journalistic texts. Proceedings of the Computational Linguistics in Netherland. 144-157. doi: 10.1.1.21.4746.

Rajesh, S., & Goyal, V. (2011). Name entity recognition systems for Hindi using CRF approach. Information Systems for Indian Languages: International Conference of Information Systems for Indian Languages, 139, p. 31-35. Springer.

Ralf, S., & Bruno, P. (2007). Cross-lingual named entity recognition. Lingvisticae Investigationes, 30, 135-162.

Ratinov, L., & Roth, D. (2009). Design challenges and misconceptions in named entity recognition. Proceedings of the Conference on Computational Natural Language Learning, 147-155.

Rushin, S., Lin, B., Gershman, A. & Frederking, R. (2010). SYNERGY: a named entity recognition system for resource-scarce languages such as Swahili using online machine translation. Proceedings of the Workshop on African Language Technology, 21-26.

Sari, Y., Hassan, M. F., & Zamin, N. (2009). A hybrid approach to semi-supervised named entity recognition in health, safety and environment reports. Proceedings of the Future Computer and Communication, 599-602.

Semenza, C. (1997). Proper-name-specific aphasias. In H.Goodglass & A.Wingfiled (Eds.), Anomia: Neuroanatomical and cognitive correlates (pp. 115-134). San Diego: Academic Press. Journal of ICT, 12, 2013, pp: 121–

Singh, A. K. (2008). Natural language processing for less privileged languages: Where do we come from? Where are we going?. Proceedings of the Workshop on NLP for Less Privileged Languages, 7–12.

Singh, A. K., Pala, K., & Surana, H. (2008). Estimating the resource adaption cost from a resource rich language to a similar resource poor language. http://jict.uum.edu.my/ Proceedings of the Language Resources and Evaluation, 3514-3519.

Srihari, R.K., Niu, C. & Li, W. (2000). A hybrid approach to named entity and sub-type tagging. Proceedings of Applied Natural Language Processing, 247-254.

Stern, R., & Benoit, S. (2010). Resources for named entity recognition and resolution in news wires. Proceedings of the Workshop on Resources and Evaluation for Identity Matching, Entity Resolution and Entity Management.

Sujan, S., Sarkar, S., & Mitra, P. (2008). A hybrid feature set based maximum entropy Hindi named entity recognition. Proceedings of the International Joint Conference in Natural Langauge Processing, 343-350.

Szarvas, G., Farkas, R., & Kocsor, A. (2006). A multilingual named entity recognition system using boosting and c4. 5 decision tree learning algorithms. Discovery Science, 267-278. Springer Berlin/Heidelberg.

Tanenblatt, M., Coden, A., & Sominsky, I. (2010). The Concept Mapper approach to named entity recognition. Proceedings of the Language Resources and Evaluation, 546-551.

Thelen, M., & Riloff, E. (2002). A bootstrapping method for learning semantic lexicons using extraction pattern contexts. Proceedings of the ACL-02 Conference on Empirical Methods in Natural Language Processing, 10, 241-221.

Tongtep, N., & Theeramunkong, T. (2011). Simultaneous character-cluster-based word segmentation and named entity recognition in Thai language. Knowledge, Information, and Creativity Support Systems: Proceedings of Knowledge, Information and Creativity Support Systems, 6746, 216-225. Springer.

Tran, Q. T., Pham, T. T., Ngo, Q. H., Dinh, D., & Collier, N. (2007). Named entity recognition in Vietnamese documents. Progress in Informatics Journal, 5, 14-17. Journal of ICT, 12, 2013, pp: 121–

Utsuro, T., & Sassano, M. (2000). Minimally supervised Japanese named entity recognition: resources and evaluation. Proceedings of the Language Resources and Evaluation, 1229-1236.

Wan, X., Zong, L., Huang, X., Ma, T., Jia, H., Wu, Y., & Xiao, J. (2011). Named entity recognition in Chinese news comments on the web. Proceedings http://jict.uum.edu.my/ of the International Joints Conference of Natural Language Processing, 856-864.

Whitelaw, C., & Patrick, J. (2003). Evaluating corpora for named entity recognition using character-level features. Proceedings of the Advances in Artificial Intelligence, 910-921.

Wu, D., Sun Lee, W., Ye, N., & Leong Chieu, H. (2009). Domain adaptive bootstrapping for named entity recognition. Proceedings of the Empirical Methods in Natural Language Processing, 1523-1532.

Wu, Y., Zhao, J., & Xu, B. (2003). Chinese named entity recognition combining a statistical model with human knowledge. Proceedings of the Multilingual and Mixed-language Named Entity Recognition, 65-72.

Yarowsky, D., Ngai, G., & Wicentowski, R. (2001). Inducing multilingual text analysis tools via robust projection across aligned corpora. Proceedings of the Human Language Technology Research, 1-8.

Zamin, N., & Oxley, A. (2011). Building a corpus-derived gazetteer for named entity recognition. Proceedings of the Communications in Computer and Information Science, 73-80.

Zamin, N., Oxley, A., Bakar, Z. A., & Farhan, S.A. (2012a). A statistical dictionary-based word alignment algorithm: An unsupervised approach. Proceedings of the International Conference of Computer and Information Sciences, 396-402.

Zamin, N., Oxley, A., Abu Bakar, Z., & Farhan, S. (2012b). A lazy man’s way to part-of-speech tagging. Proceedings of Knowledge Management and Acquisition for Intelligent Systems, 106-117.

Zhou, G., & Su, J. (2002). Named entity recognition using an HMM-based chunk tagger. Proceedings of the Annual Meeting on Association for Computational Linguistics, 473-480.

Downloads

Published

23-04-2013

How to Cite

Zamin, N., Oxley, A., & Abu Bakar, Z. (2013). Projecting Named Entity Tags from a Resource Rich Language to a Resource Poor Language. Journal of Information and Communication Technology, 12, 121-146. https://doi.org/10.32890/jict2013.12.7

Research impact

Harvested 2026-09-06
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2013.12.7