A Discourse-Based Information Retrieval for Tamil Literary Texts

Authors

  • Anita Ramalingam Department of Computer Science and Engineering SRM Institute of Science and Technology, India
  • Subalalitha Chinnaudayar Navaneethakrish Department of Computer Science and Engineering SRM Institute of Science and Technology, India

DOI:

https://doi.org/10.32890/jict2021.20.3.4

Keywords:

Discourse parser, Morphological Analyzer, Inverted indexing, Ranking, Tamil information retrieval

Abstract

Tamil literature has many valuable thoughts that can help the human community to lead a successful and a happy life. Tamil literary works are abundantly available and searched on the World Wide Web (WWW), but the existing search systems follow a keyword-based match strategy which fails to satisfy the user needs. This necessitates the demand for a focused Information Retrieval System that semantically analyses the Tamil literary text which will eventually improve the search system performance. This paper proposes a novel Information Retrieval framework that uses discourse processing techniques which aids in semantic analysis and representation of the Tamil Literary text. The proposed framework has been tested using two ancient literary works, the Thirukkural and Naladiyar, which were written during 300 BCE. The Thirukkural comprises 1330 couplets, each 7 words long, while the Naladiyar consists of 400 quatrains, each 15 words long. The proposed system, tested with all the 1330 Thirukkural couplets and 400 Naladiyar quatrains, achieved a mean average precision (MAP) score of 89%. The performance of the proposed framework has been compared with Google Tamil search and a keyword-based search which is a substandard version of the proposed framework. Google Tamil search achieved a MAP score of 56% and keyword-based method achieved a MAP score of 62% which shows that the discourse processing techniques improves the search performance of an Information Retrieval system.

References

Abraham, S. A. (2003). Chera, Chola, Pandya: Using archaeological evidence to identify the Tamil kingdoms of early historic South India. Asian Perspectives, 42(2), 207–223. https://doi.org/10.1353/asi.2003.0031

Adigalasiriyar (1985). Tolkāppiyam: Poruḷatikāram - ceyyuḷiyal [Tolkappiyam: Book of semantics – Chapter of poetry]. Tamil University Thanjavur.

Agosti, M., Marchesin, S., & Silvello, G. (2020). Learning unsupervised knowledge-enhanced representations to reduce the semantic gap in information retrieval. ACM Transactions on Information Systems (TOIS), 38(4), 1–48. https://doi.org/10.1145/3417996

Anandan, P., Saravanan, K., Parthasarathi, R., & Geetha, T. V. (2002, December). Morphological analyzer for Tamil. In International Conference on Natural Language Processing.

Anita, R., & Subalalitha, C. N. (2019a, July). An approach to cluster Tamil literatures using discourse connectives. In 2019 IEEE 1st International Conference on Energy, Systems and Information Processing (ICESIP) (pp. 1-4). IEEE. https://doi.org/10.1109/

ICESIP46348.2019.8938315 Journal of ICT, 20, No. 3 (July) 2021, pp: 353–

Anita, R., & Subalalitha, C. N. (2019b, December). Building discourse parser for Thirukkural. In Proceedings of the 16th International Conference on Natural Language Processing (ICON-2019) IIIT Hyderabad, India: NLP Association of India (pp. 18–25).

Fauzi, M. A., Arifin, A. Z., & Yuniarti, A. (2017). Arabic book retrieval using class and book index based term weighting. International Journal of Electrical & Computer Engineering, 7(6), 2088–8708. http://doi.org/10.11591/ijece.v7i6.pp3705-3710

Giridharan, R., Vellingiriraj, E. K., & Balasubramanie, P. (2016, April). Identification of Tamil ancient characters and information retrieval from temple epigraphy using image zoning. In 2016 International Conference on Recent Trends in Information Technology (ICRTIT) (pp. 1–7). IEEE. https://doi.org/10.1109/

ICRTIT.2016.7569600 ilearnTamil live online Tamil tuition: Tamil to Tamil Dictionary. (2018, June 15). Retrieved from https://ilearntamil.com/tamil-to-tamildictionary/

Liu, L., Liu, L., Fu, X., Huang, Q., Zhang, X., & Zhang, Y. (2018). A cloud-based framework for large-scale traditional Chinese medical record retrieval. Journal of Biomedical Informatics, 77, 21–33. https://doi.org/10.1016/j.jbi.2017.11.013

Mann, W. C., & Thompson, S. A. (1988). Rhetorical structure theory: Toward a functional theory of text organization. Text, 8(3), 243–281. https://doi.org/10.1515/text.1.1988.8.3.243

Meng, L., Tan, A. H., & Wunsch II, D. C. (2019). Online multimodal coindexing and retrieval of social media data. In Adaptive Resonance Theory in Social Media Data Clustering, (pp. 155–174). Springer, Cham. https://doi.org/10.1007/978-3-030-02985-2_7

Saravanan, M. S. (2020). Semantic document clustering based indexing for Tamil language information retrieval system. Journal of Critical Reviews, 7(14), 2999–3007. https://dx.doi.org/10.31838/ jcr.07.14.563

Prasath, R., Sarkar, S., & O’Reilly, P. (2015, April). Improving cross language information retrieval using corpus based query suggestion approach. In International Conference on Intelligent Text Processing and Computational Linguistics, (pp. 448 457). Springer, Cham. https://doi.org/10.1007/978-3-319-181172_33

Samia, Z., & Khaled, R. (2020). Multi-agents indexing system (MAIS) for plagiarism detection. Journal of King Saud University Journal of ICT, 20, No. 3 (July) 2021, pp: 353– Computer and Information Sciences. https://doi.org/10.1016/j. jksuci.2020.06.009 Sankaralingam, C., Rajendran, S., Kavirajan, B., Kumar, M. A., & Soman,

K. P. (2017, September). Onto-thesaurus for Tamil language: Ontology based intelligent system for information retrieval. In 2017 International Conference on Advances in Computing, Communications and Informatics (ICACCI) (pp. 2396–2396). IEEE. https://doi.org/10.1109/ICACCI.2017.8126206

Subalalitha, C. N. (2019). Information extraction framework for Kurunthogai. Sādhanā, 44(7), 1–6. https://doi.org/10.1007/ s12046-019-1140-y

Subalalitha, C. N., & Anita, R. (2016). An approach to page ranking based on discourse structures. Journal of Communications Software and Systems, 12(4), 195–200. http://dx.doi.org/10.24138/jcomss. v12i4.78

Tekli, J., Chbeir, R., Traina, A. J., & Traina Jr, C. (2019). SemIndex+: A semantic indexing scheme for structured, unstructured, and partly structured data. Knowledge-Based Systems, 164, 378–403. https://doi.org/10.1016/j.knosys.2018.11.010

Thenmozhi, D., & Aravindan, C. (2018). Ontology-based Tamil–English cross-lingual information retrieval system. Sādhanā, 43(10), 1–14. https://doi.org/10.1007/s12046-018-0942-7 Zamani, H., Dehghani, M., Croft, W. B., Learned-Miller, E., &

Kamps, J. (2018, October). From neural re-ranking to neural ranking: Learning a sparse representation for inverted indexing. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (pp. 497–506). https://doi.org/10.1145/3269206.3271800

Downloads

Published

11-06-2021

How to Cite

Ramalingam, A., & Navaneethakrish, S. C. (2021). A Discourse-Based Information Retrieval for Tamil Literary Texts. Journal of Information and Communication Technology, 20(3), 353-389. https://doi.org/10.32890/jict2021.20.3.4

Research impact

Harvested 2026-09-05
6 citations, from Scopus — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2021.20.3.4 OpenAlex W3175061013 Scopus 85109464874