Investigation of Pre-Trained Bidirectional Encoder Representations from Transformers Checkpoints for Indonesian Abstractive Text Summarization

Authors

  • Henry Lucky Computer Science Department, BINUS Graduate Program, Master of Computer Science, Bina Nusantara University, Jakarta, Indonesia
  • Derwin Suhartono Computer Science Department, School of Computer Science, Bina Nusantara University, Jakarta, Indonesia

DOI:

https://doi.org/10.32890/jict2022.21.1.4

Keywords:

Abstractive text summarization, BERT, Indonesian, Natural language processing, Transformer

Abstract

Text summarization aims to reduce text by removing less useful information to obtain information quickly and precisely. In Indonesian
abstractive text summarization, the research mostly focuses on multi-document summarization which methods will not work optimally in single-document summarization. As the public summarization datasets and works in English are focusing on single-document summarization, this study emphasized on Indonesian single-document summarization. Abstractive text summarization studies in English frequently use Bidirectional Encoder Representations from Transformers (BERT), and since Indonesian BERT checkpoint is available, it was employed in this study. This study investigated the use of Indonesian BERT in abstractive text summarization on
the IndoSum dataset using the BERTSum model. The investigation proceeded by using various combinations of model encoders, model embedding sizes, and model decoders. Evaluation results showed that models with more embedding size and used Generative Pre-Training (GPT)-like decoder could improve the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score and BERTScore of the model results.

References

Adelia, R., Suyanto, S., & Wisesty, U. N. (2019). Indonesian abstractive text summarization using bidirectional gated Journal of ICT, 21, No. 1 (January) 2022, pp: 71– recurrent unit. Procedia Computer Science, 157, 581–588. https://doi.org/10.1016/j.procs.2019.09.017

Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer normalization. ArXivPreprint ArXiv:1607.06450. https://arxiv.org/abs/1607.06450

Cai, Z., Lin, N., Ma, C., & Jiang, S. (2019). Indonesian automatic text summarization based on a new clustering method in sentence level. In Proceedings of the 2019 International Conference on Big Data Engineering (pp. 30–35). https://doi.org/10.1145/3341620.3341626

Cheng, J., & Lapata, M. (2016). Neural summarization by extracting sentences and words. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 484–494). https://dx.doi.org/10.18653/v1/ P16-1046

Christian, H., Agus, M. P., & Suhartono, D. (2016). Single document automatic text summarization using term frequency-inverse document frequency (TF-IDF). ComTech: Computer, Mathematics and Engineering Applications, 7(4), 285–294. https://doi.org/10.21512/comtech.v7i4.3746

Christie, F., & Khodra, M. L. (2016). Multi-document summarization using sentence fusion for Indonesian news articles. In 2016 International Conference On Advanced Informatics: Concepts, Theory and Application (ICAICTA) (pp. 1–6). IEEE. https://doi.org/10.1109/ICAICTA.2016.7803134

Cui, Y., Che, W., Liu, T., Qin, B., Yang, Z., Wang, S., & Hu, G. (2019). Pre-training with whole word masking for Chinese BERT. ArXiv Preprint ArXiv:1906.08101. https://arxiv.org/ abs/1906.08101

Devianti, R. S., & Khodra, M. L. (2019). Abstractive summarization using genetic semantic graph for indonesian news articles. In 2019 International Conference of Advanced Informatics: Concepts, Theory and Applications (ICAICTA) (pp. 1–6). IEEE. https://doi.org/10.1109/ICAICTA.2019.8904361

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. ArXiv Preprint ArXiv:1810.04805. https://arxiv.org/abs/1810.04805

Garmastewira, G., & Khodra, M. L. (2019). Summarizing Indonesian news articles using Graph Convolutional Network. Journal of Information and Communication Technology, 18(3), 345–365. https://doi.org/10.32890/jict2019.18.3.4675 Journal of ICT, 21, No. 1 (January) 2022, pp: 71–

Grusky, M., Naaman, M., & Artzi, Y. (2018). Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies. ArXiv Preprint ArXiv:1804.11283. https://arxiv.org/ abs/1804.11283

Halim, K., Novianus Palit, H., & Tjondrowiguno, A. N. (2020). Penerapan recurrent neural network untuk pembuatan ringkasan ekstraktif otomatis pada berita berbahasa Indonesia. Jurnal Infra, 8(1), 221–227. http://publication.petra.ac.id/ index.php/teknik-informatika/article/view/9797

Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., & Blunsom, P. (2015). Teaching machines to read and comprehend. Advances in Neural Information Processing Systems, 28, 1693–1701. https://ora.ox.ac.uk/ objects/uuid:050e7840-1ff3-49db-8d36-e83ed0adf8f7

Hidayat, E. Y., Firdausillah, F., Hastuti, K., Dewi, I. N., & Azhari, A. (2015). Automatic text summarization using latent Drichlet allocation (LDA) for document clustering. International Journal of Advances in Intelligent Informatics, 1(3), 132–139. https://doi.org/10.26555/ijain.v1i3.43

Hoang, A., Bosselut, A., Celikyilmaz, A., & Choi, Y. (2019). Efficient adaptation of pretrained transformers for abstractive summarization. ArXiv Preprint ArXiv:1906.00138. https://arxiv.org/abs/1906.00138

Ilyas, R. (2015). Peringkas otomatis dengan ekstraksi informasi untuk kumpulan berita online (Tesis Magister Institut Teknologi Bandung). Institut Teknologi Bandung, Bandung.

Koto, F., Lau, J. H., & Baldwin, T. (2020). Liputan6: A large-scale Indonesian dataset for text summarization. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 598–608). https://aclanthology.org/2020.aacl-main.60

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). https://dx.doi.org/10.18653/v1/2020.coling-main.66

Kurniawan, K., & Louvan, S. (2018). Indosum: A new benchmark dataset for Indonesian text summarization. In 2018 International Conference on Asian Language Processing (IALP) (pp. 215–220). IEEE. https://doi.org/10.1109/IALP.2018.8629109 Journal of ICT, 21, No. 1 (January) 2022, pp: 71–

Lewis, M., Aghajanyan, A., Ghazvininejad, M., Wang, S., Ghosh, G., & Zettlemoyer, L. (2020). Pre-training via paraphrasing. ArXiv Preprint ArXiv:2006.15020. https://arxiv.org/abs/2006.15020

Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2019). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. ArXiv Preprint ArXiv:1910.13461. https://doi.org/10.18653/v1/2020. acl-main.703

Lin, C.-Y. (2004). Rouge: A package for automatic evaluation of summaries. Text Summarization Branches Out, 74–81. https://www.aclweb.org/anthology/W04-1013.pdf

Liu, Y., & Lapata, M. (2020). Text summarization with pretrained encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3721–3731). https://dx.doi. org/10.18653/v1/D19-1387

Martin, L., Muller, B., Suárez, P. J. O., Dupont, Y., Romary, L., de la Clergerie, É. V., Seddah, D., & Sagot, B. (2019). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7203–7219). http://dx.doi.org/10.18653/ v1/2020.acl-main.645

Najibullah, A. (2015). Indonesian text summarization based on naïve bayes method. In Proceeding of he International Seminar and Conference on Global Issues (Vol. 1, No. 1). https://www.publikasiilmiah.unwahas.ac.id/index.php/ISC/article/ view/1265/1366

Nallapati, R., Xiang, B., & Zhou, B. (2016a). Sequence-to-sequence RNNs for text summarization. ICLR 2016, 4–7.

Nallapati, R., Zhou, B., dos Santos, C., Gulçehre, Ç., & Xiang, B. (2016b). Abstractive text summarization using sequence-to-sequence RNNs and beyond. ArXiv Preprint ArXiv:1602.06023. https://arxiv.org/abs/1602.06023

Narayan, S., Cohen, S. B., & Lapata, M. (2018). Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 1797–1807). https://dx.doi.org/10.18653/v1/ D18-1206 Journal of ICT, 21, No. 1 (January) 2022, pp: 71–

Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (pp. 2227–2237). https://dx.doi.org/10.18653/v1/N18-1202

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9. http://www.persagen.com/files/misc/ radford2019language.pdf

Rönnqvist, S., Kanerva, J., Salakoski, T., & Ginter, F. (2019). Is multilingual BERT fluent in language generation? DL4NLP 2019, 29. https://ep.liu.se/ecp/163/ecp19163.pdf#page=35

Rothe, S., Narayan, S., & Severyn, A. (2020). Leveraging pre-trained checkpoints for sequence generation tasks. Transactions of the Association for Computational Linguistics, 8, 264–280. https://doi.org/10.1162/tacl_a_00313

Rush, A. M., Chopra, S., & Weston, J. (2015). A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (pp. 379-389). https://dx.doi.org/10.18653/v1/D15-1044

Savelieva, A., Au-Yeung, B., & Ramani, V. (2020). Abstractive summarization of spoken and written instructions with BERT. ArXiv Preprint ArXiv:2008.09676. https://arxiv.org/ abs/2008.09676

See, A., Liu, P. J., & Manning, C. D. (2017). Get to the point: Summarization with pointer-generator networks. ArXiv Preprint ArXiv:1704.04368. https://arxiv.org/abs/1704.04368

Severina, V., & Khodra, M. L. (2019). Multidocument abstractive summarization using abstract meaning representation for Indonesian language. In 2019 International Conference of Advanced Informatics: Concepts, Theory and Applications (ICAITA) (pp.1-6). IEEE. https://doi.org/10.1109/ICAICTA.2019. 8904449

Shi, T., Keneshloo, Y., Ramakrishnan, N., & Reddy, C. K. (2021). Neural abstractive text summarization with sequence-to-sequence models. ACM Transactions on Data Science, 2, 1(1), 1–37. https://doi.org/10.1145/3419106 Journal of ICT, 21, No. 1 (January) 2022, pp: 71–

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. ArXiv Preprint ArXiv:1706.03762. https://arxiv.org/ abs/1706.03762

Widyassari, A. P., Affandy, A., Noersasongko, E., Fanani, A. Z., Syukur, A., & Basuki, R. S. (2019, July). Literature review of automatic text summarization: Research trend, dataset and method. In 2019 International Conference on Information and Communications Technology (ICOIACT) (pp. 491–496). IEEE. https://doi.org/10.1109/ICOIACT46704.2019.8938454

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 843–857). https://www.aclweb.org/ anthology/2020.aacl-main.

Zhang, H., Cai, J., Xu, J., & Wang, J. (2019). Pretraining-based natural language generation for text summarization. In CoNLL 2019-23rd Conference on Computational Natural Language Learning, Proceedings of the Conference (pp. 789–797). http://dx.doi.org/10.18653/v1/K19-1074

Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). BERTScore: Evaluating text generation with BERT. ArXiv Preprint ArXiv:1904.09675. https://arxiv.org/abs/1904.09675

Zhou, Q., Yang, N., Wei, F., & Zhou, M. (2017). Selective encoding for abstractive sentence summarization. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1095–1104). https://dx.doi.org/10.18653/v1/P17-1101

Downloads

Published

11-11-2021

How to Cite

Lucky, H., & Suhartono, D. (2021). Investigation of Pre-Trained Bidirectional Encoder Representations from Transformers Checkpoints for Indonesian Abstractive Text Summarization. Journal of Information and Communication Technology, 21(1), 71-94. https://doi.org/10.32890/jict2022.21.1.4

Research impact

Harvested 2026-09-05
9 citations, from Scopus — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2022.21.1.4 OpenAlex W3213696232 Scopus 85120968459