Improving <scp>ROUGE</scp>‐1 by 6%: A novel multilingual transformer for abstractive news summarization-Reference-Cited by-同舟云学术

Improving ROUGE‐1 by 6%: A novel multilingual transformer for abstractive news summarization

Published:2024-06-10 Issue:20 Volume:36 Page:
ISSN:1532-0626
Container-title:Concurrency and Computation: Practice and Experience
language:en
Short-container-title:Concurrency and Computation

Author:

Kumar Sandeep¹^ORCID,Solanki Arun¹

Affiliation:

1. Department of Computer Science and Engineering, SoICT Gautam Buddha University Greater Noida India

Abstract

SummaryNatural language processing (NLP) has undergone a significant transformation, evolving from manually crafted rules to powerful deep learning techniques such as transformers. These advancements have revolutionized various domains including summarization, question answering, and more. Statistical models like hidden Markov models (HMMs) and supervised learning have played crucial roles in laying the foundation for this progress. Recent breakthroughs in transfer learning and the emergence of large‐scale models like BERT and GPT have further pushed the boundaries of NLP research. However, news summarization remains a challenging task in NLP, often resulting in factual inaccuracies or the loss of the article's essence. In this study, we propose a novel approach to news summarization utilizing a fine‐tuned Transformer architecture pre‐trained on Google's mt‐small tokenizer. Our model demonstrates significant performance improvements over previous methods on the Inshorts English News dataset, achieving a 6% enhancement in the ROUGE‐1 score and reducing training loss by 50%. This breakthrough facilitates the generation of reliable and concise news summaries, thereby enhancing information accessibility and user experience. Additionally, we conduct a comprehensive evaluation of our model's performance using popular metrics such as ROUGE scores, with our proposed model achieving ROUGE‐1: 54.6130, ROUGE‐2: 31.1543, ROUGE‐L: 50.7709, and ROUGE‐LSum: 50.7907. Furthermore, we observe a substantial reduction in training and validation losses, underscoring the effectiveness of our proposed approach.

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/cpe.8199

Reference68 articles.

1. MayhewS TsygankovaT RothD.Ner and pos when nothing is capitalized. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP‐IJCNLP) Hong Kong China: Association for Computational Linguistics; 2019: 6255–6260. doi:10.18653/v1/D19‐1650

2. LopezMM KalitaJ.Deep Learning Applied to NLP; 2017. doi:10.48550/ARXIV.1703.03091

3. DevlinJ ChangM‐W LeeK ToutanovaK.BERT: pre‐training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers) Minneapolis Minnesota: Association for Computational Linguistics; 2019: 4171–4186. doi:10.18653/v1/N19‐1423

4. An empirical study of smoothing techniques for language modeling