Summarization assessment methodology for multiple corpora using queries and classification for functional evaluation-Reference-Cited by-同舟云学术

Summarization assessment methodology for multiple corpora using queries and classification for functional evaluation

Published:2022-06-21 Issue:3 Volume:29 Page:227-239
ISSN:1069-2509
Container-title:Integrated Computer-Aided Engineering
language:
Short-container-title:ICA

Author:

Wolyn Sam,Simske Steven J.

Abstract

Extractive summarization is an important natural language processing approach used for document compression, improved reading comprehension, key phrase extraction, indexing, query set generation, and other analytics approaches. Extractive summarization has specific advantages over abstractive summarization in that it preserves style, specific text elements, and compound phrases that might be more directly associated with the text. In this article, the relative effectiveness of extractive summarization is considered on two widely different corpora: (1) a set of works of fiction (100 total, mainly novels) available from Project Gutenberg, and (2) a large set of news articles (3000) for which a ground truthed summarization (gold standard) is provided by the authors of the news articles. Both sets were evaluated using 5 different Python Sumy algorithms and compared to randomly-generated summarizations quantitatively. Two functional approaches to assessing the efficacy of summarization using a query set on both the original documents and their summaries, and using document classification on a 12-class set to compare among different summarization approaches, are introduced. The results, unsurprisingly, show considerable differences consistent with the different nature of these two data sets. The LSA and Luhn summarization approaches were most effective on the database of fiction, while all five summarization approaches were similarly effective on the database of articles. Overall, the Luhn approach was deemed the most generally relevant among those tested.

Publisher

IOS Press

Subject

Artificial Intelligence,Computational Theory and Mathematics,Computer Science Applications,Theoretical Computer Science,Software

Reference30 articles.

1. Abstractive summarization: An overview of the state of the art;Gupta;Expert Systems with Applications.,2019

2. The automatic creation of literature abstracts;Luhn;IBM Journal of Research and Development.,1958

3. Automatic abstracting and indexing – survey and recommendations;Edmundson;Communications of the ACM.,1961

4. New methods in automatic extracting;Edmundson;Journal of the ACM (JACM).,1969

5. Automatic abstracting and indexing. II. Production of indicative abstracts by application of contextual inference and syntactic coherence criteria;Rush;Journal of the American Society for Information Science.,1971

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. IndicBART Alongside Visual Element: Multimodal Summarization in Diverse Indian Languages;Lecture Notes in Computer Science;2024

2. Algorithm Parallelism for Improved Extractive Summarization;Proceedings of the ACM Symposium on Document Engineering 2023;2023-08-22

3. Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection;International Journal of Neural Systems;2023-06-14

4. A Modified Long Short-Term Memory Cell;International Journal of Neural Systems;2023-06-09