Visual Summary Identification From Scientific Publications via Self-Supervised Learning-Reference-Cited by-同舟云学术

Visual Summary Identification From Scientific Publications via Self-Supervised Learning

Published:2021-08-19 Issue: Volume:6 Page:
ISSN:2504-0537
Container-title:Frontiers in Research Metrics and Analytics
language:
Short-container-title:Front. Res. Metr. Anal.

Author:

Yamamoto Shintaro,Lauscher Anne,Ponzetto Simone Paolo,Glavaš Goran,Morishima Shigeo

Abstract

The exponential growth of scientific literature yields the need to support users to both effectively and efficiently analyze and understand the some body of research work. This exploratory process can be facilitated by providing graphical abstracts–a visual summary of a scientific publication. Accordingly, previous work recently presented an initial study on automatic identification of a central figure in a scientific publication, to be used as the publication’s visual summary. This study, however, have been limited only to a single (biomedical) domain. This is primarily because the current state-of-the-art relies on supervised machine learning, typically relying on the existence of large amounts of labeled data: the only existing annotated data set until now covered only the biomedical publications. In this work, we build a novel benchmark data set for visual summary identification from scientific publications, which consists of papers presented at conferences from several areas of computer science. We couple this contribution with a new self-supervised learning approach to learn a heuristic matching of in-text references to figures with figure captions. Our self-supervised pre-training, executed on a large unlabeled collection of publications, attenuates the need for large annotated data sets for visual summary identification and facilitates domain transfer for this task. We evaluate our self-supervised pretraining for visual summary identification on both the existing biomedical and our newly presented computer science data set. The experimental results suggest that the proposed method is able to outperform the previous state-of-the-art without any task-specific annotations.

Funder

Japan Science and Technology Agency

Japan Society for the Promotion of Science

Ministry of Education, Culture, Sports, Science and Technology

Publisher

Frontiers Media SA

Reference46 articles.

1. SciBERT: A pretrained language model for scientific text;Beltagy,2019

2. Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references;Bornmann;J. Assoc. Inf. Sci. Tech.,2015

3. A large annotated corpus for learning natural language inference;Bowman,2015

4. Neural summarization by extracting sentences and words;Cheng,2016

5. A discourse-aware attention model for abstractive summarization of long documents;Cohan,2018

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluating efficiency and accuracy of deep-learning-based approaches on study selection for psychiatry systematic reviews;Nature Mental Health;2023-08-31

2. Cross-lingual extreme summarization of scholarly documents;International Journal on Digital Libraries;2023-08-10

3. X-SCITLDR;Proceedings of the 22nd ACM/IEEE Joint Conference on Digital Libraries;2022-06-20