Object Relation Attention for Image Paragraph Captioning-Reference-Cited by-同舟云学术

Object Relation Attention for Image Paragraph Captioning

Published:2021-05-18 Issue:4 Volume:35 Page:3136-3144
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Yang Li-Chuan,Yang Chih-Yuan,Hsu Jane Yung-jen

Abstract

Image paragraph captioning aims to automatically generate a paragraph from a given image. It is an extension of image captioning in terms of generating multiple sentences instead of a single one, and it is more challenging because paragraphs are longer, more informative, and more linguistically complicated. Because a paragraph consists of several sentences, an effective image paragraph captioning method should generate consistent sentences rather than contradictory ones. It is still an open question how to achieve this goal, and for it we propose a method to incorporate objects' spatial coherence into a language-generating model. For every two overlapping objects, the proposed method concatenates their raw visual features to create two directional pair features and learns weights optimizing those pair features as relation-aware object features for a language-generating model. Experimental results show that the proposed network extracts effective object features for image paragraph captioning and achieves promising performance against existing methods.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. FDT − Dr2T: a unified Dense Radiology Report Generation Transformer framework for X-ray images;Machine Vision and Applications;2024-05-21

2. A comprehensive literature review on image captioning methods and metrics based on deep learning technique;Multimedia Tools and Applications;2024-02-20

3. Image paragraph captioning with topic clustering and topic shift prediction;Knowledge-Based Systems;2024-02

4. DE-GAN: Text-to-image synthesis with dual and efficient fusion model;Multimedia Tools and Applications;2023-08-18

5. Evolution of visual data captioning Methods, Datasets, and evaluation Metrics: A comprehensive survey;Expert Systems with Applications;2023-07