NumCap: A Number-controlled Multi-caption Image Captioning Network-Reference-Cited by-同舟云学术

NumCap: A Number-controlled Multi-caption Image Captioning Network

Published:2023-02-27 Issue:4 Volume:19 Page:1-24
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Abdussalam Amr¹^ORCID,Ye Zhongfu¹^ORCID,Hawbani Ammar¹^ORCID,Al-Qatf Majjed¹^ORCID,Khan Rashid¹^ORCID

Affiliation:

1. University of Science and Technology of China, Hefei, China

Abstract

Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV’s embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.

Funder

CAS-TWAS President’s Fellowship for Ph.D.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture

Link

https://dl.acm.org/doi/pdf/10.1145/3576927

Reference61 articles.

1. Image Captioning with Novel Topics Guidance and Retrieval-based Topics Re-weighting

2. Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap

3. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering