Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention Alignment-Reference-Cited by-同舟云学术

Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention Alignment

Published:2024-05-10 Issue:5 Volume:23 Page:1-29
ISSN:2375-4699
Container-title:ACM Transactions on Asian and Low-Resource Language Information Processing
language:en
Short-container-title:ACM Trans. Asian Low-Resour. Lang. Inf. Process.

Author:

Rong Huan¹^ORCID,Chen Zhongfeng¹^ORCID,Lu Zhenyu¹^ORCID,Xu Fan²^ORCID,Sheng Victor S³^ORCID

Affiliation:

1. School of Artificial Intelligence, Nanjing University of Information Science & Technology, Nanjing, China

2. School of Computer and Infromation Engieering, Jiangxi Normal University, Nanchang, China

3. Department of Computer Science, Texas Tech University, Lubbock, USA

Abstract

This article focuses on the task of Multi-Modal Summarization with Multi-Modal Output for China JD.COM e-commerce product description containing both source text and source images. In the context learning of multi-modal (text and image) input, there exists a semantic gap between text and image, especially in the cross-modal semantics of text and image. As a result, capturing shared cross-modal semantics earlier becomes crucial for multi-modal summarization. However, when generating the multi-modal summarization, based on the different contributions of input text and images, the relevance and irrelevance of multi-modal contexts to the target summary should be considered, so as to optimize the process of learning cross-modal context to guide the summary generation process and to emphasize the significant semantics within each modality. To address the aforementioned challenges, Multization has been proposed to enhance multi-modal semantic information by multi-contextually relevant and irrelevant attention alignment. Specifically, a Semantic Alignment Enhancement mechanism is employed to capture shared semantics between different modalities (text and image), so as to enhance the importance of crucial multi-modal information in the encoding stage. Additionally, the IR-Relevant Multi-Context Learning mechanism is utilized to observe the summary generation process from both relevant and irrelevant perspectives, so as to form a multi-modal context that incorporates both text and image semantic information. The experimental results in the China JD.COM e-commerce dataset demonstrate that the proposed Multization method effectively captures the shared semantics between the input source text and source images, and highlights essential semantics. It also successfully generates the multi-modal summary (including image and text) that comprehensively considers the semantics information of both text and image.

Funder

National Natural Science Foundation of China

Natural Science Foundation of Jiangsu Province

Graduate Research and Innovation Projects of Jiangsu Province

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3651983

Reference48 articles.

1. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

2. Multimodal Machine Learning: A Survey and Taxonomy

3. Multimedia Summarization for Social Events in Microblog Stream

4. Abstractive Text-Image Summarization Using Multi-Modal Attentional Hierarchical RNN

5. Real20M: A Large-scale E-commerce Dataset for Cross-domain Retrieval