Cross-Modal Hybrid Feature Fusion for Image-Sentence Matching-Reference-Cited by-同舟云学术

Cross-Modal Hybrid Feature Fusion for Image-Sentence Matching

Published:2021-11-30 Issue:4 Volume:17 Page:1-23
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Xu Xing¹,Wang Yifan¹,He Yixuan¹,Yang Yang¹,Hanjalic Alan²,Shen Heng Tao¹

Affiliation:

1. University of Electronic Science and Technology of China, Chengdu, China

2. Delft University of Technology, Delft, The Netherlands

Abstract

Image-sentence matching is a challenging task in the field of language and vision, which aims at measuring the similarities between images and sentence descriptions. Most existing methods independently map the global features of images and sentences into a common space to calculate the image-sentence similarity. However, the image-sentence similarity obtained by these methods may be coarse as (1) an intermediate common space is introduced to implicitly match the heterogeneous features of images and sentences in a global level, and (2) only the inter-modality relations of images and sentences are captured while the intra-modality relations are ignored. To overcome the limitations, we propose a novel Cross-Modal Hybrid Feature Fusion (CMHF) framework for directly learning the image-sentence similarity by fusing multimodal features with inter- and intra-modality relations incorporated. It can robustly capture the high-level interactions between visual regions in images and words in sentences, where flexible attention mechanisms are utilized to generate effective attention flows within and across the modalities of images and sentences. A structured objective with ranking loss constraint is formed in CMHF to learn the image-sentence similarity based on the fused fine-grained features of different modalities bypassing the usage of intermediate common space. Extensive experiments and comprehensive analysis performed on two widely used datasets—Microsoft COCO and Flickr30K—show the effectiveness of the hybrid feature fusion framework in CMHF, in which the state-of-the-art matching performance is achieved by our proposed CMHF method.

Funder

National Natural Science Foundation of China

Fundamental Research Funds for the Central Universities

Sichuan Science and Technology Program, China

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture

Link

https://dl.acm.org/doi/pdf/10.1145/3458281

Reference67 articles.

1. Large-Scale Machine Learning with Stochastic Gradient Descent

2. Top-Down Versus Bottom-Up Control of Attention in the Prefrontal and Posterior Parietal Cortices

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Lightweight two-stage transformer for low-light image enhancement and object detection;Digital Signal Processing;2024-07

2. PAR-Net: An Enhanced Dual-Stream CNN–ESN Architecture for Human Physical Activity Recognition;Sensors;2024-03-16

3. Deep Convolutional Neural Network Compression Method: Tensor Ring Decomposition with Variational Bayesian Approach;Neural Processing Letters;2024-03-13

4. EMNet: Edge-guided multi-level network for salient object detection in low-light images;Image and Vision Computing;2024-03

5. Image–Text Cross-Modal Retrieval with Instance Contrastive Embedding;Electronics;2024-01-09