Exploiting Instance-level Relationships in Weakly Supervised Text-to-Video Retrieval-Reference-Cited by-同舟云学术

Exploiting Instance-level Relationships in Weakly Supervised Text-to-Video Retrieval

Published:2024-09-12 Issue:10 Volume:20 Page:1-21
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Yin Shukang¹^ORCID,Zhao Sirui²^ORCID,Wang Hao¹^ORCID,Xu Tong¹^ORCID,Chen Enhong³^ORCID

Affiliation:

1. School of Data Science, University of Science and Technology of China, Hefei, China

2. School of Computer Science and Technology, University of Science and Technology of China, Hefei, China and School of Computer Science and Technology, Southwest University of Science and Technology, Mianyang, China

3. School of Data Science, University of Science and Technology of China, Hefei China

Abstract

Text-to-Video Retrieval is a typical cross-modal retrieval task that has been studied extensively under a conventional supervised setting. Recently, some works have sought to extend the problem to a weakly supervised formulation, which can be more consistent with real-life scenarios and more efficient in annotation cost. In this context, a new task called Partially Relevant Video Retrieval (PRVR) is proposed, which aims to retrieve videos that are partially relevant to a given textual query, i.e., the videos containing at least one semantically relevant moment. Formulating the task as a Multiple Instance Learning (MIL) ranking problem, prior arts rely on heuristics algorithms such as a simple greedy search strategy and deal with each query independently. Although these early explorations have achieved decent performance, they may not fully utilize the bag-level label and only consider the local optimum, which could result in suboptimal solutions and inferior final retrieval performance. To address this problem, in this paper, we propose to exploit the relationships between instances to boost retrieval performance. Based on this idea, we creatively put forward: (1) a new matching scheme for pairing queries and their related moments in the video; and (2) a new loss function to facilitate cross-modal alignment between two views of an instance. Extensive validations on three publicly available datasets have demonstrated the effectiveness of our solution and verified our hypothesis that modeling instance-level relationships is beneficial in the MIL ranking setting. Our code will be publicly available at https://github.com/xjtupanda/BGM-Net .

Funder

National Natural Science Foundation of China

Young Scientists Fund of the Natural Science Foundation of Sichuan Province

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3663571

Reference82 articles.

1. Robert A. Amar, Daniel R. Dooly, Sally A. Goldman, and Qi Zhang. 2001. Multiple-instance learning of real-valued data. In ICML.

2. Localizing Moments in Video with Natural Language

3. Neural machine translation by jointly learning to align and translate;Bahdanau Dzmitry;arXiv preprint arXiv:1409.0473,2014

4. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

5. Fast bundle algorithm for multiple-instance learning;Bergeron Charles;IEEE Transactions on Pattern Analysis and Machine Intelligence,2011