Temporally Multi-Modal Semantic Reasoning with Spatial Language Constraints for Video Question Answering-Reference-Cited by-同舟云学术

Temporally Multi-Modal Semantic Reasoning with Spatial Language Constraints for Video Question Answering

Published:2022-05-31 Issue:6 Volume:14 Page:1133
ISSN:2073-8994
Container-title:Symmetry
language:en
Short-container-title:Symmetry

Author:

Liu Mingyang,Wang Ruomei^ORCID,Zhou Fan,Lin Ge

Abstract

Video question answering (QA) aims to understand the video scene and underlying plot by answering video questions. An algorithm that can competently cope with this task needs to be able to: (1) collect multi-modal information scattered in the video frame sequence while extracting, interpreting, and utilizing the potential semantic clues provided by each piece of modal information in the video, (2) integrate the multi-modal context of the above semantic clues and understand the cause and effect of the story as it evolves, and (3) identify and integrate those temporally adjacent or non-adjacent effective semantic clues implied in the above context information to provide reasonable and sufficient visual semantic information for the final question reasoning. In response to the above requirements, a novel temporally multi-modal semantic reasoning with spatial language constraints video QA solution is reported in this paper, which includes a significant feature extraction module used to extract multi-modal features according to a significant sampling strategy, a spatial language constraints module used to recognize and reason spatial dimensions in video frames under the guidance of questions, and a temporal language interaction module used to locate the temporal dimension semantic clues of the appearance features and motion features sequence. Specifically, for a question, the result processed by the spatial language constraints module is to obtain visual clues related to the question from a single image and filter out unwanted spatial information. Further, the temporal language interaction module symmetrically integrates visual clues of the appearance information and motion information scattered throughout the temporal dimensions, obtains the temporally adjacent or non-adjacent effective semantic clue, and filters out irrelevant or detrimental context information. The proposed video QA solution is validated on several video QA benchmarks. Comprehensive ablation experiments have confirmed that modeling the significant video information can improve QA ability. The spatial language constraints module and temporal language interaction module can better collect and summarize visual semantic clues.

Funder

National Key R&D Program of China

Publisher

MDPI AG

Subject

Physics and Astronomy (miscellaneous),General Mathematics,Chemistry (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2073-8994/14/6/1133/pdf

Reference53 articles.

1. Hybrid Image-Retrieval Method for Image-Splicing Validation

2. Image Caption Generation Using Multi-Level Semantic Context Information

3. Video Caption Based Searching Using End-to-End Dense Captioning and Sentence Embeddings

4. AnswerNet: Learning to Answer Questions

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Advancing Video Question Answering with a Multi-modal and Multi-layer Question Enhancement Network;Proceedings of the 31st ACM International Conference on Multimedia;2023-10-26