Hybrid-attention and frame difference enhanced network for micro-video venue recognition-Reference-Cited by-同舟云学术

Hybrid-attention and frame difference enhanced network for micro-video venue recognition

Published:2022-07-21 Issue:3 Volume:43 Page:3337-3353
ISSN:1064-1246
Container-title:Journal of Intelligent & Fuzzy Systems
language:
Short-container-title:IFS

Author:

Wang Bing¹,Huang Xianglin¹,Cao Gang¹,Yang Lifang¹,Wei Xiaolong¹,Tao Zhulin¹

Affiliation:

1. State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China

Abstract

Many micro-video related applications, such as personalized location recommendation and micro-video verification, can be benefited greatly from the venue information. Most existing works focus on integrating the information from multi-modal for exact venue category recognition. It is important to make full use of the information from different modalities. However, the performance may be limited by the lacked acoustic modality or textual descriptions in uploaded micro-videos. Therefore, in this paper visual modality is explored as the only modality according to its rich and indispensable semantic information. To this end, a hybrid-attention and frame difference enhanced network (HAFDN) is proposed to generate the comprehensive venue representation. Such network mainly contains two parallel branches: content and motion branches. Specifically, in the content branch, a domain-adaptive CNN model combined with temporal shift module (TSM) is employed to extract discriminative visual features. Then, a novel hybrid attention module (HAM) is introduced to enhance extracted features via three attention mechanisms. In HAM, channel attention, local and global spatial attention mechanisms are used to capture salient visual information from different views. In addition, convolutional Long Short-Term Memory (convLSTM) is enforced after HAM to better encode the long spatial-temporal dependency. A difference-enhanced module parallel with HAM is devised to learn the content variations among adjacent frames, which is usually ignored in prior works. Moreover, in the motion branch, 3D-CNNs and LSTM are used to capture movement variation as a supplement of content branch in a different form. Finally, the features from two branches are fused to generate robust video-level representations for predicting venue categories. Extensive experimental results on public datasets verify the effectiveness of the proposed micro-video venue recognition scheme. The source code is available at https://github.com/hs8945/HAFDN.

Publisher

IOS Press

Subject

Artificial Intelligence,General Engineering,Statistics and Probability

Reference48 articles.

1. Hierarchy-dependent cross-platform multi-view feature learning for venue category prediction;Jiang;IEEE Trans Multim,2019

2. Neural multimodal cooperative learning toward micro-video understanding;Wei;IEEE Trans Image Process,2020

3. Attention based consistent semantic learning for micro-video scene recognition;Guo;Inf Sci,2021

4. Hybrid-attention enhanced two-stream fusion network for video venue prediction;Zhang;IEEE Trans Multim,2021

5. Zhang J. , Nie L. , Wang X. , HeX, HuangX. and ChuaT., Shorter-is-better: Venue category estimation from micro-video, In Proceedings of the 2016 ACM Conference on Multimedia Conference, pages 1415–1424. ACM, 2016.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A survey of micro-video analysis;Multimedia Tools and Applications;2023-09-20