A Video Classification Method Based on Spatiotemporal Detail Attention and Feature Fusion-Reference-Cited by-同舟云学术

A Video Classification Method Based on Spatiotemporal Detail Attention and Feature Fusion

Published:2022-04-28 Issue: Volume:2022 Page:1-10
ISSN:1875-905X
Container-title:Mobile Information Systems
language:en
Short-container-title:Mobile Information Systems

Author:

Gong Xuchao¹^ORCID,Li Zongmin¹^ORCID

Affiliation:

1. School of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China

Abstract

With the explosive growth of Internet video data, demands for accurate large-scale video classification and management are increasing. In the real-world deployment, the balance between effectiveness and timeliness should be fully considered. Generally, the video classification algorithm equipped with time segment network is used in industrial deployment, and the frame extraction feature is used to classify video actions However, the issue of semantic deviation will be raised due to coarse feature description. In this paper, we propose a novel method, called image dense feature and internal significant detail description, to enhance the generalization and discrimination of feature description. Specifically, the location information layer of space-time geometric relationship is added to effectively engrave the local features of convolution layer. Moreover, the multimodal feature graph network is introduced to effectively improve the generalization ability of feature fusion. Extensive experiments show that the proposed method can effectively improve the results on two commonly used benchmarks (kinetics 400 and kinetics 600).

Funder

National Key R&D Plan of China

Publisher

Hindawi Limited

Subject

Computer Networks and Communications,Computer Science Applications

Link

http://downloads.hindawi.com/journals/misy/2022/4213335.pdf

Reference52 articles.

1. Temporal Segment Networks: Towards Good Practices for Deep Action Recognition

2. Tsm: temporal shift module for efficient video understanding;J. Lin

3. Slowfast networks for video recognition;C. Feichtenhofer

4. Movinets: Mobile video networks for efficient video recognition;D. Kondratyuk

5. AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Text Segmentation Algorithm Focused on Corpus Mining for Oilfield Exploration and Development;2024 9th International Conference on Computer and Communication Systems (ICCCS);2024-04-19