MAFormer: A cross-channel spatio-temporal feature aggregation method for human action recognition-Reference-Cited by-同舟云学术

MAFormer: A cross-channel spatio-temporal feature aggregation method for human action recognition

Published:2024-09-09 Issue: Volume: Page:1-15
ISSN:1875-8452
Container-title:AI Communications
language:
Short-container-title:AIC

Author:

Huang Hongbo¹²,Xu Longfei¹,Zheng Yaolin¹,Yan Xiaoxu¹

Affiliation:

1. Computer School, Beijing Information Science & Technology University, Beijing, China

2. Institute of Computing Intelligence, Beijing Information Science & Technology University, Beijing, China

Abstract

Human action recognition has been widely used in fields such as human–computer interaction and virtual reality. Despite significant progress, existing approaches still struggle with effectively integrating hierarchical information and processing data beyond a certain frame count. To address these challenges, we introduce the Multi-AxisFormer (MAFormer) model, which is organized in terms of spatial, temporal, and channel dimensions of the action sequence, thereby enhancing the model’s understanding of correlations and intricate structures among and within features. Drawing on the Transformer architecture, we propose the Cross-channel Spatio-temporal Aggregation (CSA) structure for more refined feature extraction and the Multi-Axis Attention (MAA) module for more comprehensive feature aggregation. Moreover, the integration of Rotary Position Embedding (RoPE) boosts the model’s extrapolation and generalization abilities. MAFormer surpasses the known state-of-the-art on multiple skeleton-based action recognition benchmarks with the accuracy of 93.2% on NTU RGB+D 60 cross-subject split, 89.9% on NTU RGB+D 120 cross-subject split, and 97.2% on N-UCLA, offering a novel paradigm for hierarchical modeling in human action recognition.

Publisher

IOS Press

Reference33 articles.

1. D. Ahn, S. Kim, H. Hong and B.C. Ko, Star-transformer: A spatio-temporal cross attention transformer for human action recognition, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 3330–3339.

2. Iris and Foot based Sustainable Biometric Identification Approach

3. Swin-fusion: Swin-transformer with feature fusion for human action recognition;Chen;Neural Processing Letters,2023

4. Y. Chen, Z. Zhang, C. Yuan, B. Li, Y. Deng and W. Hu, Channel-wise topology refinement graph convolution for skeleton-based action recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.

5. Z. Chen, S. Li, B. Yang, Q. Li and H. Liu, Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition, in: AAAI, 2021.