Multi-Modal Enhancement Transformer Network for Skeleton-Based Human Interaction Recognition-Reference-Cited by-同舟云学术

Multi-Modal Enhancement Transformer Network for Skeleton-Based Human Interaction Recognition

Published:2024-02-20 Issue:3 Volume:9 Page:123
ISSN:2313-7673
Container-title:Biomimetics
language:en
Short-container-title:Biomimetics

Author:

Hu Qianshuo¹^ORCID,Liu Haijun¹^ORCID

Affiliation:

1. School of Microelectronics and Communication Engineering, Chongqing University, Chongqing 400044, China

Abstract

Skeleton-based human interaction recognition is a challenging task in the field of vision and image processing. Graph Convolutional Networks (GCNs) achieved remarkable performance by modeling the human skeleton as a topology. However, existing GCN-based methods have two problems: (1) Existing frameworks cannot effectively take advantage of the complementary features of different skeletal modalities. There is no information transfer channel between various specific modalities. (2) Limited by the structure of the skeleton topology, it is hard to capture and learn the information about two-person interactions. To solve these problems, inspired by the human visual neural network, we propose a multi-modal enhancement transformer (ME-Former) network for skeleton-based human interaction recognition. ME-Former includes a multi-modal enhancement module (ME) and a context progressive fusion block (CPF). More specifically, each ME module consists of a multi-head cross-modal attention block (MH-CA) and a two-person hypergraph self-attention block (TH-SA), which are responsible for enhancing the skeleton features of a specific modality from other skeletal modalities and modeling spatial dependencies between joints using the specific modality, respectively. In addition, we propose a two-person skeleton topology and a two-person hypergraph representation. The TH-SA block can embed their structural information into the self-attention to better learn two-person interaction. The CPF block is capable of progressively transforming the features of different skeletal modalities from low-level features to higher-order global contexts, making the enhancement process more efficient. Extensive experiments on benchmark NTU-RGB+D 60 and NTU-RGB+D 120 datasets consistently verify the effectiveness of our proposed ME-Former by outperforming state-of-the-art methods.

Publisher

MDPI AG

Link

https://www.mdpi.com/2313-7673/9/3/123/pdf

Reference59 articles.

1. Zhao, S., Zhao, G., He, Y., Diao, Z., He, Z., Cui, Y., Jiang, L., Shen, Y., and Cheng, C. (2024). Biomimetic Adaptive Pure Pursuit Control for Robot Path Tracking Inspired by Natural Motion Constraints. Biomimetics, 9.

2. Kwon, J.Y., and Ju, D.Y. (2023). Living Lab-Based Service Interaction Design for a Companion Robot for Seniors in South Korea. Biomimetics, 8.

3. Song, F., and Li, P. (2023). YOLOv5-MS: Real-time multi-surveillance pedestrian target detection model for smart cities. Biomimetics, 8.

4. Liu, M., Meng, F., and Liang, Y. (2022). Generalized Pose Decoupled Network for Unsupervised 3d Skeleton Sequence-based Action Representation Learning. Cyborg Bionic Syst., 2022.

5. Facial Prior Guided Micro-Expression Generation;Zhang;IEEE Trans. Image Process.,2024