An attention mechanism module with spatial perception and channel information interaction-Reference-Cited by-同舟云学术

An attention mechanism module with spatial perception and channel information interaction

Published:2024-05-06 Issue:4 Volume:10 Page:5427-5444
ISSN:2199-4536
Container-title:Complex & Intelligent Systems
language:en
Short-container-title:Complex Intell. Syst.

Author:

Wang Yifan,Wang Wu,Li Yang^ORCID,Jia Yaodong,Xu Yu,Ling Yu,Ma Jiaqi

Abstract

AbstractIn the field of deep learning, the attention mechanism, as a technology that mimics human perception and attention processes, has made remarkable achievements. The current methods combine a channel attention mechanism and a spatial attention mechanism in a parallel or cascaded manner to enhance the model representational competence, but they do not fully consider the interaction between spatial and channel information. This paper proposes a method in which a space embedded channel module and a channel embedded space module are cascaded to enhance the model’s representational competence. First, in the space embedded channel module, to enhance the representational competence of the region of interest in different spatial dimensions, the input tensor is split into horizontal and vertical branches according to spatial dimensions to alleviate the loss of position information when performing 2D pooling. To smoothly process the features and highlight the local features, four branches are obtained through global maximum and average pooling, and the features are aggregated by different pooling methods to obtain two feature tensors with different pooling methods. To enable the output horizontal and vertical feature tensors to focus on different pooling features simultaneously, the two feature tensors are segmented and dimensionally transposed according to spatial dimensions, and the features are later aggregated along the spatial direction. Then, in the channel embedded space module, for the problem of no cross-channel connection between groups in grouped convolution and for which the parameters are large, this paper uses adaptive grouped banded matrices. Based on the banded matrices utilizing the mapping relationship that exists between the number of channels and the size of the convolution kernels, the convolution kernel size is adaptively computed to achieve adaptive cross-channel interaction, enhancing the correlation between the channel dimensions while ensuring that the spatial dimensions remain unchanged. Finally, the output horizontal and vertical weights are used as attention weights. In the experiment, the attention mechanism module proposed in this paper is embedded into the MobileNetV2 and ResNet networks at different depths, and extensive experiments are conducted on the CIFAR-10, CIFAR-100 and STL-10 datasets. The results show that the method in this paper captures and utilizes the features of the input data more effectively than the other methods, significantly improving the classification accuracy. Despite the introduction of an additional computational burden (0.5 M), however, the overall performance of the model still achieves the best results when the computational overhead is comprehensively considered.

Funder

Jilin Provincial Scientific and Technological Development Program

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s40747-024-01445-9.pdf

Reference34 articles.

1. Cristina Z, Eugenio MC, Enrique HV, Iyad AK, Francisco H (2023) Explainable crowd decision making methodology guided by expert natural language opinions based on sentiment analysis with attention-based deep learning and subgroup discovery. Inf Fusion 97(8):101821. https://doi.org/10.1016/j.inffus.2023.101821

2. Zhang S, Wei Z, Xu W, Zhang LL, Wang Y, Zhou X, Liu JY (2023) DSC-MVSNet: attention aware cost volume regularization based on depthwise separable convolution for multi-view stereo. Complex Intell 9:6953–6969. https://doi.org/10.1007/s40747-023-01106-3

3. Lakshmi RK, Rama SA (2023) Novel heuristic-based hybrid ResNeXt with recurrent neural network to handle multi class classification of sentiment analysis. Mach Learn: Sci Technol 4:015033. https://doi.org/10.1088/2632-2153/acc0d5

4. Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: CVPR 7132–7141. https://doi.org/10.1109/CVPR.2018.00745

5. Wang QL, Wu BG, Zhu PF, Li PH, Zuo WM; Hu QH (2020) ECA-Net: efficient channel attention for deep convolutional neural networks. 2020 IEEE/CVF conference on computer vision and pattern recognition (CVPR) 11531-11539. https://doi.org/10.1109/CVPR42600.2020.01155

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Spatial Feature Enhancement and Attention-Guided Bidirectional Sequential Spectral Feature Extraction for Hyperspectral Image Classification;Remote Sensing;2024-08-24

2. DV3-IBi_YOLOv5s: A Lightweight Backbone Network and Multiscale Neck Network Vehicle Detection Algorithm;Sensors;2024-06-11

3. A GM-JMNS-CPHD Filter for Different-Fields-of-View Stochastic Outlier Selection for Nonlinear Motion Tracking;Sensors;2024-05-16