Medical Image Segmentation with Dual-Encoding and Multi-Level Feature Adaptive Fusion-Reference-Cited by-同舟云学术

Medical Image Segmentation with Dual-Encoding and Multi-Level Feature Adaptive Fusion

Published:2024-03-30 Issue:04 Volume:38 Page:
ISSN:0218-0014
Container-title:International Journal of Pattern Recognition and Artificial Intelligence
language:en
Short-container-title:Int. J. Patt. Recogn. Artif. Intell.

Author:

Wu Shulei¹^ORCID,Yang You²^ORCID,Zhang Fanghong²^ORCID

Affiliation:

1. College of Computer and Information Science, Chongqing Normal University, Chongqing 401331, P. R. China

2. National Center for Applied Mathematics in Chongqing, Chongqing Normal University, Chongqing 401331, P. R. China

Abstract

Purpose: Accurate segmentation of medical images is critical for disease diagnosis, surgical planning and prognostic assessment. TransUNet, a hybrid CNN-Transformer-based method, extracts local features using CNN and compensates for the lack of long-range dependencies through a self-attention mechanism. However, the initial focus on extracting local features from specific regions impacts the generation of subsequent global features, thus constraining the model’s capacity to effectively capture a broader range of semantic information. Effective integration of local and global features plays a pivotal role in achieving precise and dense prediction. Therefore, we propose a novel hybrid CNN-Transformer-based method aimed at enhancing medical image segmentation. Approach: In this study, a dual-encoder parallel structure is used to enhance the feature representation of the input image. By introducing a multi-scale adaptive feature fusion module, a fine fusion of local features across perceptual domains is realized in the decoding process. The generalized convolutional block attention module helps to increase cross-channel interactions in layers with more channels, thus enabling the fusion of local features and global representations at different resolutions during the decoding process. Results: The proposed method achieves average DSC scores of 79.98%, 84.83% and 85.78% on the Synapse, ISIC2017 and Pediatric Pyelonephritis datasets, respectively. These scores are 2.5%, 0.56% and 0.42% higher than those of TransUNet. The best performance of 91.66% is observed on the ACDC dataset, representing improvements of 2.46% and 7.24% compared to HiFormer and DAE-Former, respectively. Conclusions: The experimental results show that the proposed model has a significant competitive advantage in terms of ACDC image segmentation performance.

Funder

Science and Technology Research Project of Chongqing Municipal Education Commission

Publisher

World Scientific Pub Co Pte Ltd

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0218001424540041

Reference33 articles.

1. DAE-Former: Dual Attention-Guided Efficient Transformer for Medical Image Segmentation

2. YoloMask: An Enhanced YOLO Model for Detection of Face Mask Wearing Normality, Irregularity and Spoofing

3. CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

4. Skin lesion analysis toward melanoma detection: A challenge at the 2017 International symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (ISIC)

5. Semantic segmentation in medical images through transfused convolution and transformer networks