Application of Split Coordinate Channel Attention Embedding U2Net in Salient Object Detection-Reference-Cited by-同舟云学术

Application of Split Coordinate Channel Attention Embedding U2Net in Salient Object Detection

Published:2024-03-06 Issue:3 Volume:17 Page:109
ISSN:1999-4893
Container-title:Algorithms
language:en
Short-container-title:Algorithms

Author:

Wu Yuhuan¹,Wu Yonghong¹

Affiliation:

1. School of Science, Wuhan University of Technology, Wuhan 430074, China

Abstract

Salient object detection (SOD) aims to identify the most visually striking objects in a scene, simulating the function of the biological visual attention system. The attention mechanism in deep learning is commonly used as an enhancement strategy which enables the neural network to concentrate on the relevant parts when processing input data, effectively improving the model’s learning and prediction abilities. Existing saliency object detection methods based on RGB deep learning typically treat all regions equally by using the extracted features, overlooking the fact that different regions have varying contributions to the final predictions. Based on the U2Net algorithm, this paper incorporates the split coordinate channel attention (SCCA) mechanism into the feature extraction stage. SCCA conducts spatial transformation in width and height dimensions to efficiently extract the location information of the target to be detected. While pixel-level semantic segmentation based on annotation has been successful, it assigns the same weight to each pixel which leads to poor performance in detecting the boundary of objects. In this paper, the Canny edge detection loss is incorporated into the loss calculation stage to improve the model’s ability to detect object edges. Based on the DUTS and HKU-IS datasets, experiments confirm that the proposed strategies effectively enhance the model’s detection performance, resulting in a 0.8% and 0.7% increase in the F1-score of U2Net. This paper also compares the traditional attention modules with the newly proposed attention, and the SCCA attention module achieves a top-three performance in prediction time, mean absolute error (MAE), F1-score, and model size on both experimental datasets.

Publisher

MDPI AG

Link

https://www.mdpi.com/1999-4893/17/3/109/pdf

Reference33 articles.

1. Machine learning and deep learning applications—A vision;Sharma;Glob. Transit. Proc.,2021

2. Shinde, P.P., and Shah, S. (2018, January 16–18). A review of machine learning and deep learning applications. Proceedings of the 2018 Fourth International Conference on Computing Communication Control and Automation (ICCUBEA), Pune, India.

3. Deep learning in computer vision: A critical review of emerging techniques and application scenarios;Chai;Mach. Learn. Appl.,2021

4. Computer vision techniques in construction: A critical review;Xu;Arch. Comput. Methods Eng.,2021

5. Deep learning for computer vision: A brief review;Voulodimos;Comput. Intell. Neurosci.,2018