A YOLOX Object Detection Algorithm Based on Bidirectional Cross-scale Path Aggregation-Reference-Cited by-同舟云学术

A YOLOX Object Detection Algorithm Based on Bidirectional Cross-scale Path Aggregation

Published:2024-02-15 Issue:1 Volume:56 Page:
ISSN:1573-773X
Container-title:Neural Processing Letters
language:en
Short-container-title:Neural Process Lett

Author:

Liu Qunpo,Zhang Jingwen,Zhao Yi,Bu Xuhui,Hanajima Naohiko

Abstract

AbstractTo solve the problem of insufficient feature fusion between the deep and shallow feature layers of the original YOLOX algorithm, which resulting in a loss of object semantic information, this paper proposes a YOLOX object detection algorithm based on attention and bidirectional cross-scale path aggregation. First, an efficient channel attention module is embedded in the YOLOX backbone network to reinforce the key features in the object region by distinguishing between the importance of the different channels in the feature layer, thus enhancing the detection accuracy of the network. Second, a bidirectional cross-scale path aggregation network is designed to change the information fusion circulation path while increasing the cross-scale connections. Weighted feature fusion is used to learn the importance of the different path input features for differentiated fusion, thereby improving the feature information fusion capability between the deep and shallow layers. Finally, the SIOU loss function is introduced to improve the detection performance of the network. The experimental results show that on the PASCAL VOC2007 and MS COCO2017 datasets, the algorithm in this paper improves mAP by 2.32% and 1.53% compared with the original YOLOX algorithm, and has comprehensive performance advantages compared with other algorithms. The mAP reaches 99.44% on the self-built iron ore metal foreign matter dataset, with a recognition speed of 56.90 frames/s.

Funder

National Natural Science Foundation of China

Innovative Scientists and Technicians Team of Henan Provincial High Education

Science and Technology Project of Henan Province

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s11063-024-11536-w.pdf

Reference37 articles.

1. Zhang H (2020) Research on tunnel microseismic signal processing and intelligent rock burst early warning based on deep learning. Dissertation, Chengdu University of Technology

2. Sun X L (2022) Research on generative target tracking method under deep learning framework. Dissertation, University of Chinese Academy of Sciences (Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences)

3. Girshick R, Donahue J, Darrell T et al (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 580–587

4. Girshick R (2015) Fast R-CNN. In: Proceedings of the IEEE international conference on computer vision, pp 1440–1448

5. Ren S, He K, Girshick R et al (2015) Faster R-CNN: towards real-time object detection with region proposal networks. Adv Neural Inf Process Syst 28:589–598

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. TSMDA: intelligent fault diagnosis of rolling bearing with two stage multi-source domain adaptation;Measurement Science and Technology;2024-08-08

2. Study on a Landslide Segmentation Algorithm Based on Improved High-Resolution Networks;Applied Sciences;2024-07-24