BAFusion: Bidirectional Attention Fusion for 3D Object Detection Based on LiDAR and Camera-Reference-Cited by-同舟云学术

BAFusion: Bidirectional Attention Fusion for 3D Object Detection Based on LiDAR and Camera

Published:2024-07-20 Issue:14 Volume:24 Page:4718
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Liu Min¹^ORCID,Jia Yuanjun²^ORCID,Lyu Youhao¹^ORCID,Dong Qi²,Yang Yanyu²

Affiliation:

1. Institute of Advanced Technology, University of Science and Technology of China, Hefei 230088, China

2. China Academy of Electronics and Information Technology, Beijing 100041, China

Abstract

3D object detection is a challenging and promising task for autonomous driving and robotics, benefiting significantly from multi-sensor fusion, such as LiDAR and cameras. Conventional methods for sensor fusion rely on a projection matrix to align the features from LiDAR and cameras. However, these methods often suffer from inadequate flexibility and robustness, leading to lower alignment accuracy under complex environmental conditions. Addressing these challenges, in this paper, we propose a novel Bidirectional Attention Fusion module, named BAFusion, which effectively fuses the information from LiDAR and cameras using cross-attention. Unlike the conventional methods, our BAFusion module can adaptively learn the cross-modal attention weights, making the approach more flexible and robust. Moreover, drawing inspiration from advanced attention optimization techniques in 2D vision, we developed the Cross Focused Linear Attention Fusion Layer (CFLAF Layer) and integrated it into our BAFusion pipeline. This layer optimizes the computational complexity of attention mechanisms and facilitates advanced interactions between image and point cloud data, showcasing a novel approach to addressing the challenges of cross-modal attention calculations. We evaluated our method on the KITTI dataset using various baseline networks, such as PointPillars, SECOND, and Part-A2, and demonstrated consistent improvements in 3D object detection performance over these baselines, especially for smaller objects like cyclists and pedestrians. Our approach achieves competitive results on the KITTI benchmark.

Funder

China Academy of Electronics and Information Technology

Publisher

MDPI AG

Link

https://www.mdpi.com/1424-8220/24/14/4718/pdf

Reference57 articles.

1. 3D object detection for autonomous driving: A survey;Qian;Pattern Recognit.,2022

2. Wang, L., Li, R., Shi, H., Sun, J., Zhao, L., Seah, H.S., Quah, C.K., and Tandianus, B. (2019). Multi-channel convolutional neural network based 3D object detection for indoor robot environmental perception. Sensors, 19.

3. Huang, K., Shi, B., Li, X., Li, X., Huang, S., and Li, Y. (2022). Multi-modal sensor fusion for auto driving perception: A survey. arXiv.

4. Multi-modal 3D object detection in autonomous driving: A survey;Wang;Int. J. Comput. Vis.,2023

5. Deep learning for image and point cloud fusion in autonomous driving: A review;Cui;IEEE Trans. Intell. Transp. Syst.,2021

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Object Detection and Information Perception by Fusing YOLO-SCG and Point Cloud Clustering;Sensors;2024-08-19