Progressive Multi-Modal Fusion for Robust 3D Object Detection

Mohan, Rohit; Cattaneo, Daniele; Drews, Florian; Valada, Abhinav

Computer Science > Computer Vision and Pattern Recognition

arXiv:2410.07475 (cs)

[Submitted on 9 Oct 2024]

Title:Progressive Multi-Modal Fusion for Robust 3D Object Detection

Authors:Rohit Mohan, Daniele Cattaneo, Florian Drews, Abhinav Valada

View PDF HTML (experimental)

Abstract:Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspective View (PV), thus sacrificing complementary information such as height or geometric proportions. To address this limitation, we propose ProFusion3D, a progressive fusion framework that combines features in both BEV and PV at both intermediate and object query levels. Our architecture hierarchically fuses local and global features, enhancing the robustness of 3D object detection. Additionally, we introduce a self-supervised mask modeling pre-training strategy to improve multi-modal representation learning and data efficiency through three novel objectives. Extensive experiments on nuScenes and Argoverse2 datasets conclusively demonstrate the efficacy of ProFusion3D. Moreover, ProFusion3D is robust to sensor failure, demonstrating strong performance when only one modality is available.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2410.07475 [cs.CV]
	(or arXiv:2410.07475v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2410.07475
Journal reference:	8th Annual Conference on Robot Learning, 2024

Submission history

From: Rohit Mohan [view email]
[v1] Wed, 9 Oct 2024 22:57:47 UTC (46,340 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Progressive Multi-Modal Fusion for Robust 3D Object Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Progressive Multi-Modal Fusion for Robust 3D Object Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators