PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest

Deng, Jiajun; Zhang, Sha; Dayoub, Feras; Ouyang, Wanli; Zhang, Yanyong; Reid, Ian

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.09212v1 (cs)

[Submitted on 14 Mar 2024 (this version), latest version 22 Sep 2024 (v2)]

Title:PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest

Authors:Jiajun Deng, Sha Zhang, Feras Dayoub, Wanli Ouyang, Yanyong Zhang, Ian Reid

View PDF HTML (experimental)

Abstract:In this work, we present PoIFusion, a simple yet effective multi-modal 3D object detection framework to fuse the information of RGB images and LiDAR point clouds at the point of interest (abbreviated as PoI). Technically, our PoIFusion follows the paradigm of query-based object detection, formulating object queries as dynamic 3D boxes. The PoIs are adaptively generated from each query box on the fly, serving as the keypoints to represent a 3D object and play the role of basic units in multi-modal fusion. Specifically, we project PoIs into the view of each modality to sample the corresponding feature and integrate the multi-modal features at each PoI through a dynamic fusion block. Furthermore, the features of PoIs derived from the same query box are aggregated together to update the query feature. Our approach prevents information loss caused by view transformation and eliminates the computation-intensive global attention, making the multi-modal 3D object detector more applicable. We conducted extensive experiments on the nuScenes dataset to evaluate our approach. Remarkably, our PoIFusion achieves 74.9\% NDS and 73.4\% mAP, setting a state-of-the-art record on the multi-modal 3D object detection benchmark. Codes will be made available via \url{this https URL}.

Comments:	NIL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.09212 [cs.CV]
	(or arXiv:2403.09212v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.09212

Submission history

From: Jiajun Deng [view email]
[v1] Thu, 14 Mar 2024 09:28:12 UTC (2,690 KB)
[v2] Sun, 22 Sep 2024 06:53:07 UTC (624 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators