OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Zheng, Wenzhao; Chen, Weiliang; Huang, Yuanhui; Zhang, Borui; Duan, Yueqi; Lu, Jiwen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2311.16038 (cs)

[Submitted on 27 Nov 2023]

Title:OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Authors:Wenzhao Zheng, Weiliang Chen, Yuanhui Huang, Borui Zhang, Yueqi Duan, Jiwen Lu

View PDF

Abstract:Understanding how the 3D scene evolves is vital for making decisions in autonomous driving. Most existing methods achieve this by predicting the movements of object boxes, which cannot capture more fine-grained scene information. In this paper, we explore a new framework of learning a world model, OccWorld, in the 3D Occupancy space to simultaneously predict the movement of the ego car and the evolution of the surrounding scenes. We propose to learn a world model based on 3D occupancy rather than 3D bounding boxes and segmentation maps for three reasons: 1) expressiveness. 3D occupancy can describe the more fine-grained 3D structure of the scene; 2) efficiency. 3D occupancy is more economical to obtain (e.g., from sparse LiDAR points). 3) versatility. 3D occupancy can adapt to both vision and LiDAR. To facilitate the modeling of the world evolution, we learn a reconstruction-based scene tokenizer on the 3D occupancy to obtain discrete scene tokens to describe the surrounding scenes. We then adopt a GPT-like spatial-temporal generative transformer to generate subsequent scene and ego tokens to decode the future occupancy and ego trajectory. Extensive experiments on the widely used nuScenes benchmark demonstrate the ability of OccWorld to effectively model the evolution of the driving scenes. OccWorld also produces competitive planning results without using instance and map supervision. Code: this https URL.

Comments:	Code is available at: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2311.16038 [cs.CV]
	(or arXiv:2311.16038v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2311.16038

Submission history

From: Wenzhao Zheng [view email]
[v1] Mon, 27 Nov 2023 17:59:41 UTC (8,857 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators