Pith. sign in

REVIEW 17 cited by

OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03272 v1 pith:HRBRMMZW submitted 2024-09-05 cs.CV cs.RO

OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

classification cs.CV cs.RO
keywords modelworldactionautonomousdrivingoccllamaoccupancyvisual
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world and the relations between action and world dynamics. In contrast, human beings possess world model that enables them to simulate the future states based on 3D internal visual representation and plan actions accordingly. To this end, we propose OccLLaMA, an occupancy-language-action generative world model, which uses semantic occupancy as a general visual representation and unifies vision-language-action(VLA) modalities through an autoregressive model. Specifically, we introduce a novel VQVAE-like scene tokenizer to efficiently discretize and reconstruct semantic occupancy scenes, considering its sparsity and classes imbalance. Then, we build a unified multi-modal vocabulary for vision, language and action. Furthermore, we enhance LLM, specifically LLaMA, to perform the next token/scene prediction on the unified vocabulary to complete multiple tasks in autonomous driving. Extensive experiments demonstrate that OccLLaMA achieves competitive performance across multiple tasks, including 4D occupancy forecasting, motion planning, and visual question answering, showcasing its potential as a foundation model in autonomous driving.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

    cs.RO 2026-05 unverdicted novelty 7.0

    TPS-Drive uses an agent-centric tokenizer supervised by a frozen 3D detection head to purify VLM spatial representations, enabling better scene forecasting and lower collision rates on nuScenes and NAVSIM benchmarks.

  2. GEM: Gaussian Evolution Model for Occupancy Forecasting and Motion Planning

    cs.CV 2026-05 unverdicted novelty 7.0

    GEM represents driving scenes as explicit continuous 4D Gaussian primitives with learned dynamics to enable direct querying at arbitrary timestamps for semantic occupancy forecasting and motion planning.

  3. What if? Emulative Simulation with World Models for Situated Reasoning

    cs.CV 2026-03 conditional novelty 6.5

    WanderDream supplies 15.8K panoramic mental-exploration videos and 158K QA pairs showing that world-model imagination measurably improves situated spatial reasoning without active exploration.

  4. Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0

    FocusGS localizes a 3D geometric ambiguity manifold from depth discontinuities and instantiates continuous Gaussian queries only there, yielding SOTA sparse-view driving reconstruction with far fewer Gaussians.

  5. AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

    cs.RO 2026-05 unverdicted novelty 6.0

    AnyScene is an occupancy-centric framework using a Spatial-Temporal Occupancy Diffusion Transformer and Geometry-Grounded View Expansion to generate controllable driving scenes and videos from BEV layouts.

  6. DriveFuture: Future-Aware Latent World Models for Autonomous Driving

    cs.CV 2026-05 unverdicted novelty 6.0

    DriveFuture achieves SOTA results on NAVSIM by conditioning latent world model states on future predictions to directly inform trajectory planning.

  7. HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    HERMES++ unifies 3D scene understanding and future geometry prediction in driving scenes via BEV representations, LLM-enhanced queries, a temporal link, and joint geometric optimization.

  8. Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

    cs.CV 2026-03 unverdicted novelty 6.0

    Chat-Scene++ improves 3D scene understanding in multimodal LLMs by representing scenes as context-rich object sequences with identifier tokens and grounded chain-of-thought reasoning, reaching state-of-the-art on five...

  9. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes

    cs.CV 2026-02 unverdicted novelty 6.0

    A 3D Language-Embedded Gaussians framework with opacity-aware Poisson volumetric aggregation and progressive temperature decay achieves 59.50 IoU and 21.05 mIoU on Occ-ScanNet for open-vocabulary indoor occupancy.

  10. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  11. GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

    cs.CV 2026-07 conditional novelty 5.0

    A unified framework converts surface geometry priors into sparse Gaussian occupancy predictions and extends it to multi-view and temporal inputs.

  12. CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

    cs.CV 2026-06 unverdicted novelty 5.0

    CascadeOcc uses cascaded VQ representations in an autoregressive framework with a TimeMixer for multi-scale spatial and temporal modeling, achieving top results among vision-centric methods on 4D occupancy and plannin...

  13. Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning

    cs.RO 2026-06 unverdicted novelty 5.0

    Discrete-WAM unifies world modeling and policy learning for autonomous driving by representing observations, states, decisions, and actions as tokens in one space and using hierarchical token editing for planning.

  14. Artificial Intelligence for Modeling and Simulation of Mixed Automated and Human Traffic

    cs.AI 2026-04 unverdicted novelty 5.0

    This survey synthesizes AI techniques for mixed autonomy traffic simulation and introduces a taxonomy spanning agent-level behavior models, environment-level methods, and cognitive/physics-informed approaches.

  15. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  16. SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

    cs.CV 2025-11 unverdicted novelty 5.0

    A sparse transformer predicts multi-frame 3D occupancy from images without BEV or VAE tokenization and reports SOTA results on nuScenes for 1-3s forecasting under arbitrary trajectories.

  17. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...