Pith. sign in

REVIEW 12 cited by

RenderWorld: World Model with Self-Supervised 3D Label

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11356 v2 pith:ZU4EDCKY submitted 2024-09-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords renderworldautonomousdrivingmodelworldam-vaecomparedend-to-end
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

End-to-end autonomous driving with vision-only is not only more cost-effective compared to LiDAR-vision fusion but also more reliable than traditional methods. To achieve a economical and robust purely visual autonomous driving system, we propose RenderWorld, a vision-only end-to-end autonomous driving framework, which generates 3D occupancy labels using a self-supervised gaussian-based Img2Occ Module, then encodes the labels by AM-VAE, and uses world model for forecasting and planning. RenderWorld employs Gaussian Splatting to represent 3D scenes and render 2D images greatly improves segmentation accuracy and reduces GPU memory consumption compared with NeRF-based methods. By applying AM-VAE to encode air and non-air separately, RenderWorld achieves more fine-grained scene element representation, leading to state-of-the-art performance in both 4D occupancy forecasting and motion planning from autoregressive world model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.

  2. COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    COME adds a scene-centric forecasting branch as a ControlNet-style condition to a diffusion occupancy world model, improving static-scene consistency and beating prior methods on Occ3D-nuScenes while hiding a stronger...

  3. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  4. Occupancy World Model for Robots

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RoboOccWorld predicts future 3D occupancy for indoor robots by conditioning an autoregressive transformer on the next camera pose, outperforming OccWorld on a restructured ScanNet benchmark.

  5. EventVAD: Training-Free Event-Aware Video Anomaly Detection

    cs.CV 2025-04 conditional novelty 6.0 of 10

    EventVAD improves training-free video anomaly detection by detecting event boundaries from CLIP and RAFT features and feeding coherent event segments to a 7B multimodal LLM, achieving 82.03 AUC on UCF-Crime and 64.04 ...

  6. HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    HERMES unifies BEV-based scene understanding and future point cloud generation in a single LLM-driven self-driving world model, with reported gains on nuScenes and OmniDrive-nuScenes.

  7. An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An end-to-end, non-autoregressive 3D occupancy world model warps dynamic voxels via predicted flow, moves static voxels by pose, and uses image-based rendering supervision, achieving state-of-the-art results on three ...

  8. GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GaussianFormer-2 predicts 3D semantic occupancy from cameras by multiplying Gaussian occupancy probabilities and using a Gaussian mixture for semantics, beating prior methods with far fewer Gaussians.

  9. A Definition and Roadmap for World Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

  10. ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

    cs.CV 2025-08 reject novelty 5.0 of 10

    ICM-Fusion uses a conditional VAE plus task-vector guidance to fuse multiple LoRA adapters into one model, reporting marginal average gains on vision and language benchmarks and larger gains in a few-shot long-tail setup.

  11. QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    QuadricFormer represents 3D scenes as a probabilistic mixture of superquadrics, improving accuracy and efficiency over Gaussian-based occupancy prediction on nuScenes.

  12. Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

    cs.RO 2025-09 conditional novelty 4.0 of 10

    Foundation-model perception for autonomous driving is surveyed through four capability lenses: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding.

Pith tools