REVIEW 11 cited by
Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D occupancy prediction, which estimates the detailed occupancy states and semantics of a scene, is an emerging task to overcome these limitations. To support 3D occupancy prediction, we develop a label generation pipeline that produces dense, visibility-aware labels for any given scene. This pipeline comprises three stages: voxel densification, occlusion reasoning, and image-guided voxel refinement. We establish two benchmarks, derived from the Waymo Open Dataset and the nuScenes Dataset, namely Occ3D-Waymo and Occ3D-nuScenes benchmarks. Furthermore, we provide an extensive analysis of the proposed dataset with various baseline models. Lastly, we propose a new model, dubbed Coarse-to-Fine Occupancy (CTF-Occ) network, which demonstrates superior performance on the Occ3D benchmarks. The code, data, and benchmarks are released at https://tsinghua-mars-lab.github.io/Occ3D/.
Forward citations
Cited by 11 Pith papers
-
Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI
Humanoid-OmniOcc delivers a large-scale panoramic stereo occupancy dataset for humanoid robots via Real2Sim2Real, with a model that outperforms monocular baselines in both unseen sim scenes and real settings.
-
VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.
-
SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction
SparseOcc++ decouples geometry completion (via orthogonal SCF regression on sparse anchors) from semantics, improving IoU 2.3 points and running 3.9 imes faster than SparseOcc on nuScenes while 5.9 imes faster than Oc...
-
GaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy Prediction
A hierarchical Gaussian occupancy representation with regression-based seeding predicts high-resolution 3D occupancy at lower latency than prior sparse baselines, validated on nuScenes and a new 0.1m campus dataset.
-
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
Introduces PKL to rank planning-critical occluded agents, creates a VLM-annotated benchmark on nuScenes, and shows fine-tuning on this data improves performance ~30% over random selection with smaller models outperfor...
-
VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
Offline VLM instance audits, grounded to matched object voxels and distilled via taxonomy, attribute, and graph losses, raise OccWorld and GaussianWorld closed-set occupancy mIoU without inference-time VLM cost.
-
Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy
HiPR improves 3D occupancy prediction by reparameterizing image-to-voxel projections using LiDAR-derived height priors to adapt sampling ranges to scene sparsity and height variations.
-
Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy
HiPR improves 3D occupancy prediction by adaptively reparameterizing projection sampling ranges using LiDAR height priors instead of fixed uniform pillars.
-
GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes
GOLD-BEV learns dense BEV semantic maps including dynamic agents from ego-centric sensors by using synchronized aerial imagery for training supervision and pseudo-label generation.
-
InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving
InfiniVerse reconstructs 3D occupancy from one frame, extends scenes autoregressively, converts to video via diffusion, and uses re-projection feedback to achieve SOTA FID 6.4 and FVD 67.97 on Waymo and nuScenes.
-
Principles of Robot Autonomy
A comprehensive textbook framing autonomous robots through a See-Think-Act pipeline, with exercises and notebooks, but no new technical findings.
Discussion (0). Sign in to comment.