Pith. sign in

REVIEW 18 cited by

One Million Scenes for Autonomous Driving: ONCE Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11037 v3 pith:BQTDQKI7 submitted 2021-06-21 cs.CV

classification cs.CV
keywords datadrivingautonomousdatasetmilliononcemethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected data and incrementally self-training powerful recognition models have received increasing attention and may become the solutions of next-generation industry-level powerful and robust perception models in autonomous driving. However, the research community generally suffered from data inadequacy of those essential real-world scene data, which hampers the future exploration of fully/semi/self-supervised methods for 3D perception. In this paper, we introduce the ONCE (One millioN sCenEs) dataset for 3D object detection in the autonomous driving scenario. The ONCE dataset consists of 1 million LiDAR scenes and 7 million corresponding camera images. The data is selected from 144 driving hours, which is 20x longer than the largest 3D autonomous driving dataset available (e.g. nuScenes and Waymo), and it is collected across a range of different areas, periods and weather conditions. To facilitate future research on exploiting unlabeled data for 3D detection, we additionally provide a benchmark in which we reproduce and evaluate a variety of self-supervised and semi-supervised methods on the ONCE dataset. We conduct extensive analyses on those methods and provide valuable observations on their performance related to the scale of used data. Data, code, and more information are available at https://once-for-auto-driving.github.io/index.html.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 133 citations worldwide. Full citation record

  1. Orbis 2: A Hierarchical World Model for Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical driving world model — planning in compressed DINO space at 2 Hz and rendering detailed frames at 10 Hz — achieves state-of-the-art long-horizon stability, steering response, and representation quality.

  2. OpenLongTail: Generative Scaling of Long-Tail Driving Data

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.

  3. ToosiCubix: Monocular 3D Cuboid Labeling via Vehicle Part Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A monocular annotation method estimates vehicle position, orientation, and dimensions from user clicks on parts like wheels and badges, with accurate up-to-scale 8DoF results but limited full 9DoF accuracy.

  4. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.

  5. LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LiDARDustX releases 30,000 annotated LiDAR frames, over 80% of them dust-affected, and shows that dust degrades 3D detection average precision by 9.9 to 18.1 points across six detectors.

  6. CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Post-episode multi-agent debriefing lets LLM driving agents learn concise natural-language coordination protocols that avoid collisions and merge traffic, and distillation makes the policy fast enough for near-real-time use.

  7. DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DriveEditor uses depth-aware 3D bounding box projection and single-reference appearance cues to reposition, replace, remove, and insert objects in driving videos with a single diffusion framework.

  8. Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model

    cs.RO 2024-11 conditional novelty 6.0 of 10

    AutoScenario translates multimodal real-world driving data into controllable and diverse corner-case scenarios for autonomous vehicle testing using LLMs and SUMO/CARLA simulations.

  9. An interactive enhanced driving dataset for autonomous driving

    cs.CV 2026-02 conditional novelty 5.0 of 10

    A fused, interaction-labeled dataset of 7.31M driving segments with synthetic BEV videos and VQA pairs for training/evaluating driving VLMs.

  10. Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A single universal adversarial image perturbation can route different input semantics to different attacker-defined outputs in multimodal LLMs, with up to 66% success over five targets.

  11. High-Fidelity Digital Twins for Bridging the Sim2Real Gap in LiDAR-Based ITS Perception

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A digital twin of a real intersection can generate LiDAR training data that matches the target location, and a detector trained on it reported 4.8% higher car AP than a model trained on real data, though with more syn...

  12. Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...

  13. InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A closed-loop LLM-agent framework that auto-generates research ideas and code, reported to improve baseline performance on all 12 tasks it was tested on.

  14. Street Gaussians without 3D Object Tracker

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Replacing 3D object trackers with a 2D foundation model plus LiDAR and a motion-learning correction produces state-of-the-art street-scene reconstructions without ground-truth object poses.

  15. Generating Out-Of-Distribution Scenarios Using Language Models

    cs.LG 2024-11 reject novelty 5.0 of 10

    The paper generates rare driving scenarios via an LLM-built tree, simulates them in CARLA, but its OOD-ness metric is interpreted in a way that contradicts its equation.

  16. Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.

  17. Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop

    cs.CV 2024-11 conditional novelty 4.0 of 10

    Scene Copilot is a training-free pipeline that combines an LLM with retrieval over Infinigen's codebase and human-in-the-loop Blender editing to generate customized 3D scenes and videos from text prompts.

  18. Monocular Lane Detection Based on Deep Learning: A Survey

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A structured review of 2D and 3D monocular lane detection methods, with a new four-axis taxonomy and unified FPS comparisons.

Pith tools