Pith. sign in

REVIEW 14 cited by

A2D2: Audi Autonomous Driving Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.06320 v2 pith:MSVN4VF4 submitted 2020-04-14 cs.CV cs.LGeess.IV

A2D2: Audi Autonomous Driving Dataset

classification cs.CV cs.LGeess.IV
keywords dataframesa2d2autonomousdatasetdrivingsegmentationannotations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Research in machine learning, mobile robotics, and autonomous driving is accelerated by the availability of high quality annotated data. To this end, we release the Audi Autonomous Driving Dataset (A2D2). Our dataset consists of simultaneously recorded images and 3D point clouds, together with 3D bounding boxes, semantic segmentation, instance segmentation, and data extracted from the automotive bus. Our sensor suite consists of six cameras and five LiDAR units, providing full 360 degree coverage. The recorded data is time synchronized and mutually registered. Annotations are for non-sequential frames: 41,277 frames with semantic segmentation image and point cloud labels, of which 12,497 frames also have 3D bounding box annotations for objects within the field of view of the front camera. In addition, we provide 392,556 sequential frames of unannotated sensor data for recordings in three cities in the south of Germany. These sequences contain several loops. Faces and vehicle number plates are blurred due to GDPR legislation and to preserve anonymity. A2D2 is made available under the CC BY-ND 4.0 license, permitting commercial use subject to the terms of the license. Data and further information are available at https://a2d2-dataset.github.io/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset

    cs.CV 2026-07 unverdicted novelty 7.0

    UnderOneFacade is a large-scale cross-continent 3D facade point cloud benchmark with harmonized labels that reveals existing segmentation models achieve at most 33 IoU on fine-grained architectural elements and degrad...

  2. DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

    cs.CV 2026-05 unverdicted novelty 7.0

    A reinforcement-learned vision-language agent adaptively selects and fuses monocular depth experts per sample for better performance across camera geometries.

  3. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

    cs.CV 2026-05 conditional novelty 7.0

    CARD is a new multi-modal driving dataset delivering ~500K dense depth pixels per frame from challenging road topographies using stereo cameras and fused LiDARs over 110 km.

  4. UniDAC: Universal Metric Depth Estimation for Any Camera

    cs.CV 2026-03 unverdicted novelty 7.0

    UniDAC achieves universal metric depth estimation across camera types by decoupling relative depth prediction from spatially varying scale estimation using a depth-guided module and distortion-aware positional embedding.

  5. Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

    cs.CV 2023-01 accept novelty 7.0

    Argoverse 2 introduces three new datasets with annotated sensor data, massive lidar collections, and challenging motion forecasting scenarios for autonomous driving research.

  6. ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes

    cs.CV 2026-04 unverdicted novelty 6.0

    ParkingScenes is a new multimodal dataset of 704 structured reverse and parallel parking episodes generated in CARLA with Hybrid A* and MPC trajectories, showing better model performance than unstructured simulation data.

  7. MoonSplat: Monocular Online Gaussian Splatting with Sim(3) Global Optimization

    cs.CV 2026-06 unverdicted novelty 5.0

    MoonSplat adds global Sim(3) loop closure and color residual learning to voxelized online 3D Gaussian Splatting for improved monocular camera tracking and rendering quality.

  8. High-Fidelity Digital Twins for Bridging the Sim2Real Gap in LiDAR-Based ITS Perception

    cs.CV 2025-09 conditional novelty 5.0

    A digital twin of a real intersection can generate LiDAR training data that matches the target location, and a detector trained on it reported 4.8% higher car AP than a model trained on real data, though with more syn...

  9. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

    cs.CV 2025-02 conditional novelty 5.0

    UniDepthV2 predicts metric 3D points directly from single images using a self-promptable camera module, pseudo-spherical representation, and new losses for improved cross-domain generalization.

  10. PSI: A Benchmark for Human Interpretation and Response in Traffic Interactions

    cs.CV 2021-12 unverdicted novelty 5.0

    PSI is a benchmark dataset for pedestrian intention prediction, driver decision modeling, and reasoning generation in traffic interactions, enriched with human textual explanations.

  11. Eyes All Around: Design and Analysis of 360-Degree LiDAR Perception Using Equivariant Feature Learning in Unstructured Traffic

    cs.CV 2026-05 unverdicted novelty 4.0

    A 360-degree LiDAR detection system using equivariant features achieves stable performance on vehicles in unstructured urban traffic but struggles with smaller road users.

  12. Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles

    cs.CV 2026-05 unverdicted novelty 4.0

    Adaptive confidence threshold selection improves F1 scores in explainable multi-task classification for autonomous driving and is supported by a new 958-image dataset.

  13. Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

    cs.CV 2026-02 unverdicted novelty 4.0

    L-LIO integrates audio with visual data to enhance driver safety assessment and intelligent vehicle decision-making via multimodal sensor fusion.

  14. All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles

    cs.CV 2025-10 unverdicted novelty 4.0

    A survey synthesizing sensor fusion strategies, AV datasets, and emerging LLM/VLM-powered object detection pipelines for autonomous vehicles.