Pith. sign in

REVIEW 20 cited by

Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.07484 v2 pith:BTUQRQWS submitted 2019-07-17 cs.CV cs.LGstat.ML

Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming

classification cs.CV cs.LGstat.ML
keywords benchmarkdetectionobjectimagemodelsautonomousdatadriving
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The ability to detect objects regardless of image distortions or weather conditions is crucial for real-world applications of deep learning like autonomous driving. We here provide an easy-to-use benchmark to assess how object detection models perform when image quality degrades. The three resulting benchmark datasets, termed Pascal-C, Coco-C and Cityscapes-C, contain a large variety of image corruptions. We show that a range of standard object detection models suffer a severe performance loss on corrupted images (down to 30--60\% of the original performance). However, a simple data augmentation trick---stylizing the training images---leads to a substantial increase in robustness across corruption type, severity and dataset. We envision our comprehensive benchmark to track future progress towards building robust object detection models. Benchmark, code and data are publicly available.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

    cs.CV 2025-06 accept novelty 8.0

    LAION-C supplies six novel corruptions that stay OOD for web-scale training sets and demonstrates that leading models now rival or exceed human robustness on them.

  2. Real-World Perturbation Testing of Autonomous Driving Systems

    cs.SE 2026-07 conditional novelty 7.0

    Model-level and offline robustness metrics for 72 camera/LiDAR perturbations do not reliably predict closed-loop failures on a full-scale autonomous vehicle.

  3. COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions

    cs.CV 2026-05 unverdicted novelty 7.0

    COD10K-C benchmark shows performance drops in camouflaged object detection under corruptions, with RobustCODLite retaining 92.3% of clean Dice score versus 84-88% for SINet-v2, ZoomNet, and PFNet.

  4. Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

    cs.RO 2026-05 unverdicted novelty 7.0

    Bench2Drive-Robust is a new closed-loop benchmark that evaluates end-to-end autonomous driving models under deployment perturbations from camera failures, ego-state errors, and compute delays, showing substantial perf...

  5. Benchmarking Sensor-Fault Robustness in Forecasting

    cs.LG 2026-05 conditional novelty 7.0

    SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.

  6. Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

    cs.CV 2023-12 conditional novelty 7.0

    Q-Align trains LMMs on discrete text-defined levels for visual scoring, achieving SOTA on IQA, IAA, and VQA while unifying the tasks in OneAlign.

  7. Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

    cs.CV 2026-07 conditional novelty 6.0

    In two small VLMs, internal token probability detects errors with AUROC up to 0.99 while verbalized confidence stays near 0.9 and performs near chance, except under severe low light where both fail.

  8. RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

    cs.CV 2026-05 unverdicted novelty 6.0

    RoboStressBench decomposes visual stress into four physically grounded dimensions to benchmark VLM robustness in embodied scenes and proposes a stress-aware solver.

  9. Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    SemProbe is an interactive tool for semantic robustness probing of object detectors that applies controlled diffusion inpainting to user-masked deployment images and automatically compares detection performance before...

  10. Safety-Critical Camera Reliability Monitoring for ADAS via Degradation-Aware Uncertainty Pattern Analysis

    cs.CV 2026-05 unverdicted novelty 6.0

    A multi-task network predicts degradation patterns and a multiplicative Global Sensor Health Index from RGB images to provide early warnings of camera failure in autonomous driving before downstream detection degrades.

  11. Reward-Guided Semantic Evolution for Test-time Adaptive Object Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    RGSE adapts text embeddings at test time via evolutionary search, using cosine similarity rewards from high-confidence visual proposals to improve open-vocabulary object detection under distribution shifts.

  12. Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts

    cs.SE 2026-04 unverdicted novelty 6.0

    A broad empirical benchmark shows how 15 existing test selection metrics perform for fault detection, performance estimation, and retraining under corrupted, adversarial, temporal, natural, and label shifts across ima...

  13. Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback

    cs.CV 2025-12 unverdicted novelty 6.0

    An RL-trained lightweight agent uses MLLM perceptual rewards to perform efficient label-free image restoration, matching SOTA on full-reference metrics and surpassing prior work on no-reference metrics.

  14. APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds

    cs.CV 2025-05 unverdicted novelty 6.0

    APCoTTA introduces a continual test-time adaptation method for ALS point cloud semantic segmentation using gradient-driven layer selection, entropy-based consistency loss, and random parameter interpolation, with new ...

  15. TimberVision: A Multi-Task Dataset and Framework for Log-Component Segmentation and Tracking in Autonomous Forestry Operations

    cs.CV 2025-01 unverdicted novelty 6.0

    Introduces TimberVision dataset and multi-task framework for log-component segmentation, detection, and tracking in forestry operations using RGB images.

  16. Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

    cs.CV 2026-07 conditional novelty 5.0

    In Qwen2-VL under image degradation, scale raises internal error-detection AUROC from 0.80 to 0.98 while verbalized confidence stays weak, and 4-bit hurts the confidence signal far more than accuracy, so a 7B-4bit mod...

  17. Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs

    cs.RO 2026-05 unverdicted novelty 5.0

    Sensor perturbations in driving VLAs cause Chain-of-Causation reasoning changes that correlate strongly with 5.3x higher trajectory deviation, while enabling such reasoning improves accuracy by 11.8%.

  18. StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

    cs.CV 2026-05 unverdicted novelty 5.0

    StableVLA adds an Information Bottleneck Adapter to VLA models that improves robustness to visual corruptions by 30% on average with under 10M extra parameters and no extra data, even when using a much smaller backbone.

  19. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

  20. Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs

    cs.RO 2026-05 unverdicted novelty 4.0

    Changes in Chain-of-Causation explanations under sensor perturbations correlate with 5.3× higher trajectory deviation in a driving VLA, and enabling such explanations yields 11.8% better accuracy.