REVIEW 35 cited by
Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability to detect objects regardless of image distortions or weather conditions is crucial for real-world applications of deep learning like autonomous driving. We here provide an easy-to-use benchmark to assess how object detection models perform when image quality degrades. The three resulting benchmark datasets, termed Pascal-C, Coco-C and Cityscapes-C, contain a large variety of image corruptions. We show that a range of standard object detection models suffer a severe performance loss on corrupted images (down to 30--60\% of the original performance). However, a simple data augmentation trick---stylizing the training images---leads to a substantial increase in robustness across corruption type, severity and dataset. We envision our comprehensive benchmark to track future progress towards building robust object detection models. Benchmark, code and data are publicly available.
Forward citations
Cited by 35 Pith papers
-
Real-World Perturbation Testing of Autonomous Driving Systems
Model-level and offline robustness metrics for 72 camera/LiDAR perturbations do not reliably predict closed-loop failures on a full-scale autonomous vehicle.
-
EvBS: Event-guided Blur Synthesis for Domain-adaptive Motion Deblurring
EvBS uses event streams to transfer motion patterns between image patches, synthesizing more diverse blur-sharp pairs and improving domain-adaptive deblurring.
-
Suppress and Diversify: Refining Robust Pathways for Corruption Robustness
S&D improves corruption robustness by selecting the most stable internal pathways under a synthetic corruption and diversifying them through symmetric weight tweaks, with no test-time overhead.
-
Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation
In two small VLMs, internal token probability detects errors with AUROC up to 0.99 while verbalized confidence stays near 0.9 and performs near chance, except under severe low light where both fail.
-
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
An IoU-weighted entropy objective and image-conditioned prompt selection adapt YOLO-World and Grounding DINO at test time, improving robustness on style, weather, low-light, and corruption shifts without labels.
-
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
By probing visual, projection, and response representations, the authors find that most VLM visual knowledge loss for recognition and counting occurs in the language decoder, while spatial understanding is lost in the...
-
Embodied Domain Adaptation for Object Detection
EDAOD adapts open-vocabulary object detectors to new indoor scenes via temporal instance clustering and contrastive learning, outperforming source-free baselines on a new benchmark.
-
ADAM-Dehaze: Adaptive Density-Aware Multi-Stage Dehazing for Improved Object Detection in Foggy Conditions
An adaptive, density-aware multi-stage dehazing framework routes foggy images to specialized restoration branches based on a learned fog-density score, improving dehazing quality and downstream object detection.
-
ZeroVO: Visual Odometry with Minimal Assumptions
A two-frame visual odometry model using estimated depth, language priors, and semi-supervised pseudo-label filtering achieves zero-shot metric-scale pose estimation across multiple driving datasets.
-
SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
A new benchmark shows that camera capture settings and lighting systematically change the performance of image classifiers, object detectors, and VQA models, and that common vision datasets are biased toward narrow ex...
-
On the Promise for Assurance of Differentiable Neurosymbolic Reasoning Paradigms
A systematic comparison finds that differentiable neurosymbolic systems offer better assurance mainly in arithmetic-like reasoning tasks, not across the board, and interpretable shortcuts can increase adversarial risk.
-
Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video
Robust-Ego3D benchmarks dense Neural SLAM under 124 synthetic noise settings, and CorrGS uses correspondence-guided pose initialization plus appearance restoration to improve ego-motion and 3D reconstruction on noisy ...
-
Benchmarking Image Perturbations for Testing Automated Driving Assistance Systems
A benchmark of 32 image perturbations on two ADAS shows most corruptions cause failures, and retraining on perturbed data improves robustness to simulated weather, but the retraining effect is confounded by new road data.
-
Chimera: A Block-Based Neural Architecture Search Framework for Event-Based Object Detection
Chimera uses zero-shot NAS proxies and a diversity index to search heterogeneous recurrent backbones for event cameras, reaching PEDRo mAP 64.2 with 4.9M parameters.
-
Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
HGT-Track fuses visible and thermal drone video with a heterogeneous graph transformer and reports the best MOTA and IDF1 on the authors' new VT-Tiny-MOT benchmark.
-
Visual Modality Prompt for Adapting Vision-Language Object Detectors
ModPrompt adapts vision-language object detectors to infrared and depth data with an input-dependent visual prompt and a decoupled text-embedding residual, without updating the frozen detector.
-
Benchmarking the Robustness of Optical Flow Estimation to Corruptions
Introduces KITTI-FC and GoPro-FC, the first corruption robustness benchmarks for optical flow, with 24 corruptions and 10 findings from 29 model variants.
-
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
Shallow layers are most robust to corruptions early in training; freezing them (EPS) or rewinding them (AWR) reduces corruption error on several benchmarks, with caveats.
-
Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
In Qwen2-VL under image degradation, scale raises internal error-detection AUROC from 0.80 to 0.98 while verbalized confidence stays weak, and 4-bit hurts the confidence signal far more than accuracy, so a 7B-4bit mod...
-
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.
-
Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios
A conditional diffusion model generates LoRA adapter parameters for object detectors at test time, improving continual domain adaptation accuracy by small margins over prior methods.
-
Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
A single-step diffusion feature extractor with an object-masked auxiliary branch and consistency loss improves domain-generalized and adaptive detection accuracy and speed.
-
SHARDeg: A Benchmark for Skeletal Human Action Recognition in Degraded Scenarios
A new benchmark degrades NTU-120 skeleton data three ways, shows degradation type strongly affects accuracy, and finds LogSigRNN overtakes DeGCN at 3 FPS once missing frames are interpolated.
-
SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification
A large-scale benchmark of 76 segmentation and 61 detection models shows that robustness to attacks and corruptions does not reliably track clean accuracy, and that transformer backbones generalize better under shift.
-
Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation
Training the teacher a small number of steps ahead of the student and freezing it during distillation improves student generalization by up to 3.4% on image benchmarks.
-
Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation
DUSA adapts classifiers and segmenters at test time by matching their predictions to conditional noise estimates from a pre-trained diffusion model, using a single timestep and active class selection.
-
Evaluating the Adversarial Robustness of Detection Transformers
DETR object detectors are highly vulnerable to standard adversarial attacks, transfer attacks within the DETR family, and a new attack using intermediate losses cuts accuracy with smaller perturbations.
-
Object Style Diffusion for Generalized Object Detection in Urban Scene
GoDiff improves object detection generalization across weather by generating pseudo-target images with a diffusion model and mixing style statistics during training, but the reported gains are modest and depend on kno...
-
PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection
PhysAug adds low-frequency illumination changes and Fourier-based local occlusions to training images, improving single-domain generalized object detection by about 7 mPC on DWD and Cityscapes-C.
-
OCDet: Object Center Detection via Bounding Box-Aware Heatmap Prediction on Edge Devices with NPUs
OCDet predicts object center heatmaps with Generalized Centerness and Balanced Continuous Focal Loss, and reports higher recall and CAS than YOLO11 on edge NPUs.
-
Improved Robustness from Biologically Inspired Sparse Contrast Representations
Grayscale preprocessing improves nighttime semantic segmentation robustness, but the claimed benefits of the contrast filter and of sparsification are not supported by the reported experiments.
-
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Multimodal LLM accuracy can improve on visually degraded images, and a lightweight test-time tuning module that modulates input quality yields small accuracy gains on some benchmarks.
-
Weight Averaging for Out-of-Distribution Generalization and Few-Shot Domain Adaptation
Gradient-similarity-regularized weight averaging and WA+SAM fine-tuning are tested on OOD and few-shot domain adaptation benchmarks, with mixed results that do not support the claimed improvements.
-
VENOM: Text-driven Unrestricted Adversarial Example Generation with Diffusion Models
A text-to-image diffusion attack that interleaves denoising with momentum-based adversarial gradients and an adaptive on/off switch to generate natural-looking unrestricted adversarial examples.
-
WARLearn: Weather-Adaptive Representation Learning
WARLearn adapts a clean-weather YOLO detector to fog and low light by aligning adverse-weather features to clear-weather features with a Barlow Twins loss, reaching 52.6% mAP on RTTS and 55.7% on ExDark.
Discussion (0). Continue with ORCID to comment.