LAION-C supplies six novel corruptions that stay OOD for web-scale training sets and demonstrates that leading models now rival or exceed human robustness on them.
hub
Benchmarking ro- bustness in object detection: Autonomous driving when win- ter is coming
15 Pith papers cite this work, alongside 136 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 3polarities
background 3representative citing papers
COD10K-C benchmark shows performance drops in camouflaged object detection under corruptions, with RobustCODLite retaining 92.3% of clean Dice score versus 84-88% for SINet-v2, ZoomNet, and PFNet.
Bench2Drive-Robust is a new closed-loop benchmark that evaluates end-to-end autonomous driving models under deployment perturbations from camera failures, ego-state errors, and compute delays, showing substantial performance degradation beyond image-level tests.
SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.
Q-Align trains LMMs on discrete text-defined levels for visual scoring, achieving SOTA on IQA, IAA, and VQA while unifying the tasks in OneAlign.
RoboStressBench decomposes visual stress into four physically grounded dimensions to benchmark VLM robustness in embodied scenes and proposes a stress-aware solver.
SemProbe is an interactive tool for semantic robustness probing of object detectors that applies controlled diffusion inpainting to user-masked deployment images and automatically compares detection performance before and after.
A multi-task network predicts degradation patterns and a multiplicative Global Sensor Health Index from RGB images to provide early warnings of camera failure in autonomous driving before downstream detection degrades.
RGSE adapts text embeddings at test time via evolutionary search, using cosine similarity rewards from high-confidence visual proposals to improve open-vocabulary object detection under distribution shifts.
A broad empirical benchmark shows how 15 existing test selection metrics perform for fault detection, performance estimation, and retraining under corrupted, adversarial, temporal, natural, and label shifts across image, text, and Android data.
An RL-trained lightweight agent uses MLLM perceptual rewards to perform efficient label-free image restoration, matching SOTA on full-reference metrics and surpassing prior work on no-reference metrics.
APCoTTA introduces a continual test-time adaptation method for ALS point cloud semantic segmentation using gradient-driven layer selection, entropy-based consistency loss, and random parameter interpolation, with new benchmarks showing mIoU gains of 9-14%.
Introduces TimberVision dataset and multi-task framework for log-component segmentation, detection, and tracking in forestry operations using RGB images.
StableVLA adds an Information Bottleneck Adapter to VLA models that improves robustness to visual corruptions by 30% on average with under 10M extra parameters and no extra data, even when using a much smaller backbone.
Changes in Chain-of-Causation explanations under sensor perturbations correlate with 5.3× higher trajectory deviation in a driving VLA, and enabling such explanations yields 11.8% better accuracy.
citing papers explorer
-
LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
LAION-C supplies six novel corruptions that stay OOD for web-scale training sets and demonstrates that leading models now rival or exceed human robustness on them.
-
COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions
COD10K-C benchmark shows performance drops in camouflaged object detection under corruptions, with RobustCODLite retaining 92.3% of clean Dice score versus 84-88% for SINet-v2, ZoomNet, and PFNet.
-
Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations
Bench2Drive-Robust is a new closed-loop benchmark that evaluates end-to-end autonomous driving models under deployment perturbations from camera failures, ego-state errors, and compute delays, showing substantial performance degradation beyond image-level tests.
-
Benchmarking Sensor-Fault Robustness in Forecasting
SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.
-
Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
Q-Align trains LMMs on discrete text-defined levels for visual scoring, achieving SOTA on IQA, IAA, and VQA while unifying the tasks in OneAlign.
-
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
RoboStressBench decomposes visual stress into four physically grounded dimensions to benchmark VLM robustness in embodied scenes and proposes a stress-aware solver.
-
Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection
SemProbe is an interactive tool for semantic robustness probing of object detectors that applies controlled diffusion inpainting to user-masked deployment images and automatically compares detection performance before and after.
-
Safety-Critical Camera Reliability Monitoring for ADAS via Degradation-Aware Uncertainty Pattern Analysis
A multi-task network predicts degradation patterns and a multiplicative Global Sensor Health Index from RGB images to provide early warnings of camera failure in autonomous driving before downstream detection degrades.
-
Reward-Guided Semantic Evolution for Test-time Adaptive Object Detection
RGSE adapts text embeddings at test time via evolutionary search, using cosine similarity rewards from high-confidence visual proposals to improve open-vocabulary object detection under distribution shifts.
-
Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
A broad empirical benchmark shows how 15 existing test selection metrics perform for fault detection, performance estimation, and retraining under corrupted, adversarial, temporal, natural, and label shifts across image, text, and Android data.
-
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
An RL-trained lightweight agent uses MLLM perceptual rewards to perform efficient label-free image restoration, matching SOTA on full-reference metrics and surpassing prior work on no-reference metrics.
-
APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds
APCoTTA introduces a continual test-time adaptation method for ALS point cloud semantic segmentation using gradient-driven layer selection, entropy-based consistency loss, and random parameter interpolation, with new benchmarks showing mIoU gains of 9-14%.
-
TimberVision: A Multi-Task Dataset and Framework for Log-Component Segmentation and Tracking in Autonomous Forestry Operations
Introduces TimberVision dataset and multi-task framework for log-component segmentation, detection, and tracking in forestry operations using RGB images.
-
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
StableVLA adds an Information Bottleneck Adapter to VLA models that improves robustness to visual corruptions by 30% on average with under 10M extra parameters and no extra data, even when using a much smaller backbone.
-
Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs
Changes in Chain-of-Causation explanations under sensor perturbations correlate with 5.3× higher trajectory deviation in a driving VLA, and enabling such explanations yields 11.8% better accuracy.