Vision foundation models from OpenAI and Meta are non-robust to nine categories of common perturbations, with new metrics linking robustness scores to downstream performance drops and a fine-tuning method proposed to improve stability without losing utility.
Indoor segmentation and support inference from rgbd images
2 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 2representative citing papers
Co-Me distills a confidence predictor to selectively merge low-confidence tokens in visual geometric transformers, delivering up to 21.5x speedup on VGGT and 20.4x on Pi3 while preserving spatial coverage and performance.
citing papers explorer
-
Robustness of Vision Foundation Models to Common Perturbations
Vision foundation models from OpenAI and Meta are non-robust to nine categories of common perturbations, with new metrics linking robustness scores to downstream performance drops and a fine-tuning method proposed to improve stability without losing utility.
-
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
Co-Me distills a confidence predictor to selectively merge low-confidence tokens in visual geometric transformers, delivering up to 21.5x speedup on VGGT and 20.4x on Pi3 while preserving spatial coverage and performance.