REVIEW 16 cited by
Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift
read the original abstract
Covariate shift has been shown to sharply degrade both predictive accuracy and the calibration of uncertainty estimates for deep learning models. This is worrying, because covariate shift is prevalent in a wide range of real world deployment settings. However, in this paper, we note that frequently there exists the potential to access small unlabeled batches of the shifted data just before prediction time. This interesting observation enables a simple but surprisingly effective method which we call prediction-time batch normalization, which significantly improves model accuracy and calibration under covariate shift. Using this one line code change, we achieve state-of-the-art on recent covariate shift benchmarks and an mCE of 60.28\% on the challenging ImageNet-C dataset; to our knowledge, this is the best result for any model that does not incorporate additional data augmentation or modification of the training pipeline. We show that prediction-time batch normalization provides complementary benefits to existing state-of-the-art approaches for improving robustness (e.g. deep ensembles) and combining the two further improves performance. Our findings are supported by detailed measurements of the effect of this strategy on model behavior across rigorous ablations on various dataset modalities. However, the method has mixed results when used alongside pre-training, and does not seem to perform as well under more natural types of dataset shift, and is therefore worthy of additional study. We include links to the data in our figures to improve reproducibility, including a Python notebooks that can be run to easily modify our analysis at https://colab.research.google.com/drive/11N0wDZnMQQuLrRwRoumDCrhSaIhkqjof.
Forward citations
Cited by 16 Pith papers
-
CRISP: Rank-Guided Iterative Squeezing for Robust Medical Image Segmentation under Domain Shift
CRISP uses the observed stability of positive voxel probability rankings under domain shift to build and iteratively refine high-precision and high-recall priors via latent feature perturbation, enabling parameter-fre...
-
T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models
T-VSS is a lightweight test-time defense that steers attacked visual features in VLMs using sample-specific low-rank subspaces and reliability-weighted entropy minimization to improve robustness.
-
Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm
A framework to identify and convert foldable layer normalizations to RMSNorm for exact equivalence and faster inference in deep neural networks.
-
Adaptive Camera Sensor for Vision Models
Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.
-
CRISP: Constrained Refinement via Iterative Squeezing Process for Robust Medical Image Segmentation under Domain Shift
CRISP improves source-only medical image segmentation under domain shift by iteratively refining high-precision and high-recall masks derived from perturbation-based rank stability.
-
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
A multi-level diversification wrapper for test-time adaptation that treats entropy minimization as multi-hypothesis inference to reduce underspecification and improve robustness by 1-4%.
-
Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation
TopoTTA integrates persistent homology into test-time adaptation to derive topological pseudo-labels from anomaly maps, improving segmentation by an average 15% F1 on six benchmarks while generalizing across 2D and 3D data.
-
Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging
Entropy minimization amplifies prediction bias from merged feature clusters under distribution shifts, and DSBR mitigates collapse by equalizing predicted class contributions to the unsupervised loss.
-
Quantum Tunneling-Aware Machine Learning: Physics-Derived Noise Models for Robust Deployment
QTAML derives WKB-based tunneling noise models for AI weights with affine mean drift and per-bit variance hierarchy, then uses them in TAC to achieve 95% clean accuracy with 3.4-33.6x less ECC overhead than baselines ...
-
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
Diversity-aware memory policies improve test-time adaptation performance most under constrained memory budgets and challenging non-i.i.d. streams.
-
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
TAME uses a Mixture-of-Experts prompt bank with input-dependent routing and three unsupervised objectives to adaptively defend CLIP against adversarial attacks at inference time, achieving at least 49.1% robustness ga...
-
Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation
Deployment-time Shortcut Guardrail uses unsupervised gradient attribution on a converged text encoder alone to mitigate shortcut learning and match training-time baselines under distribution shift.
-
Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation
Shortcut Guardrail mitigates token-level shortcuts in pretrained language models at deployment time via gradient-based token identification and a LoRA-trained Masked Contrastive Learning module, improving accuracy und...
-
Tent: Fully Test-time Adaptation by Entropy Minimization
Test-time entropy minimization adapts models by optimizing for confident predictions, reducing error on corrupted ImageNet-C and enabling source-free domain adaptation.
-
Threshold Modulation for Online Test-Time Adaptation of Spiking Neural Networks
Threshold Modulation dynamically adjusts firing thresholds in SNNs via neuronal dynamics-inspired normalization to enable online test-time adaptation under distribution shifts.
-
GMN4AD: Graph Matching Network for Alzheimer's Disease Diagnosis with Test-Time Domain Adaptation using Multi-centered Structure Magnetic Resonance Imaging
GMN4AD applies graph matching and test-time contrastive adaptation to improve Alzheimer's diagnosis accuracy on heterogeneous multi-center sMRI datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.