Pith. sign in

REVIEW 16 cited by

Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10963 v3 pith:XHW4WRC6 submitted 2020-06-19 cs.LG stat.ML

Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift

classification cs.LG stat.ML
keywords shiftcovariatebatchdatadatasetmodelnormalizationprediction-time
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Covariate shift has been shown to sharply degrade both predictive accuracy and the calibration of uncertainty estimates for deep learning models. This is worrying, because covariate shift is prevalent in a wide range of real world deployment settings. However, in this paper, we note that frequently there exists the potential to access small unlabeled batches of the shifted data just before prediction time. This interesting observation enables a simple but surprisingly effective method which we call prediction-time batch normalization, which significantly improves model accuracy and calibration under covariate shift. Using this one line code change, we achieve state-of-the-art on recent covariate shift benchmarks and an mCE of 60.28\% on the challenging ImageNet-C dataset; to our knowledge, this is the best result for any model that does not incorporate additional data augmentation or modification of the training pipeline. We show that prediction-time batch normalization provides complementary benefits to existing state-of-the-art approaches for improving robustness (e.g. deep ensembles) and combining the two further improves performance. Our findings are supported by detailed measurements of the effect of this strategy on model behavior across rigorous ablations on various dataset modalities. However, the method has mixed results when used alongside pre-training, and does not seem to perform as well under more natural types of dataset shift, and is therefore worthy of additional study. We include links to the data in our figures to improve reproducibility, including a Python notebooks that can be run to easily modify our analysis at https://colab.research.google.com/drive/11N0wDZnMQQuLrRwRoumDCrhSaIhkqjof.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CRISP: Rank-Guided Iterative Squeezing for Robust Medical Image Segmentation under Domain Shift

    cs.CV 2026-04 unverdicted novelty 8.0

    CRISP uses the observed stability of positive voxel probability rankings under domain shift to build and iteratively refine high-precision and high-recall priors via latent feature perturbation, enabling parameter-fre...

  2. T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0

    T-VSS is a lightweight test-time defense that steers attacked visual features in VLMs using sample-specific low-rank subspaces and reliability-weighted entropy minimization to improve robustness.

  3. Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm

    cs.LG 2026-05 conditional novelty 7.0

    A framework to identify and convert foldable layer normalizations to RMSNorm for exact equivalence and faster inference in deep neural networks.

  4. Adaptive Camera Sensor for Vision Models

    cs.CV 2025-03 unverdicted novelty 7.0

    Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.

  5. CRISP: Constrained Refinement via Iterative Squeezing Process for Robust Medical Image Segmentation under Domain Shift

    cs.CV 2026-07 conditional novelty 6.0

    CRISP improves source-only medical image segmentation under domain shift by iteratively refining high-precision and high-recall masks derived from perturbation-based rank stability.

  6. Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification

    cs.CV 2026-06 unverdicted novelty 6.0

    A multi-level diversification wrapper for test-time adaptation that treats entropy minimization as multi-hypothesis inference to reduce underspecification and improve robustness by 1-4%.

  7. Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation

    cs.CV 2026-06 unverdicted novelty 6.0

    TopoTTA integrates persistent homology into test-time adaptation to derive topological pseudo-labels from anomaly maps, improving segmentation by an average 15% F1 on six benchmarks while generalizing across 2D and 3D data.

  8. Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging

    cs.LG 2026-06 unverdicted novelty 6.0

    Entropy minimization amplifies prediction bias from merged feature clusters under distribution shifts, and DSBR mitigates collapse by equalizing predicted class contributions to the unsupervised loss.

  9. Quantum Tunneling-Aware Machine Learning: Physics-Derived Noise Models for Robust Deployment

    cs.LG 2026-05 unverdicted novelty 6.0

    QTAML derives WKB-based tunneling noise models for AI weights with affine mean drift and per-bit variance hierarchy, then uses them in TAC to achieve 95% clean accuracy with 3.4-33.6x less ECC overhead than baselines ...

  10. GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation

    cs.CV 2026-05 unverdicted novelty 6.0

    Diversity-aware memory policies improve test-time adaptation performance most under constrained memory budgets and challenging non-i.i.d. streams.

  11. TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 6.0

    TAME uses a Mixture-of-Experts prompt bank with input-dependent routing and three unsupervised objectives to adaptively defend CLIP against adversarial attacks at inference time, achieving at least 49.1% robustness ga...

  12. Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation

    cs.LG 2026-04 unverdicted novelty 6.0

    Deployment-time Shortcut Guardrail uses unsupervised gradient attribution on a converged text encoder alone to mitigate shortcut learning and match training-time baselines under distribution shift.

  13. Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation

    cs.LG 2026-04 unverdicted novelty 6.0

    Shortcut Guardrail mitigates token-level shortcuts in pretrained language models at deployment time via gradient-based token identification and a LoRA-trained Masked Contrastive Learning module, improving accuracy und...

  14. Tent: Fully Test-time Adaptation by Entropy Minimization

    cs.LG 2020-06 conditional novelty 6.0

    Test-time entropy minimization adapts models by optimizing for confident predictions, reducing error on corrupted ImageNet-C and enabling source-free domain adaptation.

  15. Threshold Modulation for Online Test-Time Adaptation of Spiking Neural Networks

    cs.CV 2025-05 unverdicted novelty 5.0

    Threshold Modulation dynamically adjusts firing thresholds in SNNs via neuronal dynamics-inspired normalization to enable online test-time adaptation under distribution shifts.

  16. GMN4AD: Graph Matching Network for Alzheimer's Disease Diagnosis with Test-Time Domain Adaptation using Multi-centered Structure Magnetic Resonance Imaging

    eess.IV 2026-06 unverdicted novelty 4.0

    GMN4AD applies graph matching and test-time contrastive adaptation to improve Alzheimer's diagnosis accuracy on heterogeneous multi-center sMRI datasets.