Pith. sign in

REVIEW 8 cited by

Diagnosing Model Performance Under Distribution Shift

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.02011 v4 pith:PB26J7ZZ submitted 2023-03-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords distributionperformancetrainingtargetconditionaldifferentdropexamples
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Prediction models can perform poorly when deployed to target distributions different from the training distribution. To understand these operational failure modes, we develop a method, called DIstribution Shift DEcomposition (DISDE), to attribute a drop in performance to different types of distribution shifts. Our approach decomposes the performance drop into terms for 1) an increase in harder but frequently seen examples from training, 2) changes in the relationship between features and outcomes, and 3) poor performance on examples infrequent or unseen during training. These terms are defined by fixing a distribution on $X$ while varying the conditional distribution of $Y \mid X$ between training and target, or by fixing the conditional distribution of $Y \mid X$ while varying the distribution on $X$. In order to do this, we define a hypothetical distribution on $X$ consisting of values common in both training and target, over which it is easy to compare $Y \mid X$ and thus predictive performance. We estimate performance on this hypothetical distribution via reweighting methods. Empirically, we show how our method can 1) inform potential modeling improvements across distribution shifts for employment prediction on tabular census data, and 2) help to explain why certain domain adaptation methods fail to improve model performance for satellite image classification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncovering Bias Mechanisms in Observational Studies

    stat.ME 2025-06 conditional novelty 7.0 of 10

    Covariances between the size of causal bias and conditional variances of treatment, selection, and outcome form a fingerprint that distinguishes transportability, confounding, and selection bias mechanisms.

  2. "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance Drift

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SHIFT is a hierarchical hypothesis-testing method that detects subgroups with large model performance decay under distribution shift and explains the decay via variable-subset-specific covariate or outcome shifts.

  3. General and Estimable Learning Bound Unifying Covariate and Concept Shifts

    stat.ML 2025-06 conditional novelty 6.0 of 10

    The authors define a total pair concept shift on the optimal transport plan between source and target covariates, yielding a Lipschitz-based target error bound that handles stochastic labels, general losses, and misma...

  4. Explaining Concept Shift with Interpretable Feature Attribution

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SGShift attributes concept shift to a sparse set of features by fitting a penalized generalized additive update term on top of the source model.

  5. When the Past Misleads: Rethinking Training Data Expansion Under Temporal Distribution Shifts

    cs.CY 2025-09 conditional novelty 5.0 of 10

    Expanding the historical training window can degrade model performance and fairness under concept shift, with the harm appearing mainly when training data are large.

  6. Realistic Evaluation of TabPFN v2 in Open Environments

    cs.LG 2025-05 conditional novelty 5.0 of 10

    TabPFN v2 underperforms tree-based models on most open-environment tabular tasks and is only preferable on small, covariate-shifted, class-balanced data.

  7. Data Curation Matters: Model Collapse and Spurious Shift Performance Prediction from Training on Uncurated Text Embeddings

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Training on LLM text embeddings can cause tabular classifiers to collapse to single-class predictions, which spuriously inflates Accuracy-on-the-Line correlations.

  8. Data Heterogeneity Modeling for Trustworthy Machine Learning

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A survey that frames heterogeneity-aware machine learning as a paradigm spanning data collection, training, evaluation, and deployment, drawing mostly on the authors' prior results.

Pith tools