Pith. sign in

REVIEW 4 cited by

Detecting and Correcting for Label Shift with Black Box Predictors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.03916 v3 pith:6MMFQNR5 submitted 2018-02-12 cs.LG cs.AIcs.NEstat.ML

classification cs.LGcs.AIcs.NEstat.ML
keywords shiftbbsepredictorstestblacklabelclassifierscorrect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Faced with distribution shift between training and test set, we wish to detect and quantify the shift, and to correct our classifiers without test set labels. Motivated by medical diagnosis, where diseases (targets) cause symptoms (observations), we focus on label shift, where the label marginal $p(y)$ changes but the conditional $p(x| y)$ does not. We propose Black Box Shift Estimation (BBSE) to estimate the test distribution $p(y)$. BBSE exploits arbitrary black box predictors to reduce dimensionality prior to shift correction. While better predictors give tighter estimates, BBSE works even when predictors are biased, inaccurate, or uncalibrated, so long as their confusion matrices are invertible. We prove BBSE's consistency, bound its error, and introduce a statistical test that uses BBSE to detect shift. We also leverage BBSE to correct classifiers. Experiments demonstrate accurate estimates and improved prediction, even on high-dimensional datasets of natural images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Estimating prevalence with precision and accuracy

    stat.ML 2025-07 conditional novelty 6.0 of 10

    A new Bayesian quantifier, PQ, produces tighter and well-calibrated prediction intervals for class prevalence estimates, beating existing methods across simulated and real datasets.

  2. Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A new evaluation metric, the DCA log score, averages cost-weighted accuracy over a bounded, logit-uniform range of class prevalences, linking calibration, label shift, and error costs in one closed-form score.

  3. Hidden-Domain Routing for All-Type Audio Deepfake Detection

    cs.SD 2026-08 accept novelty 5.0 of 10

    A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.

  4. Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring

    cs.LG 2025-06 reject novelty 5.0 of 10

    CAMA applies a three-step feature engineering procedure to LLM agents and reports 55 to 92 percent accuracy on ML monitoring report questions, outperforming six baselines.

Pith tools