Pith. sign in

REVIEW 2 cited by

On Statistical Bias In Active Learning: How and When To Fix It

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.11665 v2 pith:HXGEBUUX submitted 2021-01-27 stat.ML cs.LG

classification stat.MLcs.LG
keywords biaswhenactivedatalearninghelpfultrainingactively
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Active learning is a powerful tool when labelling data is expensive, but it introduces a bias because the training data no longer follows the population distribution. We formalize this bias and investigate the situations in which it can be harmful and sometimes even helpful. We further introduce novel corrective weights to remove bias when doing so is beneficial. Through this, our work not only provides a useful mechanism that can improve the active learning approach, but also an explanation of the empirical successes of various existing approaches which ignore this bias. In particular, we show that this bias can be actively helpful when training overparameterized models -- like neural networks -- with relatively little data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. Consensus-Driven Active Model Selection

    cs.LG 2025-07 conditional novelty 7.0 of 10

    CODA uses consensus-based priors and Bayesian updating to select the best candidate model with far fewer labels than prior active model selection methods, beating them on 18 of 26 benchmark tasks.

  2. Prediction-Powered Active Testing

    stat.ML 2026-07 accept novelty 6.0 of 10

    PPAT residualizes losses via a prediction-powered control variate inside LURE, yielding lower-variance unbiased risk estimates, tailored acquisition, and asymptotic CIs that cover with fewer labels.

Pith tools