Pith. sign in

REVIEW 3 cited by

Improved Algorithm for Deep Active Learning under Imbalance via Optimal Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09196 v4 pith:GQNXCE2N submitted 2023-12-14 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords learningactiveannotationclassdirectimbalancelabelnoise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Class imbalance severely impacts machine learning performance on minority classes in real-world applications. While various solutions exist, active learning offers a fundamental fix by strategically collecting balanced, informative labeled examples from abundant unlabeled data. We introduce DIRECT, an algorithm that identifies class separation boundaries and selects the most uncertain nearby examples for annotation. By reducing the problem to one-dimensional active learning, DIRECT leverages established theory to handle batch labeling and label noise -- another common challenge in data annotation that particularly affects active learning methods. Our work presents the first comprehensive study of active learning under both class imbalance and label noise. Extensive experiments on imbalanced datasets show DIRECT reduces annotation costs by over 60\% compared to state-of-the-art active learning methods and over 80\% versus random sampling, while maintaining robustness to label noise.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Active Learning for Lung Disease Severity Classification from Chest X-rays: Learning with Less Data in the Presence of Class Imbalance

    eess.IV 2025-08 conditional novelty 4.0 of 10

    Simple uncertainty-based active learning with a weighted loss matched full-data COVID-19 CXR severity classifiers using 15.4% (binary) and 23.1% (multi-class) of training labels.

  2. Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Weighted Task Diversity allocates the annotation budget across tasks in inverse proportion to the base model's average confidence, improving MMLU and AlpacaEval scores with up to 80% fewer labels.

  3. Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering

    cs.LG 2025-07 reject novelty 4.0 of 10

    PRRO combines signal-based data pruning and column reordering to improve the supervised learning utility of synthetic tabular data, but its evaluation is undermined by data manipulation and an ill-defined correlation measure.

Pith tools