Pith. sign in

REVIEW 1 cited by

Restoring balance: principled under/oversampling of data for optimal classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09535 v2 pith:VR35ULVJ submitted 2024-05-15 cond-mat.dis-nn cs.LG

classification cond-mat.dis-nncs.LG
keywords dataoversamplingstrategiesunderclassdependinggeneralizationimbalance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Class imbalance in real-world data poses a common bottleneck for machine learning tasks, since achieving good generalization on under-represented examples is often challenging. Mitigation strategies, such as under or oversampling the data depending on their abundances, are routinely proposed and tested empirically, but how they should adapt to the data statistics remains poorly understood. In this work, we determine exact analytical expressions of the generalization curves in the high-dimensional regime for linear classifiers (Support Vector Machines). We also provide a sharp prediction of the effects of under/oversampling strategies depending on class imbalance, first and second moments of the data, and the metrics of performance considered. We show that mixed strategies involving under and oversampling of data lead to performance improvement. Through numerical experiments, we show the relevance of our theoretical predictions on real datasets, on deeper architectures and with sampling strategies based on unsupervised probabilistic models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A solvable teacher-student perceptron model predicts that the optimal fraction of anomaly examples in training is generally away from 50%, with a sharp crossover as training noise increases.

Pith tools