Pith. sign in

REVIEW 2 cited by

Partial Resampling of Imbalanced Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.04631 v1 pith:G7HBSOYO submitted 2022-07-11 cs.LG

classification cs.LG
keywords ratiosamplingdataimbalancednumberoptimaldatasetliterature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imbalanced data is a frequently encountered problem in machine learning. Despite a vast amount of literature on sampling techniques for imbalanced data, there is a limited number of studies that address the issue of the optimal sampling ratio. In this paper, we attempt to fill the gap in the literature by conducting a large scale study of the effects of sampling ratio on classification accuracy. We consider 10 popular sampling methods and evaluate their performance over a range of ratios based on 20 datasets. The results of the numerical experiments suggest that the optimal sampling ratio is between 0.7 and 0.8 albeit the exact ratio varies depending on the dataset. Furthermore, we find that while factors such the original imbalance ratio or the number of features do not play a discernible role in determining the optimal ratio, the number of samples in the dataset may have a tangible effect.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A solvable teacher-student perceptron model predicts that the optimal fraction of anomaly examples in training is generally away from 50%, with a sharp crossover as training noise increases.

  2. Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Cost-sensitive credit models improve cost efficiency but produce less stable SHAP and LIME explanations, especially when the training data are imbalanced.

Pith tools