REVIEW 2 cited by
Partial Resampling of Imbalanced Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Imbalanced data is a frequently encountered problem in machine learning. Despite a vast amount of literature on sampling techniques for imbalanced data, there is a limited number of studies that address the issue of the optimal sampling ratio. In this paper, we attempt to fill the gap in the literature by conducting a large scale study of the effects of sampling ratio on classification accuracy. We consider 10 popular sampling methods and evaluate their performance over a range of ratios based on 20 datasets. The results of the numerical experiments suggest that the optimal sampling ratio is between 0.7 and 0.8 albeit the exact ratio varies depending on the dataset. Furthermore, we find that while factors such the original imbalance ratio or the number of features do not play a discernible role in determining the optimal ratio, the number of samples in the dataset may have a tangible effect.
Forward citations
Cited by 2 Pith papers
-
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
A solvable teacher-student perceptron model predicts that the optimal fraction of anomaly examples in training is generally away from 50%, with a sharp crossover as training noise increases.
-
Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring
Cost-sensitive credit models improve cost efficiency but produce less stable SHAP and LIME explanations, especially when the training data are imbalanced.
Discussion (0). Continue with ORCID to comment.