Pith. sign in

REVIEW 13 cited by

SMOTE: Synthetic Minority Over-sampling Technique

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1106.1813 v1 pith:GX57CDO7 submitted 2011-06-09 cs.AI

classification cs.AI
keywords classminorityclassifiermajoritymethodnormalover-samplingunder-sampling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An approach to the construction of classifiers from imbalanced datasets is described. A dataset is imbalanced if the classification categories are not approximately equally represented. Often real-world data sets are predominately composed of "normal" examples with only a small percentage of "abnormal" or "interesting" examples. It is also the case that the cost of misclassifying an abnormal (interesting) example as a normal example is often much higher than the cost of the reverse error. Under-sampling of the majority (normal) class has been proposed as a good means of increasing the sensitivity of a classifier to the minority class. This paper shows that a combination of our method of over-sampling the minority (abnormal) class and under-sampling the majority (normal) class can achieve better classifier performance (in ROC space) than only under-sampling the majority class. This paper also shows that a combination of our method of over-sampling the minority class and under-sampling the majority class can achieve better classifier performance (in ROC space) than varying the loss ratios in Ripper or class priors in Naive Bayes. Our method of over-sampling the minority class involves creating synthetic minority class examples. Experiments are performed using C4.5, Ripper and a Naive Bayes classifier. The method is evaluated using the area under the Receiver Operating Characteristic curve (AUC) and the ROC convex hull strategy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neighbor displacement-based enhanced synthetic oversampling for multiclass imbalanced data

    cs.LG 2025-01 reject novelty 6.0 of 10

    NDESO moves noisy minority points toward their class centroids before random oversampling and is claimed to beat 14 resamplers, but its cross-validation protocol casts doubt on the headline result.

  2. Uncertainty-Aware Tidal Disruption Event Classification : A Host-Agnostic Probabilistic Random Forest Approach

    astro-ph.HE 2026-07 conditional novelty 5.5 of 10

    Uncertainty-aware Probabilistic Random Forest on 11 photometric light-curve features classifies TDEs without host data more stably than XGBoost and recovers 14 new candidates from ZTF.

  3. S-PLUS Clusters And Large-scale Environments (SCALE): I. A catalog of known clusters and groups in DR5 and a pilot study of Abell 4038

    astro-ph.CO 2026-07 conditional novelty 5.0 of 10

    SCALE delivers dynamical properties for 83 known z≤0.1 systems and uses S-PLUS photo-zs plus RPM to map Abell 4038 substructure and mass-dependent color-concentration trends.

  4. Stellar flare detection in XMM-Newton with gradient boosted trees

    astro-ph.HE 2025-09 conditional novelty 5.0 of 10

    A gradient boosted classifier on X-ray light curve features detects stellar flares at 97.1% test accuracy and generates the largest public catalog of such events.

  5. A statistical study of lopsided galaxies using random forest

    astro-ph.GA 2024-11 conditional novelty 5.0 of 10

    A random forest trained on TNG50 internal galaxy properties preclassifies lopsided versus symmetric disk galaxies with about 80% balanced accuracy, and similar accuracy is reached with photometric observables alone.

  6. Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Masking the loss of head relationship classes during training raises mean recall and corruption robustness of video scene graph generation and anticipation models on Action Genome.

  7. Class Imbalance Corrections Failed to Enhance Discrimination, Model Calibration, and Prediction Stability: An Empirical Simulation Study Based on Clinical Dataset

    stat.ME 2026-06 conditional novelty 4.0 of 10

    Simulation on GUSTO-I data shows class imbalance corrections fail to boost discrimination and impair calibration plus stability in clinical prediction models.

  8. Visibility nowcasting in South Korea: a machine learning approach to class imbalance and distribution shift

    physics.ao-ph 2026-05 unverdicted novelty 4.0 of 10

    The study applies an ensemble of machine learning and deep learning models with synthetic oversampling on 2018-2020 data to nowcast visibility, finding a performance decline on 2021 test data attributed to distributio...

  9. Attentive Dilated Convolution for Automatic Sleep Staging using Force-directed Layout

    eess.SP 2024-08 unverdicted novelty 4.0 of 10

    AttDiCNN reaches 98.56%, 99.66%, and 99.08% accuracy on EDFX, HMC, and NCH sleep datasets via force-directed visibility graph EEG representations and a three-module attentive dilated CNN architecture.

  10. Beyond Synthetic Augmentation: Group-Aware Threshold Calibration for Robust Balanced Accuracy in Imbalanced Learning

    cs.LG 2025-08 conditional novelty 3.0 of 10

    On two imbalanced financial benchmarks, group-specific decision thresholds outperform or match SMOTE and CT-GAN augmentation across seven model families.

  11. STORM: Strategic Orchestration of Modalities for Rare Event Classification

    cs.CV 2024-12 conditional novelty 3.0 of 10

    STORM uses entropy imbalance and decision-tree logic to select informative modalities for rare-event classification, and reports that temporal expert features do not help SOZ detection.

  12. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

    cs.CR 2025-12 conditional novelty 2.0 of 10

    A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.

  13. Synthetic Tabular Data: Methods, Attacks and Defenses

    cs.LG 2025-06 conditional novelty 1.0 of 10

    A review of tabular synthetic data generation, privacy attacks, and defenses, whose central message is that synthetic data alone does not guarantee privacy.

Pith tools