Pith. sign in

REVIEW 1 cited by

Enhancing Synthetic Oversampling for Imbalanced Datasets Using Proxima-Orion Neighbors and q-Gaussian Weighting Technique

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15790 v1 pith:POAUJZ7C submitted 2025-01-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords instancesalgorithmclassdatasetsproposeddatasetimbalancedminority
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this article, we propose a novel oversampling algorithm to increase the number of instances of minority class in an imbalanced dataset. We select two instances, Proxima and Orion, from the set of all minority class instances, based on a combination of relative distance weights and density estimation of majority class instances. Furthermore, the q-Gaussian distribution is used as a weighting mechanism to produce new synthetic instances to improve the representation and diversity. We conduct a comprehensive experiment on 42 datasets extracted from KEEL software and eight datasets from the UCI ML repository to evaluate the usefulness of the proposed (PO-QG) algorithm. Wilcoxon signed-rank test is used to compare the proposed algorithm with five other existing algorithms. The test results show that the proposed technique improves the overall classification performance. We also demonstrate the PO-QG algorithm to a dataset of Indian patients with sarcopenia.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Kolmogorov Arnold Networks (KANs) for Imbalanced Data -- An Empirical Perspective

    cs.LG 2025-07 conditional novelty 4.0 of 10

    On ten KEEL datasets, KANs outperform MLPs on raw imbalanced data but resampling and focal loss degrade KANs while MLPs with those techniques match KAN performance at far lower cost.

Pith tools