Pith. sign in

REVIEW 2 cited by

Data Augmentation with Variational Autoencoder for Imbalanced Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07039 v1 pith:MDZGAIT7 submitted 2024-12-09 cs.LG

classification cs.LG
keywords dataimbalancedaddressapproachknownlearningmethodmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning from an imbalanced distribution presents a major challenge in predictive modeling, as it generally leads to a reduction in the performance of standard algorithms. Various approaches exist to address this issue, but many of them concern classification problems, with a limited focus on regression. In this paper, we introduce a novel method aimed at enhancing learning on tabular data in the Imbalanced Regression (IR) framework, which remains a significant problem. We propose to use variational autoencoders (VAE) which are known as a powerful tool for synthetic data generation, offering an interesting approach to modeling and capturing latent representations of complex distributions. However, VAEs can be inefficient when dealing with IR. Therefore, we develop a novel approach for generating data, combining VAE with a smoothed bootstrap, specifically designed to address the challenges of IR. We numerically investigate the scope of this method by comparing it against its competitors on simulations and datasets known for IR.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Taming Data Challenges in ML-based Security Tasks Using Generative AI

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Generative AI data augmentation, especially Nimai's sample-conditioned synthesis, improves several security classifiers and speeds drift recovery, but fails on tasks with noisy or overlapping labels.

  2. CARTGen-IR: Synthetic Tabular Data Generation for Imbalanced Regression

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A CART-based synthetic sampler with rarity-weighted resampling achieves state-of-the-art competitive results for imbalanced regression without target thresholds.

Pith tools