REVIEW 8 cited by
Modeling Tabular data using Conditional GAN
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modeling the probability distribution of rows in tabular data and generating realistic synthetic data is a non-trivial task. Tabular data usually contains a mix of discrete and continuous columns. Continuous columns may have multiple modes whereas discrete columns are sometimes imbalanced making the modeling difficult. Existing statistical and deep neural network models fail to properly model this type of data. We design TGAN, which uses a conditional generative adversarial network to address these challenges. To aid in a fair and thorough comparison, we design a benchmark with 7 simulated and 8 real datasets and several Bayesian network baselines. TGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not.
Forward citations
Cited by 8 Pith papers
-
FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents
FairDiffuseVQVAE reaches state-of-the-art fairness on the standard tabular benchmark (DPR 0.702, EOR 0.686) by uniform protected-attribute sampling at inference, paying ~15 AUC points of utility.
-
Friend or Foe
Friend or Foe is a 64-dataset compendium of 26M+ simulated bacterial interaction environments, with benchmarks showing deep tabular models classify interaction type with mean MCC of about 0.64.
-
CARTGen-IR: Synthetic Tabular Data Generation for Imbalanced Regression
A CART-based synthetic sampler with rarity-weighted resampling achieves state-of-the-art competitive results for imbalanced regression without target thresholds.
-
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.
-
From Data to Decision: A Multi-Stage Framework for Class Imbalance Mitigation in Optical Network Failure Analysis
On experimental optical-network data, threshold adjustment improves failure-detection F1 by up to 15.3%, while CTGAN data augmentation improves failure-identification F1 by up to 24.2%.
-
Pre-, In-, and Post-Processing Class Imbalance Mitigation Techniques for Failure Detection in Optical Networks
On one experimental optical network dataset, post-processing threshold adjustment improved failure-detection F1 by 15.3% over an unbalanced random forest baseline, more than any pre- or in-processing method tested.
-
CopulaSMOTE: A Copula-Based Oversampling Approach for Imbalanced Classification in Diabetes Prediction
A vine copula-based oversampler that preserves empirical margins improves minority-class F1 and recall on the CDC diabetes dataset, but not uniformly across classifiers and metrics.
-
Beyond Synthetic Augmentation: Group-Aware Threshold Calibration for Robust Balanced Accuracy in Imbalanced Learning
On two imbalanced financial benchmarks, group-specific decision thresholds outperform or match SMOTE and CT-GAN augmentation across seven model families.
Discussion (0). Sign in to comment.