Pith. sign in

REVIEW 3 cited by

TabularARGN: A Flexible and Efficient Auto-Regressive Framework for Generating High-Fidelity Synthetic Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12012 v2 pith:JF5JVZHO submitted 2025-01-21 cs.LG

classification cs.LG
keywords datadatasetsframeworkgenerationsynthetictabularargnacrossauto-regressive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce the Tabular Auto-Regressive Generative Network (TabularARGN), a flexible framework designed to handle mixed-type, multivariate, and sequential datasets. By training on all possible conditional probabilities, TabularARGN supports advanced features such as fairness-aware generation, imputation, and conditional generation on any subset of columns. The framework achieves state-of-the-art synthetic data quality while significantly reducing training and inference times, making it ideal for large-scale datasets with diverse structures. Evaluated across established benchmarks, including realistic datasets with complex relationships, TabularARGN demonstrates its capability to synthesize high-quality data efficiently. By unifying flexibility and performance, this framework paves the way for practical synthetic data generation across industries.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Static-distribution fidelity is a poor proxy for temporal fidelity in synthetic sequential tabular data; measuring timestamp, trajectory, cross-sectional, and relational structure over time changes model rankings.

  2. Disjoint Generation of Synthetic Data

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A framework that generates tabular synthetic data by partitioning columns, training separate generative models, and joining the outputs with a learned validator improves empirical privacy while keeping utility competitive.

  3. Achieving Hilbert-Schmidt Independence Under R\'enyi Differential Privacy for Fair and Private Data Generation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    FLIP combines a VAE, latent diffusion, Rényi DP, and CKA alignment across protected groups to produce tabular data with substantially reduced predictability of the protected attribute.

Pith tools