Pith. sign in

REVIEW 4 cited by

TabPFGen -- Tabular Data Generation with TabPFN

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05216 v1 pith:SEKNHFE6 submitted 2024-06-07 cs.LG

classification cs.LG
keywords datatabulargenerativetabpfnmodelstabpfgendiscriminativeenergy-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in deep generative modelling have not translated well to tabular data. We argue that this is caused by a mismatch in structure between popular generative models and discriminative models of tabular data. We thus devise a technique to turn TabPFN -- a highly performant transformer initially designed for in-context discriminative tabular tasks -- into an energy-based generative model, which we dub TabPFGen. This novel framework leverages the pre-trained TabPFN as part of the energy function and does not require any additional training or hyperparameter tuning, thus inheriting TabPFN's in-context learning capability. We can sample from TabPFGen analogously to other energy-based models. We demonstrate strong results on standard generative modelling tasks, including data augmentation, class-balancing, and imputation, unlocking a new frontier of tabular data generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CORE: In-Context Reconstruction for Unified Tabular Anomaly Detection

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CORE detects anomalies in new tabular datasets by reconstructing each test sample from the nearest normal context samples in a learned, feature-aligned space; it is proposed as the first reconstruction-based unified t...

  2. Risk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    LLM-based tabular generators reproduce seed rows often enough that membership-inference attacks succeed more against them than against GAN, VAE, or diffusion baselines.

  3. A Closer Look on Memorization in Tabular Diffusion Model: A Data-Centric Perspective

    cs.LG 2025-05 reject novelty 6.0 of 10

    A small subset of training samples drives most memorization in tabular diffusion models, and pruning them based on early memorization signals reduces measured leakage, though the evaluation metric makes part of the ga...

  4. Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    TabularARGN is a discretization-based auto-regressive network claimed to generate high-fidelity, privacy-robust synthetic tabular data, competitive with diffusion and GAN baselines.

Pith tools