Pith. sign in

REVIEW 1 cited by

Generating and Imputing Tabular Data via Diffusion and Flow-based Gradient-Boosted Trees

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09968 v3 pith:Y6MDVYYE submitted 2023-09-18 cs.LG

classification cs.LG
keywords datatabulargenerationapproachdiffusiongeneratinggradient-boostedimputation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tabular data is hard to acquire and is subject to missing values. This paper introduces a novel approach for generating and imputing mixed-type (continuous and categorical) tabular data utilizing score-based diffusion and conditional flow matching. In contrast to prior methods that rely on neural networks to learn the score function or the vector field, we adopt XGBoost, a widely used Gradient-Boosted Tree (GBT) technique. To test our method, we build one of the most extensive benchmarks for tabular data generation and imputation, containing 27 diverse datasets and 9 metrics. Through empirical evaluation across the benchmark, we demonstrate that our approach outperforms deep-learning generation methods in data generation tasks and remains competitive in data imputation. Notably, it can be trained in parallel using CPUs without requiring a GPU. Our Python and R code is available at https://github.com/SamsungSAILMontreal/ForestDiffusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CFMI: Flow Matching for Missing Data Imputation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A conditional flow-matching model trained only on observed portions of data imputes missing entries competitively across 24 tabular and two time-series datasets.

Pith tools