Pith. sign in

REVIEW 2 minor 2 cited by

Diffusion and Flow Matching Models for Tabular Data: A Survey

T0 review · 0 major / 2 minor · reviewed 2026-05-25 · grok-4.3

Pith's one-line read This is the first survey dedicated to diffusion and flow matching models for tabular data.

desk verdict This is the first survey on diffusion and flow matching for tabular data and it organizes the scattered literature around practical challenges. read the letter →

arxiv 2502.17119 v2 pith:RXK2QDN2 submitted 2025-02-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords diffusionmodelsflowmatchingtabulardatagenerativesurveysynthesisimputationanomalydetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Tabular data generation faces persistent difficulties from mixed numerical and categorical features, missing values, imbalances, and domain constraints that earlier GAN and VAE approaches often handle unstably. Diffusion models address this through iterative noising and denoising, while flow matching learns direct transport fields, both offering more stable training for tasks like synthesis, imputation, and anomaly detection. The paper collects and organizes the scattered literature on these methods, identifies why direct comparisons remain elusive, and flags open issues in scalability, privacy, and constraint handling. A reader would care because tabular records dominate real-world datasets where reliable generative tools could improve data sharing and augmentation.

What carries the argument

The survey's four-way organizational structure around data-engineering challenges, tasks, design choices, and evaluation dimensions.

What would settle it

Discovery of any earlier survey whose scope is limited to diffusion and flow matching models applied to tabular data.

Watch

Extended reading notes

Core claim

To the best of our knowledge, this is the first survey dedicated specifically to diffusion and flow matching models for tabular data. We review work from June 2015 to May 2026, organize it around data-engineering challenges, tasks, design choices, and evaluation dimensions, and discuss open problems in scalability, feature dependency modeling, privacy, fairness, benchmarking, and constraint-aware generation.

Load-bearing premise

The literature on diffusion and flow matching models for tabular data remains difficult to compare because methods target different tasks and rely on different representations, objectives, evaluation protocols, and domain assumptions.

Editorial extensions

If this is right

  • Researchers can use the organization to locate methods for specific tabular tasks such as synthesis or imputation.
  • Future work must address the documented gaps in scalability and constraint-aware generation.
  • Standardized benchmarks would reduce the current fragmentation in evaluation protocols.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A shared evaluation protocol across tasks could accelerate progress by making incremental improvements visible.
  • Constraint-aware variants may prove essential for regulated domains where synthetic data must obey hard rules.
  • Privacy and fairness analyses could be integrated into the generative process rather than applied after the fact.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The manuscript is a survey of diffusion and flow matching models for tabular data, claiming to be the first dedicated review of the topic. It reviews literature from June 2015 to May 2026, organizes existing work around data-engineering challenges, tasks, design choices, and evaluation dimensions, and discusses open problems including scalability, feature dependency modeling, privacy, fairness, benchmarking, and constraint-aware generation. The authors state that they maintain updates in a GitHub repository.

Significance. If the coverage is comprehensive and free of selection bias, the survey would be significant for organizing an emerging, heterogeneous literature on generative models for structured data. The explicit maintenance of a GitHub repository for updates strengthens the work by providing a mechanism for ongoing relevance and community contribution.

minor comments (2)
  1. [Abstract] The review period is stated as extending to May 2026. The authors should clarify whether this is a projected cutoff, a typographical error, or the intended scope, as the current date of the manuscript appears to precede this endpoint.
  2. [Abstract] The abstract refers to a GitHub repository for updates but does not provide the URL. Including the repository link in the manuscript (and ideally in the abstract) would improve accessibility.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the constructive review and the recommendation of minor revision. The assessment correctly identifies the survey's scope, organization around data-engineering challenges and tasks, coverage of open problems, and the value of the maintained GitHub repository. No specific major comments were provided in the report.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity in survey paper

full rationale

This manuscript is explicitly a literature survey with no derivations, equations, predictions, or technical claims whose validity depends on internal self-reference. The sole novel assertion (being the first dedicated survey) is a factual statement about external literature coverage rather than a result derived from the paper's own inputs. No self-citation chains, fitted parameters renamed as predictions, or ansatzes are present. The work is therefore self-contained against external benchmarks with score 0.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is a literature survey with no free parameters, axioms, or invented entities introduced by the authors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Diffusion and Flow Matching Models for Tabular Data: A Survey." pith.science (2026). https://pith.science/paper/RXK2QDN2

@misc{pith2026250217119,
  author       = {Pith},
  title        = {Pith review of: Diffusion and Flow Matching Models for Tabular Data: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RXK2QDN2}},
  note         = {Machine review of arXiv:2502.17119}
}
read the original abstract

Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured records. For tabular data, however, generative modeling remains difficult: a dataset may contain numerical and categorical attributes, missing values, sensitive fields, imbalanced categories, complex feature dependencies, and domain constraints. Earlier tabular data modeling methods based on GANs or VAEs have achieved useful results, but they can suffer from unstable training, mode collapse, weak modeling of multimodal distributions, and fragile handling of mixed-type features. Diffusion models have therefore attracted growing interest because their noising-and-denoising formulation provides a flexible and stable way to model complex data distributions, and has been adapted to tabular synthesis, missing-value imputation, trustworthy data generation, and anomaly detection. Flow matching offers a closely related route by learning transport vector fields along probability paths, often with more direct control over path design and sampling efficiency. Despite this progress, the literature on diffusion and flow matching models for tabular data remains difficult to compare because methods target different tasks and rely on different representations, objectives, evaluation protocols, and domain assumptions. To the best of our knowledge, this is the first survey dedicated specifically to diffusion and flow matching models for tabular data. We review work from June 2015 to May 2026, organize it around data-engineering challenges, tasks, design choices, and evaluation dimensions, and discuss open problems in scalability, feature dependency modeling, privacy, fairness, benchmarking, and constraint-aware generation. We maintain updates in a GitHub repository.

Figures

Figures reproduced from arXiv: 2502.17119 by the authors.

Figure 1
Figure 1. Timeline of Generative Models for Tabular Data: Below the timeline, key advancements in traditional machine learning models and deep generative [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of Diffusion Models for Tabular Data. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Alternating clustering and GAN-based imputation in a feedback loop yields more accurate missing-value recovery on heterogeneous data than single-distribution methods.

  2. Diffusion Models in Finance: A Survey

    q-fin.CP 2026-08 conditional novelty 4.0 of 10

    A structured survey of diffusion-family generative models in finance, organized by financial data type, with an open-source reference repository.

Reference graph

Works this paper leans on

154 extracted references · 154 canonical work pages · cited by 2 Pith papers

  1. [1]

    Data mining in healthcare and biomedicine: a survey of the literature,

    I. Yoo, P. Alafaireet, M. Marinov, K. Pena-Hernandez, R. Gopidi, J.- F. Chang, and L. Hua, “Data mining in healthcare and biomedicine: a survey of the literature,” Journal of medical systems , vol. 36, pp. 2431–2448, 2012

  2. [2]

    M. F. Dixon, I. Halperin, and P. Bilokon, Machine learning in finance. Springer, 2020, vol. 1170

  3. [3]

    Data mining in education,

    A. Algarni, “Data mining in education,” International Journal of Advanced Computer Science and Applications , vol. 7, no. 6, pp. 456– 461, 2016

  4. [4]

    An extensive review on data mining methods and clustering models for intelligent transportation system,

    S. Anand, P. Padmanabham, A. Govardhan, and R. H. Kulkarni, “An extensive review on data mining methods and clustering models for intelligent transportation system,” Journal of Intelligent Systems , vol. 27, no. 2, pp. 263–273, 2018

  5. [5]

    Data mining in psychological treatment research: a primer on classification and regression trees

    M. W. King and P. A. Resick, “Data mining in psychological treatment research: a primer on classification and regression trees.” Journal of consulting and clinical psychology , vol. 82, no. 5, p. 895, 2014

  6. [6]

    General data protection regulation,

    G. GDPR, “General data protection regulation,” Regulation (EU), vol. 679, 2016

  7. [7]

    California consumer privacy act of 2018 (ccpa),

    C. S. Legislature, “California consumer privacy act of 2018 (ccpa),” 2018, accessed: 2024-12-27. [Online]. Available: https: //oag.ca.gov/privacy/ccpa

  8. [8]

    Tabd- dpm: Modelling tabular data with diffusion models,

    A. Kotelnikov, D. Baranchuk, I. Rubachev, and A. Babenko, “Tabd- dpm: Modelling tabular data with diffusion models,” in International Conference on Machine Learning . PMLR, 2023, pp. 17 564–17 579

Show all 154 references
  1. [9]

    Miwae: Deep generative modelling and imputation of incomplete data sets,

    P.-A. Mattei and J. Frellsen, “Miwae: Deep generative modelling and imputation of incomplete data sets,” in International conference on machine learning. PMLR, 2019, pp. 4413–4423

  2. [10]

    A systematic review on imbalanced data challenges in machine learning: Applications and solutions,

    H. Kaur, H. S. Pannu, and A. K. Malhi, “A systematic review on imbalanced data challenges in machine learning: Applications and solutions,” ACM computing surveys (CSUR) , vol. 52, no. 4, pp. 1–36, 2019

  3. [11]

    On oversampling imbalanced data with deep conditional generative models,

    V . A. Fajardo, D. Findlay, C. Jaiswal, X. Yin, R. Houmanfar, H. Xie, J. Liang, X. She, and D. B. Emerson, “On oversampling imbalanced data with deep conditional generative models,” Expert Systems with Applications, vol. 169, p. 114463, 2021

  4. [12]

    Generating synthetic data in finance: opportunities, challenges and pitfalls,

    S. A. Assefa, D. Dervovic, M. Mahfouz, R. E. Tillman, P. Reddy, and M. Veloso, “Generating synthetic data in finance: opportunities, challenges and pitfalls,” in Proceedings of the First ACM International Conference on AI in Finance , 2020, pp. 1–8

  5. [13]

    Synthetic data generation for tabular health records: A systematic review,

    M. Hernandez, G. Epelde, A. Alberdi, R. Cilla, and D. Rankin, “Synthetic data generation for tabular health records: A systematic review,”Neurocomputing, vol. 493, pp. 28–45, 2022

  6. [14]

    Handling missing data with graph representation learning,

    J. You, X. Ma, Y . Ding, M. J. Kochenderfer, and J. Leskovec, “Handling missing data with graph representation learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 19 075–19 087, 2020

  7. [15]

    Gain: Missing data imputation using generative adversarial nets,

    J. Yoon, J. Jordon, and M. Schaar, “Gain: Missing data imputation using generative adversarial nets,” in International conference on machine learning. PMLR, 2018, pp. 5689–5698

  8. [16]

    Tabular and latent space synthetic data generation: a literature review,

    J. Fonseca and F. Bacao, “Tabular and latent space synthetic data generation: a literature review,” Journal of Big Data , vol. 10, no. 1, p. 115, 2023

  9. [17]

    A tutorial on energy-based learning,

    Y . LeCun, S. Chopra, R. Hadsell, M. Ranzato, F. Huang et al. , “A tutorial on energy-based learning,” Predicting structured data , vol. 1, no. 0, 2006

  10. [18]

    Auto-encoding variational bayes,

    D. P. Kingma, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013

  11. [19]

    Generative adversarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014

  12. [20]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,”

  13. [21]

    Available: https://arxiv.org/abs/1706.03762

    [Online]. Available: https://arxiv.org/abs/1706.03762

  14. [22]

    Normalizing flows: An introduction and review of current methods,

    I. Kobyzev, S. J. Prince, and M. A. Brubaker, “Normalizing flows: An introduction and review of current methods,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 11, pp. 3964– 3979, 2020. MANUSCRIPT SUBMITTED TO IEEE FOR POSSIBLE PUBLICATION 21 TA...

  15. [23]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning . PMLR, 2015, pp. 2256–2265

  16. [24]

    Catastrophic forgetting and mode collapse in gans,

    H. Thanh-Tung and T. Tran, “Catastrophic forgetting and mode collapse in gans,” in 2020 international joint conference on neural networks (ijcnn). IEEE, 2020, pp. 1–10

  17. [25]

    Diagnosing and enhancing vae models,

    B. Dai and D. Wipf, “Diagnosing and enhancing vae models,” in International Conference on Learning Representations , 2019

  18. [26]

    Hitchhiker’s guide on energy-based models: a compre- hensive review on the relation with other generative models, sampling and statistical physics,

    D. Carbone, “Hitchhiker’s guide on energy-based models: a compre- hensive review on the relation with other generative models, sampling and statistical physics,” arXiv preprint arXiv:2406.13661 , 2024

  19. [27]

    Limitations of autoregressive models and their alternatives,

    C.-C. Lin, A. Jaech, X. Li, M. R. Gormley, and J. Eisner, “Limitations of autoregressive models and their alternatives,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL- HLT), 2021

  20. [28]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020

  21. [29]

    Score-based generative modeling through stochastic differential equations,

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in International Conference on Learning Rep- resentations

  22. [30]

    Wavegrad: Estimating gradients for waveform generation,

    N. Chen, Y . Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “Wavegrad: Estimating gradients for waveform generation,” in Inter- national Conference on Learning Representations , 2020

  23. [31]

    Diffwave: A versatile diffusion model for audio synthesis,

    Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,” in International Conference on Learning Representations , 2020

  24. [32]

    Argmax flows and multinomial diffusion: Learning categorical distributions,

    E. Hoogeboom, D. Nielsen, P. Jaini, P. Forr ´e, and M. Welling, “Argmax flows and multinomial diffusion: Learning categorical distributions,” Advances in Neural Information Processing Systems , vol. 34, pp. 12 454–12 465, 2021

  25. [33]

    Structured denoising diffusion models in discrete state-spaces,

    J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg, “Structured denoising diffusion models in discrete state-spaces,” Ad- vances in Neural Information Processing Systems , vol. 34, pp. 17 981– 17 993, 2021

  26. [34]

    A survey on video diffusion models,

    Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y .-G. Jiang, “A survey on video diffusion models,”ACM Computing Surveys, vol. 57, no. 2, pp. 1–42, 2024

  27. [35]

    Generative diffusion models on graphs: methods and applications,

    C. Liu, W. Fan, Y . Liu, J. Li, H. Li, H. Liu, J. Tang, and Q. Li, “Generative diffusion models on graphs: methods and applications,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 2023, pp. 6702–6711

  28. [36]

    Stasy: Score-based tabular data synthe- sis,

    J. Kim, C. Lee, and N. Park, “Stasy: Score-based tabular data synthe- sis,” in The Eleventh International Conference on Learning Represen- tations, 2023

  29. [37]

    Autodiff: combining auto-encoder and diffusion model for tabular data synthe- sizing,

    N. Suh, X. Lin, D.-Y . Hsieh, M. Honarkhah, and G. Cheng, “Autodiff: combining auto-encoder and diffusion model for tabular data synthe- sizing,” in NeurIPS 2023 Workshop on Synthetic Data Generation with Generative AI

  30. [38]

    Codi: Co-evolving contrastive diffusion models for mixed-type tabular synthesis,

    C. Lee, J. Kim, and N. Park, “Codi: Co-evolving contrastive diffusion models for mixed-type tabular synthesis,” in International Conference on Machine Learning . PMLR, 2023, pp. 18 940–18 956

  31. [39]

    Mixed-type tabular data synthesis with score-based diffusion in latent space,

    H. Zhang, J. Zhang, Z. Shen, B. Srinivasan, X. Qin, C. Faloutsos, H. Rangwala, and G. Karypis, “Mixed-type tabular data synthesis with score-based diffusion in latent space,” in The Twelfth International Conference on Learning Representations , 2024

  32. [40]

    Generating and imputing tabular data via diffusion and flow-based gradient-boosted trees,

    A. Jolicoeur-Martineau, K. Fatras, and T. Kachman, “Generating and imputing tabular data via diffusion and flow-based gradient-boosted trees,” in International Conference on Artificial Intelligence and Statis- tics. PMLR, 2024, pp. 1288–1296

  33. [41]

    Diffusion models: A comprehensive survey of methods and applications,

    L. Yang, Z. Zhang, Y . Song, S. Hong, R. Xu, Y . Zhao, W. Zhang, B. Cui, and M.-H. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys, vol. 56, no. 4, pp. 1–39, 2023

  34. [42]

    A survey on generative diffusion models,

    H. Cao, C. Tan, Z. Gao, Y . Xu, G. Chen, P.-A. Heng, and S. Z. Li, “A survey on generative diffusion models,” IEEE Transactions on Knowledge and Data Engineering , 2024

  35. [43]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 9, pp. 10 850–10 869, 2023

  36. [44]

    Diffusion models in nlp: A survey,

    Y . Zhu and Y . Zhao, “Diffusion models in nlp: A survey,”arXiv preprint arXiv:2303.07576, 2023

  37. [45]

    Diffusion models for time- MANUSCRIPT SUBMITTED TO IEEE FOR POSSIBLE PUBLICATION 22 series applications: a survey,

    L. Lin, Z. Li, R. Li, X. Li, and J. Gao, “Diffusion models for time- MANUSCRIPT SUBMITTED TO IEEE FOR POSSIBLE PUBLICATION 22 series applications: a survey,” Frontiers of Information Technology & Electronic Engineering, vol. 25, no. 1, pp. 19–41, 2024

  38. [46]

    Challenges and opportunities of generative models on tabular data,

    A. X. Wang, S. S. Chukova, C. R. Simpson, and B. P. Nguyen, “Challenges and opportunities of generative models on tabular data,” Applied Soft Computing , p. 112223, 2024

  39. [47]

    Generative models for tabular data: A review,

    D.-K. Kim, D. Ryu, Y . Lee, and D.-H. Choi, “Generative models for tabular data: A review,”Journal of Mechanical Science and Technology, vol. 38, no. 9, pp. 4989–5005, 2024

  40. [48]

    A comprehensive survey on generative diffusion models for structured data,

    H. Koo and T. E. Kim, “A comprehensive survey on generative diffusion models for structured data,” arXiv e-prints, pp. arXiv–2306, 2023

  41. [49]

    An introduction to variational autoencoders,

    D. P. Kingma, M. Welling et al. , “An introduction to variational autoencoders,”Foundations and Trends® in Machine Learning, vol. 12, no. 4, pp. 307–392, 2019

  42. [50]

    Random variables, joint distribution functions, and copulas,

    A. Sklar, “Random variables, joint distribution functions, and copulas,” Kybernetika, vol. 9, no. 6, pp. 449–460, 1973

  43. [51]

    Gaussian mixture models

    D. A. Reynolds et al. , “Gaussian mixture models.” Encyclopedia of biometrics, vol. 741, no. 659-663, 2009

  44. [52]

    Clinical reasoning over tabular data and text with bayesian networks,

    P. Rabaey, J. Deleu, S. Heytens, and T. Demeester, “Clinical reasoning over tabular data and text with bayesian networks,” in International Conference on Artificial Intelligence in Medicine . Springer, 2024, pp. 229–250

  45. [53]

    Smote: synthetic minority over-sampling technique,

    N. V . Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “Smote: synthetic minority over-sampling technique,” Journal of ar- tificial intelligence research, vol. 16, pp. 321–357, 2002

  46. [54]

    Borderline-smote: a new over- sampling method in imbalanced data sets learning,

    H. Han, W.-Y . Wang, and B.-H. Mao, “Borderline-smote: a new over- sampling method in imbalanced data sets learning,” in International conference on intelligent computing . Springer, 2005, pp. 878–887

  47. [55]

    Synthetic minority oversampling using edited displacement-based k-nearest neighbors,

    A. X. Wang, S. S. Chukova, and B. P. Nguyen, “Synthetic minority oversampling using edited displacement-based k-nearest neighbors,” Applied Soft Computing , vol. 148, p. 110895, 2023

  48. [56]

    Smote-enc: A novel smote-based method to generate synthetic data for nominal and continuous features,

    M. Mukherjee and M. Khushi, “Smote-enc: A novel smote-based method to generate synthetic data for nominal and continuous features,” Applied system innovation , vol. 4, no. 1, p. 18, 2021

  49. [57]

    Adasyn: Adaptive synthetic sampling approach for imbalanced learning,

    H. He, Y . Bai, E. A. Garcia, and S. Li, “Adasyn: Adaptive synthetic sampling approach for imbalanced learning,” in 2008 IEEE interna- tional joint conference on neural networks (IEEE world congress on computational intelligence). Ieee, 2008, pp. 1322–1328

  50. [58]

    synthpop: Bespoke creation of synthetic data in r,

    B. Nowok, G. M. Raab, and C. Dibben, “synthpop: Bespoke creation of synthetic data in r,” Journal of statistical software, vol. 74, pp. 1–26, 2016

  51. [59]

    Modeling tabular data using conditional gan,

    L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, “Modeling tabular data using conditional gan,” Advances in neural information processing systems , vol. 32, 2019

  52. [60]

    Goggle: Generative modelling for tabular data by learning relational structure,

    T. Liu, Z. Qian, J. Berrevoets, and M. van der Schaar, “Goggle: Generative modelling for tabular data by learning relational structure,” in The Eleventh International Conference on Learning Representations, 2023

  53. [61]

    Ctab-gan: Effective table data synthesizing,

    Z. Zhao, A. Kunar, R. Birke, and L. Y . Chen, “Ctab-gan: Effective table data synthesizing,” in Asian Conference on Machine Learning . PMLR, 2021, pp. 97–112

  54. [62]

    Ctab- gan+: Enhancing tabular data synthesis,

    Z. Zhao, A. Kunar, R. Birke, H. Van der Scheer, and L. Y . Chen, “Ctab- gan+: Enhancing tabular data synthesis,” Frontiers in big Data, vol. 6, p. 1296508, 2024

  55. [63]

    Large language models: A survey,

    S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Ama- triain, and J. Gao, “Large language models: A survey,” arXiv preprint arXiv:2402.06196, 2024

  56. [64]

    Language models are realistic tabular data generators,

    V . Borisov, K. Sessler, T. Leemann, M. Pawelczyk, and G. Kasneci, “Language models are realistic tabular data generators,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=cEygmQNOeI

  57. [65]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  58. [66]

    Diffusion models beat gans on image synthesis,

    P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Advances in neural information processing systems, vol. 34, pp. 8780–8794, 2021

  59. [67]

    Sos: Score-based oversampling for tabular data,

    J. Kim, C. Lee, Y . Shin, S. Park, M. Kim, N. Park, and J. Cho, “Sos: Score-based oversampling for tabular data,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 762–772

  60. [68]

    Large language models (LLMs) on tabular data: Prediction, generation, and understanding - a survey,

    X. Fang, W. Xu, F. A. Tan, Z. Hu, J. Zhang, Y . Qi, S. H. Sengamedu, and C. Faloutsos, “Large language models (LLMs) on tabular data: Prediction, generation, and understanding - a survey,”Transactions on Machine Learning Research , 2024. [Online]. Available: https://openreview...

  61. [69]

    Diffusion models for missing value imputation in tabular data,

    S. Zheng and N. Charoenphakdee, “Diffusion models for missing value imputation in tabular data,” inNeurIPS 2022 First Table Representation Workshop

  62. [70]

    What do we really know about wages? the importance of nonreporting and census imputation,

    L. Lillard, J. P. Smith, and F. Welch, “What do we really know about wages? the importance of nonreporting and census imputation,”Journal of Political Economy, vol. 94, no. 3, Part 1, pp. 489–506, 1986

  63. [71]

    Strategies for handling missing data in electronic health record derived data,

    B. J. Wells, K. M. Chagin, A. S. Nowacki, and M. W. Kattan, “Strategies for handling missing data in electronic health record derived data,” Egems, vol. 1, no. 3, 2013

  64. [72]

    A survey on missing data in machine learning,

    T. Emmanuel, T. Maupong, D. Mpoeleng, T. Semong, B. Mphago, and O. Tabona, “A survey on missing data in machine learning,” Journal of Big data , vol. 8, pp. 1–37, 2021

  65. [73]

    Inference and missing data,

    D. B. Rubin, “Inference and missing data,” Biometrika, vol. 63, no. 3, pp. 581–592, 1976

  66. [74]

    Tabdiff: a unified diffusion model for multi-modal tabular data generation,

    J. Shi, M. Xu, H. Hua, H. Zhang, S. Ermon, and J. Leskovec, “Tabdiff: a unified diffusion model for multi-modal tabular data generation,” in NeurIPS 2024 Third Table Representation Learning Workshop

  67. [75]

    Generative modeling by estimating gradients of the data distribution,

    Y . Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Advances in neural information processing systems, vol. 32, 2019

  68. [76]

    P. E. Kloeden, E. Platen, P. E. Kloeden, and E. Platen, Stochastic differential equations. Springer, 1992

  69. [77]

    Neural ordinary differential equations,

    R. T. Chen, Y . Rubanova, J. Bettencourt, and D. K. Duvenaud, “Neural ordinary differential equations,” Advances in neural information pro- cessing systems, vol. 31, 2018

  70. [78]

    Classifier-free diffusion guidance,

    J. Ho and T. Salimans, “Classifier-free diffusion guidance,” in NeurIPS 2021 Workshop on Deep Generative Models and Downstream Appli- cations, 2021

  71. [79]

    Tabular data aug- mentation for machine learning: Progress and prospects of embracing generative ai,

    L. Cui, H. Li, K. Chen, L. Shou, and G. Chen, “Tabular data aug- mentation for machine learning: Progress and prospects of embracing generative ai,” arXiv preprint arXiv:2407.21523 , 2024

  72. [80]

    Missdiff: Training diffusion models on tabular data with missing values,

    Y . Ouyang, L. Xie, C. Li, and G. Cheng, “Missdiff: Training diffusion models on tabular data with missing values,” in ICML 2023 Workshop on Structured Probabilistic Inference {\&} Generative Modeling , 2023

  73. [81]

    Synthetic health-related lon- gitudinal data with mixed-type variables generated using diffusion models,

    I. Nicholas, H. Kuo, F. Garcia, A. Sonnerborg, M. Bohm, R. Kaiser, M. Zazzi, L. Jorm, and S. Barbieri, “Synthetic health-related lon- gitudinal data with mixed-type variables generated using diffusion models,” in NeurIPS 2023 Workshop on Synthetic Data Generation with Generati...

  74. [82]

    Findiff: Diffusion models for financial tabular data generation,

    T. Sattarov, M. Schreyer, and D. Borth, “Findiff: Diffusion models for financial tabular data generation,” in Proceedings of the Fourth ACM International Conference on AI in Finance , 2023, pp. 64–72

  75. [83]

    Meddiff: Generating electronic health records using accelerated denoising diffusion model,

    H. He, S. Zhao, Y . Xi, and J. C. Ho, “Meddiff: Generating electronic health records using accelerated denoising diffusion model,” arXiv preprint arXiv:2302.04355, 2023

  76. [84]

    Synthesizing mixed-type electronic health records using diffusion models,

    T. Ceritli, G. O. Ghosheh, V . K. Chauhan, T. Zhu, A. P. Creagh, and D. A. Clifton, “Synthesizing mixed-type electronic health records using diffusion models,” arXiv preprint arXiv:2302.14679 , 2023

  77. [85]

    A flexible generative model for heterogeneous tabular ehr with missing modality,

    H. He, Y . Xi, Y . Chen, B. Malin, J. Ho et al. , “A flexible generative model for heterogeneous tabular ehr with missing modality,” in The Twelfth International Conference on Learning Representations , 2024

  78. [86]

    Ehrdiff: Exploring realistic ehr synthesis with diffusion models,

    H. Yuan, S. Zhou, and S. Yu, “Ehrdiff: Exploring realistic ehr synthesis with diffusion models,” Transactions on Machine Learning Research , 2024

  79. [87]

    Entity-based financial tabular data synthesis with diffusion models,

    C. Liu and C. Liu, “Entity-based financial tabular data synthesis with diffusion models,” in Proceedings of the 5th ACM International Conference on AI in Finance , 2024, pp. 547–554

  80. [88]

    Imb-findiff: Conditional diffusion models for class imbalance synthesis of financial tabular data,

    M. Schreyer, T. Sattarov, A. Sim, and K. Wu, “Imb-findiff: Conditional diffusion models for class imbalance synthesis of financial tabular data,” in Proceedings of the 5th ACM International Conference on AI in Finance, 2024, pp. 617–625

  81. [89]

    Guided discrete diffusion for electronic health record generation,

    J. Han, Z. Chen, Y . Li, Y . Kou, E. Halperin, R. E. Tillman, and Q. Gu, “Guided discrete diffusion for electronic health record generation,” arXiv preprint arXiv:2404.12314 , 2024

  82. [90]

    Tabunite: Efficient encoding schemes for flow and diffusion tabular generative models,

    J. Si, Z. Ou, M. Qu, and Y . Li, “Tabunite: Efficient encoding schemes for flow and diffusion tabular generative models,” 2024. [Online]. Available: https://openreview.net/forum?id=Zoli4UAQVZ

  83. [91]

    Continuous diffusion for mixed-type tabular data,

    M. Mueller, K. Gruber, and D. Fok, “Continuous diffusion for mixed-type tabular data,” 2024. [Online]. Available: https: //arxiv.org/abs/2312.10431

  84. [92]

    Extracting training data from diffusion models,

    N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V . Sehwag, F. Tramer, B. Balle, D. Ippolito, and E. Wallace, “Extracting training data from diffusion models,” in 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 5253–5270. MANUSCRIPT SUBMITTED TO IEEE FOR POSSIBLE P...

  85. [93]

    Repaint: Inpainting using denoising diffusion probabilis- tic models,

    A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “Repaint: Inpainting using denoising diffusion probabilis- tic models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 461–11 471

  86. [94]

    Multilayer feedforward networks are universal approximators,

    K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks, vol. 2, no. 5, pp. 359–366, 1989

  87. [95]

    Xgboost: A scalable tree boosting system,

    T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining , 2016, pp. 785–794

  88. [96]

    Encoding categorical data: Is there yet anything’hotter’than one-hot encoding?

    E. Poslavskaya and A. Korolev, “Encoding categorical data: Is there yet anything’hotter’than one-hot encoding?” arXiv preprint arXiv:2312.16930, 2023

  89. [97]

    On the challenges of learning with inference networks on sparse, high-dimensional data,

    R. Krishnan, D. Liang, and M. Hoffman, “On the challenges of learning with inference networks on sparse, high-dimensional data,” in International conference on artificial intelligence and statistics . PMLR, 2018, pp. 143–151

  90. [98]

    Analog bits: Generating discrete data using diffusion models with self-conditioning,

    T. Chen, R. ZHANG, and G. Hinton, “Analog bits: Generating discrete data using diffusion models with self-conditioning,” in The Eleventh International Conference on Learning Representations

  91. [99]

    Vaem: a deep generative model for heterogeneous mixed type data,

    C. Ma, S. Tschiatschek, R. Turner, J. M. Hern ´andez-Lobato, and C. Zhang, “Vaem: a deep generative model for heterogeneous mixed type data,” Advances in Neural Information Processing Systems , vol. 33, pp. 11 237–11 247, 2020

  92. [100]

    Estimation of non-normalized statistical models by score matching

    A. Hyv ¨arinen and P. Dayan, “Estimation of non-normalized statistical models by score matching.” Journal of Machine Learning Research , vol. 6, no. 4, 2005

  93. [101]

    Contin- uous diffusion for categorical data,

    S. Dieleman, L. Sartran, A. Roshannai, N. Savinov, Y . Ganin, P. H. Richemond, A. Doucet, R. Strudel, C. Dyer, C. Durkan et al., “Contin- uous diffusion for categorical data,” arXiv preprint arXiv:2211.15089 , 2022

  94. [102]

    Mining electronic health records (ehrs) a survey,

    P. Yadav, M. Steinbach, V . Kumar, and G. Simon, “Mining electronic health records (ehrs) a survey,” ACM Computing Surveys (CSUR) , vol. 50, no. 6, pp. 1–40, 2018

  95. [104]

    Iterative procedures for nonlinear integral equations,

    D. G. Anderson, “Iterative procedures for nonlinear integral equations,” Journal of the ACM (JACM) , vol. 12, no. 4, pp. 547–560, 1965

  96. [105]

    Elucidating the design space of diffusion-based generative models,

    T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” Advances in neural infor- mation processing systems , vol. 35, pp. 26 565–26 577, 2022

  97. [106]

    Clavaddpm: Multi- relational data synthesis with cluster-guided diffusion models,

    W. Pang, M. Shafieinejad, L. Liu, and X. He, “Clavaddpm: Multi- relational data synthesis with cluster-guided diffusion models,” Ad- vances in Neural Information Processing Systems , 2024

  98. [107]

    Relational data generation with graph neural networks and latent diffusion models,

    V . Hudovernik, “Relational data generation with graph neural networks and latent diffusion models,” in NeurIPS 2024 Third Table Represen- tation Learning Workshop, 2024

  99. [108]

    Benchmarking the fidelity and utility of synthetic relational data,

    V . Hudovernik, M. Jurkovi ˇc, and E. ˇStrumbelj, “Benchmarking the fidelity and utility of synthetic relational data,” arXiv preprint arXiv:2410.03411, 2024

  100. [109]

    Missing value imputation: a review and analysis of the literature (2006–2017),

    W.-C. Lin and C.-F. Tsai, “Missing value imputation: a review and analysis of the literature (2006–2017),” Artificial Intelligence Review , vol. 53, pp. 1487–1509, 2020

  101. [110]

    Hyperimpute: Generalized iterative imputation with automatic model selection,

    D. Jarrett, B. C. Cebere, T. Liu, A. Curth, and M. van der Schaar, “Hyperimpute: Generalized iterative imputation with automatic model selection,” in International Conference on Machine Learning. PMLR, 2022, pp. 9916–9937

  102. [111]

    Multivariate imputation by chained equations,

    S. Van Buuren and C. G. Oudshoorn, “Multivariate imputation by chained equations,” 2000

  103. [112]

    Mida: Multiple imputation using denoising autoencoders,

    L. Gondara and K. Wang, “Mida: Multiple imputation using denoising autoencoders,” in Advances in Knowledge Discovery and Data Min- ing: 22nd Pacific-Asia Conference, PAKDD 2018, Melbourne, VIC, Australia, June 3-6, 2018, Proceedings, Part III 22 . Springer, 2018, pp. 260–272

  104. [113]

    Handling incomplete heterogeneous data using vaes,

    A. Nazabal, P. M. Olmos, Z. Ghahramani, and I. Valera, “Handling incomplete heterogeneous data using vaes,” Pattern Recognition, vol. 107, p. 107501, 2020

  105. [114]

    Self-supervision im- proves diffusion models for tabular data imputation,

    Y . Liu, T. Ajanthan, H. Husain, and V . Nguyen, “Self-supervision im- proves diffusion models for tabular data imputation,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024, pp. 1513–1522

  106. [115]

    Diffusion models for tabular data imputation and synthetic data generation,

    M. Villaiz ´an-Vallelado, M. Salvatori, C. Segura, and I. Arapakis, “Diffusion models for tabular data imputation and synthetic data generation,” arXiv preprint arXiv:2407.02549 , 2024

  107. [116]

    Natural generative noise diffusion model imputation,

    A. Wibisono, P. Mursanto, S. See et al. , “Natural generative noise diffusion model imputation,” Knowledge-Based Systems , vol. 301, p. 112310, 2024

  108. [117]

    Rethinking the diffusion models for missing data imputation: A gradient flow perspective,

    Z. Chen, H. Li, F. Wang, O. Zhang, H. Xu, X. Jiang, Z. Song, and H. Wang, “Rethinking the diffusion models for missing data imputation: A gradient flow perspective,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024

  109. [118]

    Unleashing the potential of diffusion models for incomplete data imputation,

    H. Zhang, L. Fang, and P. S. Yu, “Unleashing the potential of diffusion models for incomplete data imputation,” 2024. [Online]. Available: https://arxiv.org/abs/2405.20690

  110. [119]

    Csdi: Conditional score-based diffusion models for probabilistic time series imputation,

    Y . Tashiro, J. Song, Y . Song, and S. Ermon, “Csdi: Conditional score-based diffusion models for probabilistic time series imputation,” Advances in Neural Information Processing Systems , vol. 34, pp. 24 804–24 816, 2021

  111. [120]

    Revisiting deep learning models for tabular data,

    Y . Gorishniy, I. Rubachev, V . Khrulkov, and A. Babenko, “Revisiting deep learning models for tabular data,” Advances in Neural Information Processing Systems, vol. 34, pp. 18 932–18 943, 2021

  112. [121]

    Large-scale wasserstein gradient flows,

    P. Mokrov, A. Korotin, L. Li, A. Genevay, J. M. Solomon, and E. Burnaev, “Large-scale wasserstein gradient flows,” Advances in Neural Information Processing Systems , vol. 34, pp. 15 243–15 256, 2021

  113. [122]

    Maximum likelihood from incomplete data via the em algorithm,

    A. P. Dempster, N. M. Laird, and D. B. Rubin, “Maximum likelihood from incomplete data via the em algorithm,” Journal of the royal statistical society: series B (methodological) , vol. 39, no. 1, pp. 1–22, 1977

  114. [123]

    SiloFuse: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models ,

    A. Shankar, H. Brouwer, R. Hai, and L. Chen, “ SiloFuse: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models ,” in 2024 IEEE 40th International Conference on Data Engineering (ICDE) . Los Alamitos, CA, USA: IEEE Computer Society, May 2024, pp. 110–123. [O...

  115. [124]

    Fedtabdiff: Federated learning of diffusion probabilistic models for synthetic mixed-type tabular data generation,

    T. Sattarov, M. Schreyer, and D. Borth, “Fedtabdiff: Federated learning of diffusion probabilistic models for synthetic mixed-type tabular data generation,” arXiv preprint arXiv:2401.06263 , 2024

  116. [125]

    Balanced mixed- type tabular data synthesis with diffusion models,

    Z. Yang, P. Guo, K. Zanna, and A. Sano, “Balanced mixed- type tabular data synthesis with diffusion models,” arXiv preprint arXiv:2404.08254, 2024

  117. [126]

    Differentially private federated learning of diffusion models for synthetic tabular data generation,

    T. Sattarov, M. Schreyer, and D. Borth, “Differentially private federated learning of diffusion models for synthetic tabular data generation,” arXiv preprint arXiv:2412.16083 , 2024

  118. [127]

    Federated learning: Collaborative machine learning without centralized training data,

    B. McMahan and D. Ramage, “Federated learning: Collaborative machine learning without centralized training data,” Google Research Blog, vol. 3, 2017

  119. [128]

    The algorithmic foundations of differential privacy,

    C. Dwork, A. Roth et al., “The algorithmic foundations of differential privacy,”Foundations and Trends® in Theoretical Computer Science , vol. 9, no. 3–4, pp. 211–407, 2014

  120. [129]

    Tabadm: Unsupervised tabular anomaly detection with diffusion models,

    G. Zamberg, M. Salhov, O. Lindenbaum, and A. Averbuch, “Tabadm: Unsupervised tabular anomaly detection with diffusion models,” arXiv preprint arXiv:2307.12336, 2023

  121. [130]

    On diffusion modeling for anomaly detection,

    V . Livernoche, V . Jain, Y . Hezaveh, and S. Ravanbakhsh, “On diffusion modeling for anomaly detection,” in The Twelfth International Confer- ence on Learning Representations , 2024

  122. [131]

    Self-supervised enhanced denoising diffusion for anomaly detection,

    S. Li, J. Yu, Y . Lu, G. Yang, X. Du, and S. Liu, “Self-supervised enhanced denoising diffusion for anomaly detection,” Information Sciences, vol. 669, p. 120612, 2024

  123. [132]

    Anomaly detection by estimating gradients of the tabular data distribution,

    Anonymous, “Anomaly detection by estimating gradients of the tabular data distribution,” in Submitted to The Thirteenth International Conference on Learning Representations, 2024, under review. [Online]. Available: https://openreview.net/forum?id=7QDIFrtAsB

  124. [133]

    Frauddiffuse: Diffusion-aided synthetic fraud augmentation for improved fraud detection,

    R. Roy, D. Tiwari, and A. Pandey, “Frauddiffuse: Diffusion-aided synthetic fraud augmentation for improved fraud detection,” in Pro- ceedings of the 5th ACM International Conference on AI in Finance , 2024, pp. 90–98

  125. [134]

    Synthetic data generation for fraud detection using diffusion models,

    Y . Pushkarenko and V . Zaslavskyi, “Synthetic data generation for fraud detection using diffusion models,” Information & Security: An International Journal , vol. 55, no. 2, pp. 185–198, 2024. [Online]. Available: https://doi.org/10.11610/isij.5534

  126. [135]

    Simple and effective masked diffusion language models,

    S. S. Sahoo, M. Arriola, Y . Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V . Kuleshov, “Simple and effective masked diffusion language models,” arXiv preprint arXiv:2406.07524 , 2024

  127. [136]

    Likelihood-based diffusion language models,

    I. Gulrajani and T. B. Hashimoto, “Likelihood-based diffusion language models,” Advances in Neural Information Processing Systems , vol. 36, 2024

  128. [137]

    Ssd-lm: Semi-autoregressive simplex-based diffusion language model for text generation and mod- ular control,

    X. Han, S. Kumar, and Y . Tsvetkov, “Ssd-lm: Semi-autoregressive simplex-based diffusion language model for text generation and mod- ular control,” in The 61st Annual Meeting Of The Association For Computational Linguistics, 2023

  129. [138]

    Diffusion-lm improves controllable text generation,

    X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto, “Diffusion-lm improves controllable text generation,” Advances in Neural Information Processing Systems, vol. 35, pp. 4328–4343, 2022. MANUSCRIPT SUBMITTED TO IEEE FOR POSSIBLE PUBLICATION 24

  130. [139]

    Latent diffusion for language generation,

    J. Lovelace, V . Kishore, C. Wan, E. Shekhtman, and K. Q. Weinberger, “Latent diffusion for language generation,” Advances in Neural Infor- mation Processing Systems , vol. 36, 2024

  131. [140]

    Self- conditioned embedding diffusion for text generation,

    R. Strudel, C. Tallec, F. Altch ´e, Y . Du, Y . Ganin, A. Mensch, W. Grathwohl, N. Savinov, S. Dieleman, L. Sifre et al. , “Self- conditioned embedding diffusion for text generation,” arXiv preprint arXiv:2211.04236, 2022

  132. [141]

    Diffusing gaussian mixtures for generating categorical data,

    F. Regol and M. Coates, “Diffusing gaussian mixtures for generating categorical data,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 8, 2023, pp. 9570–9578

  133. [142]

    Denoising diffusion implicit mod- els,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit mod- els,” in International Conference on Learning Representations , 2021

  134. [143]

    Simplified and generalized masked diffusion for discrete data,

    J. Shi, K. Han, Z. Wang, A. Doucet, and M. K. Titsias, “Simplified and generalized masked diffusion for discrete data,” arXiv preprint arXiv:2406.04329, 2024

  135. [144]

    Autoregressive diffusion models,

    E. Hoogeboom, A. A. Gritsenko, J. Bastings, B. Poole, R. van den Berg, and T. Salimans, “Autoregressive diffusion models,” in International Conference on Learning Representations , 2021

  136. [145]

    Diffusion language mod- els can perform many tasks with scaling and instruction-finetuning,

    J. Ye, Z. Zheng, Y . Bao, L. Qian, and Q. Gu, “Diffusion language mod- els can perform many tasks with scaling and instruction-finetuning,” arXiv preprint arXiv:2308.12219 , 2023

  137. [146]

    Gotta go fast when generating data with score-based models,

    A. Jolicoeur-Martineau, K. Li, R. Pich ´e-Taillefer, T. Kachman, and I. Mitliagkas, “Gotta go fast when generating data with score-based models,” arXiv preprint arXiv:2105.14080 , 2021

  138. [147]

    Diffuser: Discrete diffu- sion via edit-based reconstruction,

    M. Reid, V . J. Hellendoorn, and G. Neubig, “Diffuser: Discrete diffu- sion via edit-based reconstruction,” arXiv preprint arXiv:2210.16886 , 2022

  139. [148]

    A continuous time framework for discrete denoising models,

    A. Campbell, J. Benton, V . De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet, “A continuous time framework for discrete denoising models,” Advances in Neural Information Processing Systems , vol. 35, pp. 28 266–28 279, 2022

  140. [149]

    Score-based continuous-time discrete diffusion models,

    H. Sun, L. Yu, B. Dai, D. Schuurmans, and H. Dai, “Score-based continuous-time discrete diffusion models,” in The Eleventh Interna- tional Conference on Learning Representations , 2023

  141. [150]

    Fast sampling via de-randomization for discrete diffusion models,

    Z. Chen, H. Yuan, Y . Li, Y . Kou, J. Zhang, and Q. Gu, “Fast sampling via de-randomization for discrete diffusion models,” 2024. [Online]. Available: https://openreview.net/forum?id=m4Ya9RkEEW

  142. [151]

    How faithful is your synthetic data? sample-level metrics for evaluating and auditing generative models,

    A. Alaa, B. Van Breugel, E. S. Saveliev, and M. van der Schaar, “How faithful is your synthetic data? sample-level metrics for evaluating and auditing generative models,” in International Conference on Machine Learning. PMLR, 2022, pp. 290–306

  143. [152]

    Gen- erating multi-label discrete patient records using generative adversarial networks,

    E. Choi, S. Biswal, B. Malin, J. Duke, W. F. Stewart, and J. Sun, “Gen- erating multi-label discrete patient records using generative adversarial networks,” in Machine learning for healthcare conference . PMLR, 2017, pp. 286–305

  144. [153]

    {AttriGuard}: A practical defense against attribute inference attacks via adversarial machine learning,

    J. Jia and N. Z. Gong, “ {AttriGuard}: A practical defense against attribute inference attacks via adversarial machine learning,” in 27th USENIX Security Symposium (USENIX Security 18) , 2018, pp. 513– 529

  145. [154]

    Membership inference attacks against machine learning models,

    R. Shokri, M. Stronati, C. Song, and V . Shmatikov, “Membership inference attacks against machine learning models,” in 2017 IEEE symposium on security and privacy (SP) . IEEE, 2017, pp. 3–18

  146. [155]

    Adbench: Anomaly detection benchmark,

    S. Han, X. Hu, H. Huang, M. Jiang, and Y . Zhao, “Adbench: Anomaly detection benchmark,” in Thirty-Sixth Conference on Neural Informa- tion Processing Systems Datasets and Benchmarks Track , 2022

Pith tools

Reviewed May 25, 2026 · model on record in the stance chip above.