Pith. sign in

REVIEW 18 cited by

Machine Learning for Synthetic Data Generation: A Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.04062 v10 pith:3RSVITZG submitted 2023-02-08 cs.LG

classification cs.LG
keywords datasyntheticgenerationlearningmachinemodelsreviewapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning heavily relies on data, but real-world applications often encounter various data-related issues. These include data of poor quality, insufficient data points leading to under-fitting of machine learning models, and difficulties in data access due to concerns surrounding privacy, safety, and regulations. In light of these challenges, the concept of synthetic data generation emerges as a promising alternative that allows for data sharing and utilization in ways that real-world data cannot facilitate. This paper presents a comprehensive systematic review of existing studies that employ machine learning models for the purpose of generating synthetic data. The review encompasses various perspectives, starting with the applications of synthetic data generation, spanning computer vision, speech, natural language processing, healthcare, and business domains. Additionally, it explores different machine learning methods, with particular emphasis on neural network architectures and deep generative models. The paper also addresses the crucial aspects of privacy and fairness concerns related to synthetic data generation. Furthermore, this study identifies the challenges and opportunities prevalent in this emerging field, shedding light on the potential avenues for future research. By delving into the intricacies of synthetic data generation, this paper aims to contribute to the advancement of knowledge and inspire further exploration in synthetic data generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?

    stat.ML 2026-06 unverdicted novelty 7.0 of 10

    Synthetic minority augmentation improves threshold-integrated and optimized classification metrics only under model misspecification by correcting ranking errors, while providing no fundamental benefit beyond possible...

  2. When Does Model Collapse Occur in Structured Interactive Learning?

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Model collapse occurs in structured interactive learning if and only if the directed interaction graph satisfies a specific topological condition, with finite-sample guarantees for linear regression and asymptotic res...

  3. Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    TabGRAA enables self-improving tabular language models through iterative group-relative advantage alignment using modular automated quality signals like distinguishability classifiers.

  4. HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop Questions

    cs.CL 2025-05 unverdicted novelty 7.0 of 10

    HopWeaver automatically synthesizes authentic bridge and comparison multi-hop questions from cross-document sources via a pipeline that identifies complementary documents and builds reasoning paths.

  5. SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SADGE is a new fused similarity metric combining DINOv3 appearance and MASt3R geometry via constrained bilinear interaction that correlates with downstream synthetic-to-real performance at Pearson r=0.88 across multip...

  6. What Makes Synthetic Data Effective in Image Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Dense scene composition and instance fidelity in synthetic diffusion images drive better segmentation performance; SENSE framework exploits this to improve models on Cityscapes, COCO, and ADE20K.

  7. Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

    cs.SE 2026-05 unverdicted novelty 6.0 of 10

    FireFly inverts task synthesis by exploring real MCP servers first via pairwise tool graphs and sub-DAG sampling, then generates 5,144 verified tasks backward from outcomes to train a 4B model that matches Claude Sonn...

  8. Generative AI-Based Monte Carlo Simulation for Method Evaluation Using Synthetic Multilevel Data

    stat.ME 2026-05 unverdicted novelty 6.0 of 10

    A framework using generative AI to produce synthetic multilevel data for Monte Carlo simulations that evaluate the performance and parameter recovery of quantitative methods.

  9. FUTURE: Flexible Unlearning for Tree Ensemble

    cs.LG 2025-08 conditional novelty 6.0 of 10

    FUTURE forgets training samples from tree ensembles by optimizing sigmoid-smoothed split thresholds and copying them back to the original discrete trees.

  10. Synthetic CVs To Build and Test Fairness-Aware Hiring Tools

    cs.CY 2025-08 conditional novelty 6.0 of 10

    A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.

  11. PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A DSL plus SMT solver generates and validates 83,657 logic puzzles, and fine-tuning on them improves a 7B model's scores on several reasoning benchmarks.

  12. Stochastic dynamics learning with state-space systems

    stat.ML 2025-08 unverdicted novelty 6.0 of 10

    Establishes that fading memory and solution stability hold generically in state-space systems for reservoir computing even without the echo state property, with a distributional attractor perspective for stochastic cases.

  13. CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    CoX-MoE achieves up to 7.1x higher throughput than FlexGen for MoE inference via coalesced expert execution and AMX-enabled CPU-GPU orchestration with static expert stratification.

  14. Fundamental Trade-Offs in Multi-Bit Watermarking of Stochastic Processes

    cs.IT 2026-05 unverdicted novelty 5.0 of 10

    Derives matched converse and achievability bounds that characterize optimal trade-offs among false-alarm probability, detection error probability, distortion, and information rate for multi-bit watermarking of station...

  15. Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    TabGRAA applies group-relative advantage alignment in an iterative reward-guided post-training loop to improve tabular language model generators on fidelity, utility, and privacy trade-offs across five benchmarks.

  16. Can Synthetic Data be Fair and Private? A Comparative Study of Synthetic Data Generation and Fairness Algorithms

    cs.LG 2025-01 unverdicted novelty 5.0 of 10

    DECAF synthetic data generator best balances privacy and fairness while fairness pre-processing improves outcomes more on synthetic data than real data, though at some cost to predictive accuracy.

  17. Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records

    cs.CL 2025-09 conditional novelty 3.0 of 10

    Synthetic patient-doctor dialogues generated by ChatGPT-4o and Gemini and mixed with 20,000 real Arabic records improved fine-tuned LLM BERTScore F1 scores, with ChatGPT-4o data giving larger gains than Gemini data.

  18. Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A survey that categorizes tabular data synthesis by generation objectives and adds a benchmark comparison of six models on Adult and CreditRisk.

Pith tools