Pith. sign in

REVIEW 5 cited by

Synthcity: facilitating innovative use cases of synthetic data in different data modalities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.07573 v1 pith:KPLV5II2 submitted 2023-01-18 cs.LG cs.AI

classification cs.LGcs.AI
keywords datasynthcitysyntheticcasescommunitygithubhttpsinnovative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthcity is an open-source software package for innovative use cases of synthetic data in ML fairness, privacy and augmentation across diverse tabular data modalities, including static data, regular and irregular time series, data with censoring, multi-source data, composite data, and more. Synthcity provides the practitioners with a single access point to cutting edge research and tools in synthetic data. It also offers the community a playground for rapid experimentation and prototyping, a one-stop-shop for SOTA benchmarks, and an opportunity for extending research impact. The library can be accessed on GitHub (https://github.com/vanderschaarlab/synthcity) and pip (https://pypi.org/project/synthcity/). We warmly invite the community to join the development effort by providing feedback, reporting bugs, and contributing code.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 26 citations worldwide. Full citation record

  1. SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

    cs.CR 2025-06 conditional novelty 7.0 of 10

    A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...

  2. Large Language Models for Imbalanced Classification: Diversity makes the difference

    cs.LG 2025-10 conditional novelty 6.0 of 10

    ImbLLM improves LLM-based oversampling for imbalanced tabular data by conditioning prompts on label plus features, fixing the label at the start of fine-tuning, and interpolating minority samples.

  3. TAGAL: Tabular Data Generation using Agentic LLM Methods

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TAGAL uses an agentic LLM loop, generation plus feedback, to produce synthetic tabular data without LLM training, matching trained models on some datasets and beating the training-free EPIC baseline.

  4. Ensembling Membership Inference Attacks Against Tabular Generative Models

    cs.CR 2025-09 conditional novelty 6.0 of 10

    No single membership inference attack dominates across tabular generative models, and unsupervised ensembles of attacks achieve better average rankings.

  5. A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation

    cs.CR 2025-07 reject novelty 3.0 of 10

    A review that lists 17 privacy metrics for synthetic data with their formulas and assumptions, rescaling each to a common privacy-risk direction.

Pith tools