REVIEW 5 cited by
Synthcity: facilitating innovative use cases of synthetic data in different data modalities
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Synthcity is an open-source software package for innovative use cases of synthetic data in ML fairness, privacy and augmentation across diverse tabular data modalities, including static data, regular and irregular time series, data with censoring, multi-source data, composite data, and more. Synthcity provides the practitioners with a single access point to cutting edge research and tools in synthetic data. It also offers the community a playground for rapid experimentation and prototyping, a one-stop-shop for SOTA benchmarks, and an opportunity for extending research impact. The library can be accessed on GitHub (https://github.com/vanderschaarlab/synthcity) and pip (https://pypi.org/project/synthcity/). We warmly invite the community to join the development effort by providing feedback, reporting bugs, and contributing code.
Forward citations
Cited by 5 Pith papers
-
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...
-
Large Language Models for Imbalanced Classification: Diversity makes the difference
ImbLLM improves LLM-based oversampling for imbalanced tabular data by conditioning prompts on label plus features, fixing the label at the start of fine-tuning, and interpolating minority samples.
-
TAGAL: Tabular Data Generation using Agentic LLM Methods
TAGAL uses an agentic LLM loop, generation plus feedback, to produce synthetic tabular data without LLM training, matching trained models on some datasets and beating the training-free EPIC baseline.
-
Ensembling Membership Inference Attacks Against Tabular Generative Models
No single membership inference attack dominates across tabular generative models, and unsupervised ensembles of attacks achieve better average rankings.
-
A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation
A review that lists 17 privacy metrics for synthetic data with their formulas and assumptions, rescaling each to a common privacy-risk direction.
Discussion (0). Sign in to comment.