Pith. sign in

REVIEW 2 cited by

Synthetic Data for Social Good

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.08874 v1 pith:OCJLFFGC submitted 2017-10-24 cs.CY

classification cs.CY
keywords datadatasetsyntheticdatasynthesizergenerationgoodprivacysharing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data for good implies unfettered access to data. But data owners must be conservative about how, when, and why they share data or risk violating the trust of the people they aim to help, losing their funding, or breaking the law. Data sharing agreements can help prevent privacy violations, but require a level of specificity that is premature during preliminary discussions, and can take over a year to establish. We consider the generation and use of synthetic data to facilitate ad hoc collaborations involving sensitive data. A good synthetic dataset has two properties: it is representative of the original data, and it provides strong guarantees about privacy. In this paper, we discuss important use cases for synthetic data that challenge the state of the art in privacy-preserving data generation, and describe DataSynthesizer, a dataset generation tool that takes a sensitive dataset as input and generates a structurally and statistically similar synthetic dataset, with strong privacy guarantees, as output. The data owners need not release their data, while potential collaborators can begin developing models and methods with some confidence that their results will work similarly on the real dataset. The distinguishing feature of DataSynthesizer is its usability - in most cases, the data owner need not specify any parameters to start generating and sharing data safely and effectively. The code implementing DataSynthesizer is publicly available on GitHub at https://github.com/DataResponsibly. The work on DataSynthesizer is part of the Data, Responsibly project, where the goal is to operationalize responsibility in data sharing, integration, analysis and use.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic CVs To Build and Test Fairness-Aware Hiring Tools

    cs.CY 2025-08 conditional novelty 6.0 of 10

    A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.

  2. SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework

    cs.CR 2025-07 conditional novelty 4.0 of 10

    SynthGuard combines Kubernetes/Kubeflow orchestration with SDV/SynthGauge evaluations into a modular, owner-controlled pipeline framework for synthetic data generation.

Pith tools