Pith. sign in

REVIEW 1 cited by

Synthetic Data: Methods, Use Cases, and Risks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01230 v3 pith:UFBJ6NN6 submitted 2023-03-01 cs.CR cs.AIcs.CY

classification cs.CRcs.AIcs.CY
keywords datasyntheticcasesdatasetsoftenprivacysharingactual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sharing data can often enable compelling applications and analytics. However, more often than not, valuable datasets contain information of a sensitive nature, and thus, sharing them can endanger the privacy of users and organizations. A possible alternative gaining momentum in both the research community and industry is to share synthetic data instead. The idea is to release artificially generated datasets that resemble the actual data -- more precisely, having similar statistical properties. In this article, we provide a gentle introduction to synthetic data and discuss its use cases, the privacy challenges that are still unaddressed, and its inherent limitations as an effective privacy-enhancing technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A three-method auditing framework detects with roughly 87 to 97 percent accuracy whether classifiers, generators, and t-SNE plots were trained on or derived from LLM-generated synthetic data.

Pith tools