Pith. sign in

REVIEW

Measuring the quality of Synthetic data for use in competitions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.11345 v1 pith:QAYB5BLY submitted 2018-06-29 cs.LG stat.ML

classification cs.LGstat.ML
keywords datasyntheticdatasetlearningmachineorderperformancepotential
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning has the potential to assist many communities in using the large datasets that are becoming more and more available. Unfortunately, much of that potential is not being realized because it would require sharing data in a way that compromises privacy. In order to overcome this hurdle, several methods have been proposed that generate synthetic data while preserving the privacy of the real data. In this paper we consider a key characteristic that synthetic data should have in order to be useful for machine learning researchers - the relative performance of two algorithms (trained and tested) on the synthetic dataset should be the same as their relative performance (when trained and tested) on the original dataset.

Discussion (0). Sign in to comment.

Pith tools