Pith. sign in

REVIEW 2 cited by

Generating Synthetic but Plausible Healthcare Record Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.01514 v1 pith:RETVFJYW submitted 2018-07-04 stat.ML cs.LG

classification stat.MLcs.LG
keywords datasetsmethodsyntheticcontainingdatasetgansgenerategenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating datasets that "look like" given real ones is an interesting tasks for healthcare applications of ML and many other fields of science and engineering. In this paper we propose a new method of general application to binary datasets based on a method for learning the parameters of a latent variable moment that we have previously used for clustering patient datasets. We compare our method with a recent proposal (MedGan) based on generative adversarial methods and find that the synthetic datasets we generate are globally more realistic in at least two senses: real and synthetic instances are harder to tell apart by Random Forests, and the MMD statistic. The most likely explanation is that our method does not suffer from the "mode collapse" which is an admitted problem of GANs. Additionally, the generative models we generate are easy to interpret, unlike the rather obscure GANs. Our experiments are performed on two patient datasets containing ICD-9 diagnostic codes: the publicly available MIMIC-III dataset and a dataset containing admissions for congestive heart failure during 7 years at Hospital de Sant Pau in Barcelona.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stage-Diff: Stage-wise Long-Term Time Series Generation Based on Diffusion Models

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Stage-Diff generates long multivariate time series in stages, decomposing each stage into multi-scale trends and using multi-channel convolution to carry information between stages.

  2. Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A survey that categorizes tabular data synthesis by generation objectives and adds a benchmark comparison of six models on Adult and CreditRisk.

Pith tools