Pith. sign in

REVIEW 2 cited by

Generation of Synthetic Electronic Health Records Using a Federated GAN

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.02543 v1 pith:7CZEX54W submitted 2021-09-06 cs.LG

classification cs.LG
keywords datasyntheticcentraldata-setevaluationfederatedmedicalpatients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sensitive medical data is often subject to strict usage constraints. In this paper, we trained a generative adversarial network (GAN) on real-world electronic health records (EHR). It was then used to create a data-set of "fake" patients through synthetic data generation (SDG) to circumvent usage constraints. This real-world data was tabular, binary, intensive care unit (ICU) patient diagnosis data. The entire data-set was split into separate data silos to mimic real-world scenarios where multiple ICU units across different hospitals may have similarly structured data-sets within their own organisations but do not have access to each other's data-sets. We implemented federated learning (FL) to train separate GANs locally at each organisation, using their unique data silo and then combining the GANs into a single central GAN, without any siloed data ever being exposed. This global, central GAN was then used to generate the synthetic patients data-set. We performed an evaluation of these synthetic patients with statistical measures and through a structured review by a group of medical professionals. It was shown that there was no significant reduction in the quality of the synthetic EHR when we moved between training a single central model and training on separate data silos with individual models before combining them into a central model. This was true for both the statistical evaluation (Root Mean Square Error (RMSE) of 0.0154 for single-source vs. RMSE of 0.0169 for dual-source federated) and also for the medical professionals' evaluation (no quality difference between EHR generated from a single source and EHR generated from multiple sources).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Federated Timeline Synthesis: Scalable and Private Methodology For Model Training and Deployment

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A federated approach that trains a global generator on synthetic patient timelines produced by local models, preserving most but not all zero-shot prediction performance.

  2. Imputation of Longitudinal Data Using GANs: Challenges and Implications for Classification

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A systematic review of GAN-based longitudinal data imputation that categorizes methods and shows that most ignore missingness mechanisms, static features, and mixed data types.

Pith tools