Pith. sign in

REVIEW 2 major objections 5 minor 19 references

SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking

T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read A no-code tool makes synthetic rare-disease patient records that differ from common diseases by a controllable amount for fair ML testing.

desk verdict Useful no-code Synthea GUI for controllable rare-disease-like EHRs, but the 'definable degree' claim is still just free knobs plus one PCA sketch. read the letter →

arxiv 2607.09404 v1 pith:ROHMJNBL submitted 2026-07-10 cs.LG

classification cs.LG
keywords syntheticdatagenerationrarediseaseelectronichealthrecordsmachinelearningbenchmarkingSyntheatrajectorysimulationBMIcomorbiditymodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Rare-disease diagnosis is slow because symptoms overlap common conditions, and privacy rules block easy access to real electronic health records for machine-learning work. This paper introduces SYNRARE, a graphical interface built on an existing rule-based patient simulator, that lets researchers take ordinary disease modules and globally tweak their event distributions so a minority of patients become a close rare-disease-like variant. The degree of difference is set by the user (variance, range shift, and probability of change), and an optional body-mass-index layer adds comorbidity risk multipliers drawn from published evidence. The resulting synthetic cohorts keep full control over trajectories and demographics, avoid biases from real data, and give a controlled test bed for classifiers, outlier detectors, and other algorithms before anyone seeks real patient records. The authors present the tool as a practical way to run technical benchmarking under known difficulty, not as a source of clinical insight into any particular rare disease.

What carries the argument

Global random modification of module states: every clinical event distribution (uniform, Gaussian, or exponential) is simultaneously adjusted by user-chosen variance scale, range shift, and change probability, optionally combined with fixed BMI-category multipliers for blood pressure, heart rate, glucose, cholesterol, and symptom severity.

What would settle it

Generate matched common and modified cohorts at several declared dissimilarity levels, train standard classifiers and outlier detectors, and check whether measured performance degrades smoothly and predictably with the declared difference; if performance is random or fails to track the knobs, the proxy is invalid.

Watch

Extended reading notes

Core claim

SYNRARE is a graphical interface on a rule-based patient generator that lets users create synthetic electronic health records of rare-disease patients that differ from common-disease patients only by a user-defined amount, so machine-learning algorithms can be benchmarked under controlled technical conditions without real patient data.

Load-bearing premise

That sliding the numbers that define each synthetic clinical event, plus a fixed table of body-mass-index risk multipliers, produces patient groups whose controlled difference is a fair stand-in for the real rare-versus-common diagnostic problem.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript presents SYNRARE, a GUI built on the Synthea rule-based synthetic EHR generator. Its central claim is that researchers can more easily create a minority cohort of simulated rare-disease patients that differ from a common-disease majority by a user-controlled technical degree (via global random edits to state distributions and BMI-driven comorbidity multipliers), thereby providing a privacy-preserving testbed for ML benchmarking of rare-disease detection methods. The paper describes installation requirements, three generation modes (single/multiple/all modules), state viewing and legacy-module migration, a Random Modification workflow (variance, range shift, application probability over Uniform/Gaussian/Exponential parameters), a fixed BMI comorbidity table drawn from literature, and an illustrative bronchitis-vs-perturbed-variant PCA example. Code is stated to be available on GitLab.

Significance. If the tool works as described, it lowers the barrier to creating controlled, fully synthetic imbalanced EHR cohorts for technical evaluation of classifiers and outlier detectors before real-data access is obtained. Strengths that should be credited include: open availability of the software, a no-code interface over Synthea modules, explicit legacy-module migration support, and literature-grounded BMI multipliers (Table 1). These are useful engineering contributions for the rare-disease ML community. The significance is primarily practical rather than methodological; the paper does not introduce a new generative model or a validated difficulty metric, so impact depends on whether the GUI knobs demonstrably produce reproducible, measurable cohort separation usable for benchmarking.

major comments (2)
  1. Abstract/Results and contribution 1 claim that SYNRARE generates RD patients that 'differ only to a definable degree' from common-disease patients, enabling controlled technical ML benchmarking. The only empirical support is the qualitative PCA sketch in Figure 1D (bronchitis vs. a 1.1/1.2/100% random modification) and a narrative remark that logistic regression becomes hard. There is no quantitative definition of 'degree' (no distance, overlap, separability, or difficulty metric), no demonstration that the free parameters (variance, range shift, probability) map monotonically or reproducibly onto classifier performance, and no ML benchmark table. Without that link, the enabling claim for controlled benchmarking is not yet secured by evidence in the manuscript.
  2. Module Modification and Discussion: the paper itself notes that Synthea produces idealized trajectories lacking the missingness and confounding typical of real RD EHRs. The global random-edit mechanism plus fixed BMI multipliers (Table 1) therefore produce technical dissimilarity whose clinical coherence and relevance as a proxy for the real rare-vs-common diagnostic problem are unvalidated. At minimum, the manuscript should either (a) report a small controlled experiment showing that knob settings produce graded, measurable separation on standard ML tasks, or (b) substantially qualify the claim so that 'definable degree' is clearly presented as a free user setting rather than a demonstrated property of the generated cohorts.
minor comments (5)
  1. Figure 1D is described as a PCA of original vs. modified bronchitis but no sample sizes, feature construction, or preprocessing are stated; a short methods note would make the illustration reproducible.
  2. Table 1: 'multipler' is misspelled (should be 'multiplier'); units and sources for each row could be cited more explicitly next to the table.
  3. Introduction and Synthea Overview: the Cystic Fibrosis example is helpful but could briefly note which Synthea module fields are being idealized, to set expectations for users.
  4. Availability: the GitLab link is given; a short statement of license, version of Synthea targeted, and whether example modified modules are shipped would aid adoption.
  5. Several incomplete citations appear (e.g., Li et al.; Ren et al. without full bibliographic detail in the reference list as presented).

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: SYNRARE is a GUI tool for user-controlled Synthea perturbations, not a first-principles prediction that folds inputs into claimed outputs.

full rationale

The paper presents a software contribution (GUI over Synthea) whose central claim is that users can apply global Random Modification (variance, range shift, probability) and fixed BMI multipliers (Table 1, from Guh et al. 2009) to produce synthetic cohorts with a user-chosen degree of dissimilarity for ML benchmarking. There is no derivation chain, no fitted parameter renamed as a prediction, no uniqueness theorem, and no self-citation that is load-bearing for a mathematical result. The only empirical illustration is a qualitative PCA of bronchitis vs. a 1.1/1.2/100% variant (Figure 1D); the 'definable degree' is explicitly a free user setting, not a quantity claimed to be derived from first principles or recovered from data. Weaknesses of evidence (lack of quantitative difficulty metrics, idealized Synthea trajectories) are correctness/validation issues, not circularity. Score 0 is therefore the correct finding.

Assumptions & free parameters 2 free parameters · 4 assumptions · 1 invented entities

The central claim rests on Synthea’s rule-based modules as a faithful enough generator, on user-chosen global perturbation knobs as a definition of ‘definable dissimilarity,’ and on literature-derived BMI multipliers as realistic comorbidity effects. No free parameters are fitted to new data; the invented entity is the SYNRARE tool itself.

free parameters (2)
  • variance / range-shift / modification-probability knobs = user-set (example 1.1 / 1.2 / 100%)
    User-chosen global multipliers (example values 1.1, 1.2, 100%) that define how far the rare variant is pushed from the base module; they are free controls, not data fits.
  • BMI category vital-sign and risk multipliers (Table 1) = Table 1 fixed factors
    Fixed increments and multipliers (e.g., +25 mmHg systolic for Obesity III, ×2.10 disease probability) taken from or inspired by meta-analyses and applied as hard-coded factors rather than re-estimated on new data.
assumptions (4)
  • domain assumption Rule-based Synthea modules generate fully synthetic, bias-free trajectories suitable for controlled ML evaluation of rare-vs-common separation.
    Stated in Introduction and Synthea Overview as the reason for choosing Synthea over data-driven generators.
  • ad hoc to paper Global random edits to state distributions (Uniform/Gaussian/Exponential parameters and transition probabilities) produce a minority class that differs from the majority by a user-definable technical degree.
    Core of Module Modification and the bronchitis example; the paper treats this as sufficient for benchmarking difficulty.
  • domain assumption BMI-category multipliers for BP, heart rate, glucose, lipids, liver enzymes, symptom severity, and disease probability (Table 1) adequately capture obesity-related comorbidity effects for synthetic generation.
    Justified by citation to Guh et al. 2009 and related clinical statements in Module Modification.
  • domain assumption Idealized Synthea trajectories without real-world missingness or confounding are still useful for delineating ML capability limits under controlled difficulty.
    Explicitly acknowledged in Discussion as a limitation that does not invalidate technical benchmarking use.
invented entities (1)
  • SYNRARE GUI and global module-modification workflow
    purpose: Lower the barrier to creating rare-disease-like synthetic EHR cohorts with controllable dissimilarity from common-disease modules.
    The paper’s primary contribution; independent evidence is the public repository and described features rather than external clinical validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking." pith.science (2026). https://pith.science/paper/ROHMJNBL

@misc{pith2026260709404,
  author       = {Pith},
  title        = {Pith review of: SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ROHMJNBL}},
  note         = {Machine review of arXiv:2607.09404}
}
read the original abstract

Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants. Machine Learning Algorithms applied to Electronic Health Records show promise for accelerating the diagnosis; however, legal and privacy concerns pose significant barriers. To address these issues, Synthetic Data Generation is an alternative method for obtaining Electronic Health Records and can be applied with any Machine Learning algorithm for benchmarking and development purposes. Despite the availability of Synthetic Data Generation algorithms, support for generating a subset of patients that differ in a definable degree from the majority to simulate patients with RD is often lacking. Results: We present SYNRARE, a graphical user interface based on the Synthea framework that enables easier modification and generation of synthetic Electronic Health Records of RD patients, which differ only to a definable degree from patients with common diseases, thereby enabling the benchmarking and testing of algorithms under controlled technical conditions. SYNRARE enables researchers to rapidly benchmark their Machine Learning algorithms across any scenario. Availability and implementation: SYNRARE, including detailed instructions for installing, is available at https://gitlab.sdu.dk/screen4care/synrare.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

19 extracted references · 2 canonical work pages

  1. [1]

    Nature Communications , year=

    A machine learning model for identifying patients at risk for wild-type transthyretin amyloid cardiomyopathy , author=. Nature Communications , year=

  2. [2]

    PLoS ONE , year=

    Detecting rare diseases in electronic health records using machine learning and knowledge engineering: Case study of acute hepatic porphyria , author=. PLoS ONE , year=

  3. [3]

    Orphanet Journal of Rare Diseases , year=

    Development of a rare disease algorithm to identify persons at risk of Gaucher disease using electronic health records in the United States , author=. Orphanet Journal of Rare Diseases , year=

  4. [4]

    Journal of the American Medical Informatics Association , volume=

    Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record , author=. Journal of the American Medical Informatics Association , volume=. 2018 , publisher=

  5. [5]

    BMC Public Health , year=

    The incidence of co-morbidities related to obesity and overweight: A systematic review and meta-analysis , author=. BMC Public Health , year=

  6. [6]

    Synthetic Data Engine , year =

  7. [7]

    Malin and Jon Duke and Walter F

    Edward Choi and Siddharth Biswal and Bradley A. Malin and Jon Duke and Walter F. Stewart and Jimeng Sun , title =. CoRR , volume =. 2017 , url =. 1703.06490 , timestamp =

  8. [8]

    Sensors , VOLUME =

    Alabsi, Basim Ahmad and Anbar, Mohammed and Rihan, Shaza Dawood Ahmed , TITLE =. Sensors , VOLUME =. 2023 , NUMBER =

Show all 19 references
  1. [9]

    and Chaudhary, Durgesh and Avula, Venkatesh and Mudiganti, Satish and Husby, Hannah and Shahjouei, Shima and Afshar, Ardavan and Stewart, Walter F

    Li, Jiang and Yan, Xiaowei S. and Chaudhary, Durgesh and Avula, Venkatesh and Mudiganti, Satish and Husby, Hannah and Shahjouei, Shima and Afshar, Ardavan and Stewart, Walter F. and Yeasin, Mohammed and Zand, Ramin and Abedi, Vida , date =. Imputation of missing values for ele...

  2. [10]

    Moving Beyond Medical Statistics: A Systematic Review on Missing Data Handling in Electronic Health Records

    Ren, Wenhui and Liu, Zheng and Wu, Yanqiu and Zhang, Zhilong and Hong, Shenda and Liu, Huixin , date =. Moving Beyond Medical Statistics: A Systematic Review on Missing Data Handling in Electronic Health Records. , volume =. doi:10.34133/hds.0176 , abstract =

  3. [11]

    doi:10.1016/S0140-6736(13)61836-X , pages =

    Metabolic mediators of the effects of body-mass index, overweight, and obesity on coronary heart disease and stroke: a pooled analysis of 97 prospective cohorts with 1·8 million participants , volume =. doi:10.1016/S0140-6736(13)61836-X , pages =

  4. [12]

    , title=

    European Commission, Directorate-General for Health & Safety. , title=. 2025 , url=

  5. [13]

    Preventing Chronic Disease , year=

    Public Health and Rare Diseases: Oxymoron No More , author=. Preventing Chronic Disease , year=

  6. [14]

    , author=

    The landscape for rare diseases in 2024. , author=. The Lancet. Global health , year=

  7. [15]

    2024 , month =

    Dubief, Jessie and Gross, Edith Sky and Faye, Fatoumata , title =. 2024 , month =

  8. [16]

    Frontiers in Artificial Intelligence , year=

    Leveraging Open Electronic Health Record Data and Environmental Exposures Data to Derive Insights Into Rare Pulmonary Disease , author=. Frontiers in Artificial Intelligence , year=

  9. [17]

    The Journal of Pediatrics , year=

    Diagnosis of Cystic Fibrosis: Consensus Guidelines from the Cystic Fibrosis Foundation , author=. The Journal of Pediatrics , year=

  10. [18]

    , author=

    Disparities in first evaluation of infants with cystic fibrosis since implementation of newborn screening. , author=. Journal of cystic fibrosis : official journal of the European Cystic Fibrosis Society , year=

  11. [19]

    , author=

    Cystic Fibrosis: A Review. , author=. JAMA , year=

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.