REVIEW 2 major objections 5 minor 19 references
SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking
T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read A no-code tool makes synthetic rare-disease patient records that differ from common diseases by a controllable amount for fair ML testing.
desk verdict Useful no-code Synthea GUI for controllable rare-disease-like EHRs, but the 'definable degree' claim is still just free knobs plus one PCA sketch. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Global random modification of module states: every clinical event distribution (uniform, Gaussian, or exponential) is simultaneously adjusted by user-chosen variance scale, range shift, and change probability, optionally combined with fixed BMI-category multipliers for blood pressure, heart rate, glucose, cholesterol, and symptom severity.
What would settle it
Generate matched common and modified cohorts at several declared dissimilarity levels, train standard classifiers and outlier detectors, and check whether measured performance degrades smoothly and predictably with the declared difference; if performance is random or fails to track the knobs, the proxy is invalid.
Extended reading notes
Core claim
SYNRARE is a graphical interface on a rule-based patient generator that lets users create synthetic electronic health records of rare-disease patients that differ from common-disease patients only by a user-defined amount, so machine-learning algorithms can be benchmarked under controlled technical conditions without real patient data.
Load-bearing premise
That sliding the numbers that define each synthetic clinical event, plus a fixed table of body-mass-index risk multipliers, produces patient groups whose controlled difference is a fair stand-in for the real rare-versus-common diagnostic problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents SYNRARE, a GUI built on the Synthea rule-based synthetic EHR generator. Its central claim is that researchers can more easily create a minority cohort of simulated rare-disease patients that differ from a common-disease majority by a user-controlled technical degree (via global random edits to state distributions and BMI-driven comorbidity multipliers), thereby providing a privacy-preserving testbed for ML benchmarking of rare-disease detection methods. The paper describes installation requirements, three generation modes (single/multiple/all modules), state viewing and legacy-module migration, a Random Modification workflow (variance, range shift, application probability over Uniform/Gaussian/Exponential parameters), a fixed BMI comorbidity table drawn from literature, and an illustrative bronchitis-vs-perturbed-variant PCA example. Code is stated to be available on GitLab.
Significance. If the tool works as described, it lowers the barrier to creating controlled, fully synthetic imbalanced EHR cohorts for technical evaluation of classifiers and outlier detectors before real-data access is obtained. Strengths that should be credited include: open availability of the software, a no-code interface over Synthea modules, explicit legacy-module migration support, and literature-grounded BMI multipliers (Table 1). These are useful engineering contributions for the rare-disease ML community. The significance is primarily practical rather than methodological; the paper does not introduce a new generative model or a validated difficulty metric, so impact depends on whether the GUI knobs demonstrably produce reproducible, measurable cohort separation usable for benchmarking.
major comments (2)
- Abstract/Results and contribution 1 claim that SYNRARE generates RD patients that 'differ only to a definable degree' from common-disease patients, enabling controlled technical ML benchmarking. The only empirical support is the qualitative PCA sketch in Figure 1D (bronchitis vs. a 1.1/1.2/100% random modification) and a narrative remark that logistic regression becomes hard. There is no quantitative definition of 'degree' (no distance, overlap, separability, or difficulty metric), no demonstration that the free parameters (variance, range shift, probability) map monotonically or reproducibly onto classifier performance, and no ML benchmark table. Without that link, the enabling claim for controlled benchmarking is not yet secured by evidence in the manuscript.
- Module Modification and Discussion: the paper itself notes that Synthea produces idealized trajectories lacking the missingness and confounding typical of real RD EHRs. The global random-edit mechanism plus fixed BMI multipliers (Table 1) therefore produce technical dissimilarity whose clinical coherence and relevance as a proxy for the real rare-vs-common diagnostic problem are unvalidated. At minimum, the manuscript should either (a) report a small controlled experiment showing that knob settings produce graded, measurable separation on standard ML tasks, or (b) substantially qualify the claim so that 'definable degree' is clearly presented as a free user setting rather than a demonstrated property of the generated cohorts.
minor comments (5)
- Figure 1D is described as a PCA of original vs. modified bronchitis but no sample sizes, feature construction, or preprocessing are stated; a short methods note would make the illustration reproducible.
- Table 1: 'multipler' is misspelled (should be 'multiplier'); units and sources for each row could be cited more explicitly next to the table.
- Introduction and Synthea Overview: the Cystic Fibrosis example is helpful but could briefly note which Synthea module fields are being idealized, to set expectations for users.
- Availability: the GitLab link is given; a short statement of license, version of Synthea targeted, and whether example modified modules are shipped would aid adoption.
- Several incomplete citations appear (e.g., Li et al.; Ren et al. without full bibliographic detail in the reference list as presented).
Circularity Check
No circular derivation: SYNRARE is a GUI tool for user-controlled Synthea perturbations, not a first-principles prediction that folds inputs into claimed outputs.
full rationale
The paper presents a software contribution (GUI over Synthea) whose central claim is that users can apply global Random Modification (variance, range shift, probability) and fixed BMI multipliers (Table 1, from Guh et al. 2009) to produce synthetic cohorts with a user-chosen degree of dissimilarity for ML benchmarking. There is no derivation chain, no fitted parameter renamed as a prediction, no uniqueness theorem, and no self-citation that is load-bearing for a mathematical result. The only empirical illustration is a qualitative PCA of bronchitis vs. a 1.1/1.2/100% variant (Figure 1D); the 'definable degree' is explicitly a free user setting, not a quantity claimed to be derived from first principles or recovered from data. Weaknesses of evidence (lack of quantitative difficulty metrics, idealized Synthea trajectories) are correctness/validation issues, not circularity. Score 0 is therefore the correct finding.
Assumptions & free parameters
free parameters (2)
- variance / range-shift / modification-probability knobs =
user-set (example 1.1 / 1.2 / 100%)
- BMI category vital-sign and risk multipliers (Table 1) =
Table 1 fixed factors
assumptions (4)
- domain assumption Rule-based Synthea modules generate fully synthetic, bias-free trajectories suitable for controlled ML evaluation of rare-vs-common separation.
- ad hoc to paper Global random edits to state distributions (Uniform/Gaussian/Exponential parameters and transition probabilities) produce a minority class that differs from the majority by a user-definable technical degree.
- domain assumption BMI-category multipliers for BP, heart rate, glucose, lipids, liver enzymes, symptom severity, and disease probability (Table 1) adequately capture obesity-related comorbidity effects for synthetic generation.
- domain assumption Idealized Synthea trajectories without real-world missingness or confounding are still useful for delineating ML capability limits under controlled difficulty.
invented entities (1)
-
SYNRARE GUI and global module-modification workflow
Cite this review
Pith. "Pith review of SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking." pith.science (2026). https://pith.science/paper/ROHMJNBL
@misc{pith2026260709404,
author = {Pith},
title = {Pith review of: SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking},
year = {2026},
howpublished = {\url{https://pith.science/paper/ROHMJNBL}},
note = {Machine review of arXiv:2607.09404}
}
read the original abstract
Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants. Machine Learning Algorithms applied to Electronic Health Records show promise for accelerating the diagnosis; however, legal and privacy concerns pose significant barriers. To address these issues, Synthetic Data Generation is an alternative method for obtaining Electronic Health Records and can be applied with any Machine Learning algorithm for benchmarking and development purposes. Despite the availability of Synthetic Data Generation algorithms, support for generating a subset of patients that differ in a definable degree from the majority to simulate patients with RD is often lacking. Results: We present SYNRARE, a graphical user interface based on the Synthea framework that enables easier modification and generation of synthetic Electronic Health Records of RD patients, which differ only to a definable degree from patients with common diseases, thereby enabling the benchmarking and testing of algorithms under controlled technical conditions. SYNRARE enables researchers to rapidly benchmark their Machine Learning algorithms across any scenario. Availability and implementation: SYNRARE, including detailed instructions for installing, is available at https://gitlab.sdu.dk/screen4care/synrare.
Reference graph
Works this paper leans on
-
[1]
Nature Communications , year=
A machine learning model for identifying patients at risk for wild-type transthyretin amyloid cardiomyopathy , author=. Nature Communications , year=
-
[2]
PLoS ONE , year=
Detecting rare diseases in electronic health records using machine learning and knowledge engineering: Case study of acute hepatic porphyria , author=. PLoS ONE , year=
-
[3]
Orphanet Journal of Rare Diseases , year=
Development of a rare disease algorithm to identify persons at risk of Gaucher disease using electronic health records in the United States , author=. Orphanet Journal of Rare Diseases , year=
-
[4]
Journal of the American Medical Informatics Association , volume=
Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record , author=. Journal of the American Medical Informatics Association , volume=. 2018 , publisher=
2018
-
[5]
BMC Public Health , year=
The incidence of co-morbidities related to obesity and overweight: A systematic review and meta-analysis , author=. BMC Public Health , year=
-
[6]
Synthetic Data Engine , year =
-
[7]
Malin and Jon Duke and Walter F
Edward Choi and Siddharth Biswal and Bradley A. Malin and Jon Duke and Walter F. Stewart and Jimeng Sun , title =. CoRR , volume =. 2017 , url =. 1703.06490 , timestamp =
arXiv 2017
-
[8]
Sensors , VOLUME =
Alabsi, Basim Ahmad and Anbar, Mohammed and Rihan, Shaza Dawood Ahmed , TITLE =. Sensors , VOLUME =. 2023 , NUMBER =
2023
Show all 19 references
-
[9]
and Chaudhary, Durgesh and Avula, Venkatesh and Mudiganti, Satish and Husby, Hannah and Shahjouei, Shima and Afshar, Ardavan and Stewart, Walter F
Li, Jiang and Yan, Xiaowei S. and Chaudhary, Durgesh and Avula, Venkatesh and Mudiganti, Satish and Husby, Hannah and Shahjouei, Shima and Afshar, Ardavan and Stewart, Walter F. and Yeasin, Mohammed and Zand, Ramin and Abedi, Vida , date =. Imputation of missing values for ele...
-
[10]
Moving Beyond Medical Statistics: A Systematic Review on Missing Data Handling in Electronic Health Records
Ren, Wenhui and Liu, Zheng and Wu, Yanqiu and Zhang, Zhilong and Hong, Shenda and Liu, Huixin , date =. Moving Beyond Medical Statistics: A Systematic Review on Missing Data Handling in Electronic Health Records. , volume =. doi:10.34133/hds.0176 , abstract =
-
[11]
doi:10.1016/S0140-6736(13)61836-X , pages =
Metabolic mediators of the effects of body-mass index, overweight, and obesity on coronary heart disease and stroke: a pooled analysis of 97 prospective cohorts with 1·8 million participants , volume =. doi:10.1016/S0140-6736(13)61836-X , pages =
-
[12]
, title=
European Commission, Directorate-General for Health & Safety. , title=. 2025 , url=
2025
-
[13]
Preventing Chronic Disease , year=
Public Health and Rare Diseases: Oxymoron No More , author=. Preventing Chronic Disease , year=
-
[14]
, author=
The landscape for rare diseases in 2024. , author=. The Lancet. Global health , year=
2024
-
[15]
2024 , month =
Dubief, Jessie and Gross, Edith Sky and Faye, Fatoumata , title =. 2024 , month =
2024
-
[16]
Frontiers in Artificial Intelligence , year=
Leveraging Open Electronic Health Record Data and Environmental Exposures Data to Derive Insights Into Rare Pulmonary Disease , author=. Frontiers in Artificial Intelligence , year=
-
[17]
The Journal of Pediatrics , year=
Diagnosis of Cystic Fibrosis: Consensus Guidelines from the Cystic Fibrosis Foundation , author=. The Journal of Pediatrics , year=
-
[18]
, author=
Disparities in first evaluation of infants with cystic fibrosis since implementation of newborn screening. , author=. Journal of cystic fibrosis : official journal of the European Cystic Fibrosis Society , year=
-
[19]
, author=
Cystic Fibrosis: A Review. , author=. JAMA , year=
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.