REVIEW 4 major objections 5 minor 35 references
Combining DXA T-scores with routine electronic health record data ranks fragility-fracture risk better than the FRAX score recorded in the DXA report, in two independent US cohorts.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:42 UTC pith:COPM5NNK
load-bearing objection Two-cohort external validation is the real contribution; the FRAX comparison is endpoint-mismatched and the claimed gap is likely an upper bound, but this is a solid clinical prediction paper that deserves peer review. the 4 major comments →
Fracture Risk Prediction in Adults Over 50 Years Old Using DXA and EHR: Comparison of Traditional and Machine Learning Models in Two Large Cohorts
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that a penalized survival-regression model combining DXA-derived T-scores with structured electronic health record (EHR) data—prior fracture, demographics, comorbidity counts, medications with negative skeletal effects, and osteoporosis treatment history—ranks incident fragility-fracture risk better than the FRAX 10-year major osteoporotic fracture probability recorded in the DXA report. In the development cohort this expanded model reached a concordance index (C-index) of 0.779 versus 0.653 for FRAX in internal validation; transported without refitting to an independent health-system cohort it reached 0.714 versus 0.590 for FRAX, with gradient-boosting survival slightly hig
What carries the argument
The carrying mechanism is the expanded (Setting A) penalized survival-regression model, whose risk score is built from the minimum T-score across three DXA sites plus structured EHR predictors aggregated into counts; it is benchmarked against the FRAX probability extracted from the same DXA report. The minimum T-score—the lowest of lumbar spine, femoral neck, and total hip—is the paper's key DXA-derived variable, capturing skeletal-site discordance that femoral-neck-only tools miss. The time-to-event framework with paired internal and external validation allows direct C-index comparison on the same patients.
Load-bearing premise
The load-bearing premise is that FRAX, a 10-year probability for a narrower fracture set, is a fair benchmark for models predicting a broader set of fragility fractures over a median follow-up of about three years; if endpoint and horizon differences inflate the gap, the 'better than FRAX' claim is not established.
What would settle it
Compute both models on the exact endpoint FRAX was designed for—hip, clinical spine, forearm, and humerus—over a 10-year horizon with identical censoring and the same FRAX inputs; if the EHR/DXA model no longer shows a C-index advantage, the central claim would be undercut.
If this is right
- A DXA-tested patient's fracture risk can be ranked more accurately than the FRAX number in the report, using data already in the chart.
- The lowest T-score across the three central DXA sites carries more information than femoral neck T-score alone, suggesting femoral-neck-only tools may underuse DXA data.
- An interpretable survival-regression model can match or beat machine-learning survival models in development, so the gain does not require black-box methods.
- The model retained most of its discrimination when transported to a different health system, indicating the approach is not specific to one site's data.
- Shorter-horizon (1- and 2-year) discrimination is especially high, which matters for near-term treatment decisions.
Where Pith is reading between the lines
- Editorial extension: A direct translation of these results is that health systems could deploy an interpretable survival model at the point of DXA reporting, using structured EHR fields already present, while handling missing data with explicit workflows.
- Editorial extension: The minimum-T-score finding points to a simple DXA report change—displaying the lowest site T-score—that could improve risk communication even before model deployment.
- Editorial extension: A natural next study is prospective collection of both the model score and FRAX at the same visit, with falls and treatment decisions recorded, to test whether the ranking gain reduces undertreatment of high-risk patients.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript develops time-to-event fracture prediction models for adults aged 50 years and older using DXA-derived T-scores and structured EHR data. The models were developed at NYP/WCM (n=11,510, 858 fractures), internally validated in 20 repeated 80/20 splits, and externally validated in an INPC cohort restricted to 1,932 FRAX-available patients. Predictor settings A (expanded EHR) and B (FRAX-like) are compared across penalized Cox regression, random survival forest, gradient-boosting survival, and XGBoost survival, with FRAX probabilities extracted from DXA reports as the benchmark. The central claim is improved discrimination over FRAX (Harrell's C-index 0.779 vs 0.653 internal; 0.714 vs 0.590 external) and that a transparent Cox model is competitive with machine-learning alternatives. The authors explicitly state that calibration, prospective evaluation, and implementation assessment remain before clinical use.
Significance. If the FRAX comparison is endpoint-matched, the study makes a useful contribution: two independent health systems, prespecified feature sets, paired internal and external comparisons, and a balanced message that ML is not uniformly better than Cox regression. The strongest internal result is a parsimonious penalized Cox model, and the external result points to gradient-boosting survival, which is a credible and nontrivial finding. However, the paper's headline claim—that EHR/DXA models are better than FRAX—cannot be evaluated from the current results because the comparator differs in endpoint and time horizon. The manuscript is candid about limitations, but the abstract and the headline numerical comparisons need to be conditioned on that mismatch.
major comments (4)
- [§2.1.1, §2.2.4, §3.4] The FRAX benchmark is not endpoint- or horizon-matched. eTable 1 shows the EHR outcome includes lower leg/ankle/foot fractures (S82*) and all vertebral/spine codes, whereas FRAX's major osteoporotic fracture endpoint is hip, clinical spine, forearm, and humerus; FRAX is also a 10-year probability, while the models are evaluated over observed follow-up (median 35 months) and at 1-, 2-, and 5-year horizons. The C-index gaps reported in §3.2.1 and §3.4 (0.779 vs 0.653; 0.714 vs 0.590) therefore do not establish that the EHR/DXA models outperform FRAX for the same clinical prediction task. Please add a sensitivity analysis restricted to the canonical FRAX fracture codes, and evaluate both the models and FRAX at a common 5-year IPCW horizon, or explicitly qualify every 'better than FRAX' statement in the Abstract as a comparison against a different endpoint and horizon.
- [§2.1, Figure 1, §3.3] The external validation cohort is only the FRAX-available subset of INPC (1,932 of 2,595 DXA patients). Because FRAX reporting in DXA reports is likely associated with clinical referral patterns, this restriction may make the external results nonrepresentative of all DXA-tested patients. The Discussion acknowledges the restriction, but the Abstract and §3.3 present the INPC result as external validation without this caveat. Please provide model discrimination on the full INPC cohort without the paired-FRAX requirement, and compare baseline characteristics and fracture rates of included versus excluded INPC patients. Without this, the external validation claim is limited to a subgroup whose reports contain FRAX.
- [§2.1, §2.1.3] The complete-case analysis excludes participants with missing or implausible T-scores and all T-scores greater than +1. The flowchart excludes 4,017 of 15,527 NYP/WCM participants and 663 of 2,595 INPC participants, but the number excluded specifically because of the T-score > +1 rule is not reported. This exclusion could truncate the low-risk tail of the population and inflate discrimination. Please report counts and baseline characteristics by exclusion reason, and repeat the primary analysis with T-scores > +1 retained (or with a sensitivity threshold) to verify the advantage over FRAX is not an artifact of this exclusion rule.
- [§2.1.1, §4] The outcome is ascertained from ICD-10 diagnosis codes, including vertebral codes such as M48.4/M48.5/M49.5 that can represent prevalent or clinically silent vertebral fractures rather than incident events. Because prior fracture and T-score are the strongest predictors, differential misclassification correlated with baseline characteristics could inflate the apparent discrimination of the models. Please report the proportion of events in each fracture category in the analysis cohort and provide a sensitivity analysis excluding vertebral fractures, or requiring imaging or encounter confirmation for vertebral fracture events, to assess whether the central C-index gap is robust to outcome definition.
minor comments (5)
- [eTable 1] The 'No Fracture' count for NYP/WCM is 10,663 in eTable 1 but 10,652 in the text and Table 1; reconcile the discrepancy.
- [§3.3] Section 3.3 appears to be an empty heading; subsequent subsections are numbered 3.4 and 3.4.1. Re-number the results sections.
- [Figures 2 and 3] The significance asterisks in the figures are not fully specified in the captions; state explicitly which comparison each significance marker refers to (e.g., vs FRAX, vs Cox Setting A).
- [§2.2.3] The paired t-tests across 20 repeated random splits ignore the non-independence of the test sets within the same cohort; consider reporting participant-level bootstrap confidence intervals or stating that the split-level tests are descriptive.
- [§2.2.2] The fixed penalization parameter alpha = 0.01 for the Cox models is prespecified, which is a strength, but no sensitivity analysis is reported for this choice; a brief robustness check would be useful.
Circularity Check
No significant circularity: empirical validation study with independent held-out splits, external cohort, and FRAX benchmark; endpoint mismatch is a validity caveat, not a circular derivation.
full rationale
The paper makes no mathematical derivation claim that could reduce to its inputs. It fits time-to-event models to baseline DXA/EHR predictors and evaluates them on 20 repeated 80/20 internal splits and on an independent INPC cohort without refitting (Sections 2.2.3, 3.2-3.4). FRAX probabilities are extracted from DXA reports and used only as a benchmark comparator, not as a predictor (Section 2.1.3), so the comparator is not an output of the models. The outcome (incident fragility fracture after index DXA) is defined separately from the predictors and is not used to construct predictor values. The acknowledged mismatch between the broader EHR-defined fracture endpoint and the classic FRAX MOF endpoint (Section 2.1.1) is a question of benchmark fairness and external validity, not circularity: it does not make the C-index gaps equal to the models' inputs by construction. No load-bearing self-citation, imported uniqueness theorem, or renamed known result was found. Therefore the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- Cox penalization parameter (alpha) =
0.01
- GBS hyperparameters =
n_estimators=1500, learning_rate=0.05, max_depth=3, min_samples_split=10, min_samples_leaf=5, subsample=0.8
- RSF hyperparameters =
n_estimators=800, max_features='sqrt', min_samples_split=10, min_samples_leaf=5
- XGBoost survival hyperparameters =
learning_rate=0.05, max_depth=3, min_child_weight=1, subsample=0.8, colsample_bytree=0.8, reg_lambda=1, reg_alpha=0
- T-score exclusion threshold =
T-score > +1 excluded
axioms (6)
- domain assumption ICD-10 codes identify incident fragility fractures with acceptable sensitivity and specificity.
- domain assumption Censoring is independent (non-informative) after conditioning on covariates.
- domain assumption Proportional hazards assumption holds for the Cox models.
- domain assumption FRAX benchmark is comparable to model scores despite differing endpoints and time horizons.
- domain assumption Complete-case analysis and missingness mechanisms do not qualitatively change conclusions.
- ad hoc to paper T-scores greater than +1 are artifactually elevated and can be excluded without biasing risk estimates.
read the original abstract
Accurate fracture risk prediction is important for osteoporosis management, but commonly used clinical tools may not fully use information available in electronic health records (EHRs) and dual-energy X-ray absorptiometry (DXA) reports. We developed and externally validated time-to-event fracture prediction models among adults aged 50 years or older with clinically obtained DXA reports in 2 US health care systems. The development cohort was derived from NewYork-Presbyterian/Weill Cornell Medical Center and the external validation cohort from the Indiana Network for Patient Care. Predictors included demographics, lifestyle factors, prior fracture, comorbidities, medication exposures, osteoporosis treatment history, and DXA-derived T-scores extracted from radiology reports. The outcome was time from index DXA to first incident fragility fracture identified from structured diagnosis codes. We evaluated penalized Cox regression, random survival forest, gradient-boosting survival, and XGBoost survival models using 2 prespecified predictor settings and compared discrimination with clinically reported FRAX major osteoporotic fracture probabilities. The development cohort included 11,510 adults, of whom 858 sustained incident fragility fractures; the external validation cohort included 1,932 adults, of whom 180 sustained fractures. In internal validation, the expanded Cox model achieved a mean Harrell C-index of 0.779, compared with 0.653 for FRAX. In external validation, the corresponding Cox model achieved a Harrell C-index of 0.714, compared with 0.590 for FRAX; gradient-boosting survival had the highest external discrimination (0.725). EHR- and DXA-enhanced models showed better discrimination than clinically reported FRAX scores in this DXA-tested population, but calibration assessment, prospective evaluation, and implementation workflow assessment are needed before clinical use.
Figures
Reference graph
Works this paper leans on
-
[1]
M. S. LeBoff, S. L. Greenspan, K. L. Insogna, E. M. Lewiecki, K. G. Saag, A. J. Singer, and E. S. Siris. The clinician’s guide to prevention and treatment of osteoporosis.Osteoporosis International, 33(10): 2049–2102, 2022. doi: 10.1007/s00198-022-06457-7
-
[2]
N. C. Wright, A. C. Looker, K. G. Saag, J. R. Curtis, E. S. Delzell, S. Randall, and B. Dawson-Hughes. The recent prevalence of osteoporosis and low bone mass in the united states based on bone mineral density at the femoral neck or lumbar spine.Journal of Bone and Mineral Research, 29(11):2520–2526,
-
[3]
C. W. Gillespie and P. E. Morin. Trends and disparities in osteoporosis screening among women in the united states, 2008–2014.The American Journal of Medicine, 130(3):306–316, 2017. doi: 10.1016/j. amjmed.2016.10.018
doi:10.1016/j 2008
-
[4]
C. A. Brauer, M. Coca-Perraillon, D. M. Cutler, and A. B. Rosen. Incidence and mortality of hip fractures in the united states.JAMA, 302(14):1573–1579, 2009. doi: 10.1001/jama.2009.1462
arXiv 2009
-
[5]
B. L. Moreland, J. K. Legha, K. E. Thomas, and E. R. Burns. Hip fracture-related emergency department visits, hospitalizations and deaths by mechanism of injury among adults aged 65 and older, united states 2019.Journal of Aging and Health, 35(5-6):345–355, 2023. doi: 10.1177/08982643231155902
-
[6]
R. Burge, B. Dawson-Hughes, D. H. Solomon, J. B. Wong, A. King, and A. Tosteson. Incidence and 17 economic burden of osteoporosis-related fractures in the united states, 2005–2025.Journal of Bone and Mineral Research, 22(3):465–475, 2007. doi: 10.1359/jbmr.061113
-
[7]
E. M. Lewiecki, N. C. Wright, J. R. Curtis, E. Siris, R. F. Gagel, K. G. Saag, A. J. Singer, and R. A. Adler. Hip fracture trends in the united states, 2002 to 2015.Osteoporosis International, 29(3):717–722, 2018. doi: 10.1007/s00198-017-4345-0
-
[8]
R. J. Desai, M. Mahesri, Y. Abdia, J. Barberio, A. Tong, D. Zhang, P. Mavros, S. C. Kim, and J. M. Franklin. Association of osteoporosis medication use after hip fracture with prevention of subsequent nonvertebral fractures.JAMA Network Open, 1(3):e180826, 2018. doi: 10.1001/jamanetworkopen.2018. 0826
-
[9]
P. Choksi, B. L. Gay, M. R. Haymart, and M. Papaleontiou. Physician-reported barriers to osteoporosis screening: a nationwide survey.Endocrine Practice, 29(8):606–611, 2023. doi: 10.1016/j.eprac.2023.04. 006
-
[10]
C. M. Naso, S. Y. Lin, G. Song, and H. Xue. Time trend analysis of osteoporosis prevalence among adults 50 years of age and older in the usa, 2005–2018.Osteoporosis International, 36(3):547–554, 2025. doi: 10.1007/s00198-024-07041-1
-
[11]
J. A. Kanis, O. Johnell, A. Oden, H. Johansson, and E. McCloskey. Frax and the assessment of fracture probability in men and women from the uk.Osteoporosis International, 19(4):385–397, 2008. doi: 10.1007/s00198-007-0543-5
-
[12]
Panday, A
K. Panday, A. Gona, and M. B. Humphrey. Medication-induced osteoporosis: screening and treat- ment strategies.Therapeutic Advances in Musculoskeletal Disease, 6(5):185–202, 2014. doi: 10.1177/ 1759720X14546350
2014
-
[13]
O. Lehmann, O. Mineeva, D. Veshchezerova, H. Hauselmann, L. Guyer, S. Reichenbach, T. Lehmann, O. Demler, and J. Everts-Graber. Fracture risk prediction in postmenopausal women with traditional and machine learning models in a nationwide, prospective cohort study in switzerland with validation in the uk biobank.Journal of Bone and Mineral Research, 39(8):...
-
[14]
Y. A. Almog, A. Rai, P. Zhang, A. Moulaison, R. Powell, A. Mishra, K. Weinberg, C. Hamilton, M. Oates, E. McCloskey, and S. R. Cummings. Deep learning with electronic health records for short-term fracture risk identification: Crystal bone algorithm development and validation.Journal of Medical Internet Research, 22(10):e22550, 2020. doi: 10.2196/22550
-
[15]
Regression models and life-tables.J
David R Cox. Regression models and life-tables.J. R. Stat. Soc. Series B Stat. Methodol., 34(2):187–202,
-
[16]
Frank E. Harrell, Kerry L. Lee, and Daniel B. Mark. Multivariable prognostic models: issues in devel- oping models, evaluating assumptions and adequacy, and measuring and reducing errors.Statistics in Medicine, 15(4):361–387, 1996. doi: 10.1002/(SICI)1097-0258(19960229)15:4<361::AID-SIM168>3.0. CO;2-4
-
[17]
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, and Michael S. Lauer. Random survival forests.The Annals of Applied Statistics, 2(3):841–860, 2008. doi: 10.1214/08-AOAS169
-
[18]
Jaeger, Thomas Lumley, Sara Bergstra, et al
Byron C. Jaeger, Thomas Lumley, Sara Bergstra, et al. Oblique random survival forests.The Annals of Applied Statistics, 13(3):1847–1883, 2019. doi: 10.1214/19-AOAS1261
-
[19]
Torsten Hothorn, Peter Bühlmann, Sandrine Dudoit, Annette Molinaro, and Mark J. van der Laan. Survival ensembles.Biostatistics, 7(3):355–373, 2006. doi: 10.1093/biostatistics/kxj011
-
[20]
Yuanqing Chen, Zeng Jia, Dan Mercola, and Xuefeng Xie. A gradient boosting algorithm for survival 18 analysis via direct optimization of concordance index.Computational and Mathematical Methods in Medicine, 2013:873595, 2013. doi: 10.1155/2013/873595
-
[21]
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 785–794,
-
[22]
Hajime Uno, Tianxi Cai, Michael J. Pencina, Ralph B. D’Agostino, and L. J. Wei. On the c-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data.Statistics in Medicine, 30(10):1105–1117, 2011. doi: 10.1002/sim.4154
doi:10.1002/sim.4154 2011
-
[23]
Heagerty, Thomas Lumley, and Margaret S
Patrick J. Heagerty, Thomas Lumley, and Margaret S. Pepe. Time-dependent roc curves for censored survival data and a diagnostic marker.Biometrics, 56(2):337–344, 2000. doi: 10.1111/j.0006-341X.2000. 00337.x
-
[24]
An update on the fracture risk assessment tool: What have we learned over 15+ years?Endocrinol
Laura T Dickens and Rajesh K Jain. An update on the fracture risk assessment tool: What have we learned over 15+ years?Endocrinol. Metab. Clin. North Am., 53(4):531–545, December 2024. doi: 10.1016/j.ecl.2024.08.001
-
[25]
W. D. Leslie, S. R. Majumdar, L. M. Lix, S. N. Morin, H. Johansson, A. Oden, E. V. McCloskey, and J. A. Kanis. Can change in frax score be used to “treat to target”? a population-based cohort study.Journal of Bone and Mineral Research, 29(5):1074–1080, 2014. doi: 10.1002/jbmr.2116
-
[26]
Evangelia Christodoulou, Jie Ma, Gary S. Collins, Ewout W. Steyerberg, Jan Y. Verbakel, and Ben Van Calster. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models.Journal of Clinical Epidemiology, 110:12–22, 2019. doi: 10.1016/j.jclinepi.2019.02.004
-
[27]
Simon Nusinovici, Yih Chung Tham, Michelle Yu Chak Yan, Daniel Shu Wei Ting, Jia Li, Charumathi Sabanayagam, Tien Yin Wong, and Ching-Yu Cheng. Logistic regression was as good as machine learning for predicting major chronic diseases.Journal of Clinical Epidemiology, 122:56–69, 2020. doi: 10.1016/j.jclinepi.2020.03.002
-
[28]
Gloria Hoi Yee Li, Ching Lung Cheung, Kathryn Choon Beng Tan, et al. Development and vali- dation of sex-specific hip fracture prediction models using electronic health records: a retrospective, population-based cohort study.EClinicalMedicine, 58:101876, 2023. doi: 10.1016/j.eclinm.2023.101876
arXiv 2023
-
[29]
Raju Jaiswal, Aldina Pivodic, et al. Prediction of hip fracture by high-resolution peripheral quanti- tative computed tomography in older swedish women.Journal of Bone and Mineral Research, 40(6): 779–789, 2025. doi: 10.1093/jbmr/zjaf020
-
[30]
B. C. S. de Vries, J. H. Hegeman, W. Nijmeijer, J. Geerdink, C. Seifert, and C. G. M. Groothuis- Oudshoorn. Comparing three machine learning approaches to design a risk assessment tool for future fractures: predicting a subsequent major osteoporotic fracture in fracture patients with osteopenia and osteoporosis.Osteoporosis International, 32(3):437–449, 2...
-
[31]
Christian Kruse, Pia Eiken, and Peter Vestergaard. Machine learning principles can improve hip frac- ture prediction.Calcified Tissue International, 100(4):348–360, 2017. doi: 10.1007/s00223-017-0238-7
-
[32]
neck” or “total,
Aaron Fisher, Cynthia Rudin, and Francesca Dominici. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177):1–81, 2019. 19 Supplementary materials eTable 1: Distribution of incident fragility fracture categories in the NYP/WC...
2019
-
[1972]
doi: 10.1007/978-1-4612-4380-9\_37
-
[2014]
doi: 10.1002/jbmr.2269
-
[2016]
doi: 10.1145/2939672.2939785
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.