{"id":"c29170ca-a776-4d1f-82a4-7fc4d4be231f","arxiv_id":"2607.28671","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"EHR- and DXA-based survival models outperform clinically reported FRAX scores for ranking fragility fracture risk in DXA-tested adults (internal C-index 0.779 vs 0.653; external 0.714 vs 0.590).","lead":"Doctors can improve fracture risk prediction by combining bone density scans with electronic health record data, outperforming the standard FRAX calculator in two large US health systems. The best model ranked patients correctly more often, but the study acknowledges that calibration and prospective testing are still needed before clinical use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The FRAX comparison is not endpoint-matched: the models predict a broader EHR-defined fragility-fracture outcome over ~3 years, while FRAX is a 10-year MOF probability. The reported C-index gap may therefore be inflated by the benchmark rather than by model superiority.","rationale":"The reader's weakest assumption—fairness of the FRAX benchmark—is exactly the load-bearing concern I identify. The paper is transparently self-aware: Section 2.1.1 states the outcome is not identical to FRAX's MOF endpoint, and the Discussion calls FRAX a clinically relevant benchmark rather than a perfectly matched comparator. However, the abstract's central claim ('showed better discrimination than clinically reported FRAX scores') is not qualified by this mismatch, so the conditional verdict is appropriate. A matched-endpoint analysis is feasible with existing data and would settle whether the C-index gap reflects true superiority or a benchmark artifact. No other concern—complete-case analysis, missing code, or FRAX extraction details—is as directly load-bearing, because the endpoint/horizon mismatch is the primary basis for the headline comparison. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":19196,"tokens_out":5866,"duration_ms":58445,"concrete_test":"Re-run both internal and external validation with a FRAX-matched endpoint: incident hip (S72*), humerus (S42*), forearm (S52*), and clinically diagnosed vertebral fractures only—excluding lower-leg/ankle/foot (S82*) and non-clinical M48.x/M49.5 codes—and evaluate FRAX and the models over the same follow-up horizon (e.g., 5-year restricted C-index and time-dependent AUC). If the Setting A Cox model's advantage over FRAX narrows to <0.05 or reverses, the 'better than FRAX' claim is not established for the FRAX endpoint. Also report the proportion of DXA reports with FRAX computed without femoral-neck BMD; if substantial, repeat the benchmark using BMD-inclusive FRAX.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1.1 defines the outcome as first incident fragility fracture using ICD-10 prefixes that include lower leg/ankle/foot (S82*) and all spine/vertebral codes (M48.4/M48.5/M49.5, S22*, S32*, T08*). Section 2.2.4 evaluates Harrell's C-index over observed follow-up (median ~35 months) and at 1/2/5 years. FRAX, extracted from DXA reports, is a 10-year major osteoporotic fracture probability whose canonical endpoint is hip, clinical spine, forearm, and humerus. The central claim—that EHR/DXA models are better than FRAX—is based on C-index gaps (0.779 vs 0.653 internal; 0.714 vs 0.590 external) that compare a model trained and evaluated on a broader, shorter-horizon endpoint against a fixed 10-year risk score for a narrower endpoint. If the added fracture types (e.g., ankle/foot) are more predictable from EHR variables, or if the vertebral codes capture silent/prevalent fractures, the apparent superiority could be a benchmark artifact rather than true model superiority. The authors acknowledge this mismatch in Section 2.1.1 and the Discussion, but the abstract's headline statement is not conditioned on it. Thus the load-bearing assumption—that the FRAX comparison is fair despite endpoint and horizon differences—remains unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript develops time-to-event fracture prediction models for adults aged 50 years and older using DXA-derived T-scores and structured EHR data. The models were developed at NYP/WCM (n=11,510, 858 fractures), internally validated in 20 repeated 80/20 splits, and externally validated in an INPC cohort restricted to 1,932 FRAX-available patients. Predictor settings A (expanded EHR) and B (FRAX-like) are compared across penalized Cox regression, random survival forest, gradient-boosting survival, and XGBoost survival, with FRAX probabilities extracted from DXA reports as the benchmark. The central claim is improved discrimination over FRAX (Harrell's C-index 0.779 vs 0.653 internal; 0.714 vs 0.590 external) and that a transparent Cox model is competitive with machine-learning alternatives. The authors explicitly state that calibration, prospective evaluation, and implementation assessment remain before clinical use.","tokens_in":19576,"tokens_out":7226,"duration_ms":71789,"significance":"If the FRAX comparison is endpoint-matched, the study makes a useful contribution: two independent health systems, prespecified feature sets, paired internal and external comparisons, and a balanced message that ML is not uniformly better than Cox regression. The strongest internal result is a parsimonious penalized Cox model, and the external result points to gradient-boosting survival, which is a credible and nontrivial finding. However, the paper's headline claim—that EHR/DXA models are better than FRAX—cannot be evaluated from the current results because the comparator differs in endpoint and time horizon. The manuscript is candid about limitations, but the abstract and the headline numerical comparisons need to be conditioned on that mismatch.","major_comments":[{"comment":"The FRAX benchmark is not endpoint- or horizon-matched. eTable 1 shows the EHR outcome includes lower leg/ankle/foot fractures (S82*) and all vertebral/spine codes, whereas FRAX's major osteoporotic fracture endpoint is hip, clinical spine, forearm, and humerus; FRAX is also a 10-year probability, while the models are evaluated over observed follow-up (median 35 months) and at 1-, 2-, and 5-year horizons. The C-index gaps reported in §3.2.1 and §3.4 (0.779 vs 0.653; 0.714 vs 0.590) therefore do not establish that the EHR/DXA models outperform FRAX for the same clinical prediction task. Please add a sensitivity analysis restricted to the canonical FRAX fracture codes, and evaluate both the models and FRAX at a common 5-year IPCW horizon, or explicitly qualify every 'better than FRAX' statement in the Abstract as a comparison against a different endpoint and horizon.","section":"§2.1.1, §2.2.4, §3.4"},{"comment":"The external validation cohort is only the FRAX-available subset of INPC (1,932 of 2,595 DXA patients). Because FRAX reporting in DXA reports is likely associated with clinical referral patterns, this restriction may make the external results nonrepresentative of all DXA-tested patients. The Discussion acknowledges the restriction, but the Abstract and §3.3 present the INPC result as external validation without this caveat. Please provide model discrimination on the full INPC cohort without the paired-FRAX requirement, and compare baseline characteristics and fracture rates of included versus excluded INPC patients. Without this, the external validation claim is limited to a subgroup whose reports contain FRAX.","section":"§2.1, Figure 1, §3.3"},{"comment":"The complete-case analysis excludes participants with missing or implausible T-scores and all T-scores greater than +1. The flowchart excludes 4,017 of 15,527 NYP/WCM participants and 663 of 2,595 INPC participants, but the number excluded specifically because of the T-score > +1 rule is not reported. This exclusion could truncate the low-risk tail of the population and inflate discrimination. Please report counts and baseline characteristics by exclusion reason, and repeat the primary analysis with T-scores > +1 retained (or with a sensitivity threshold) to verify the advantage over FRAX is not an artifact of this exclusion rule.","section":"§2.1, §2.1.3"},{"comment":"The outcome is ascertained from ICD-10 diagnosis codes, including vertebral codes such as M48.4/M48.5/M49.5 that can represent prevalent or clinically silent vertebral fractures rather than incident events. Because prior fracture and T-score are the strongest predictors, differential misclassification correlated with baseline characteristics could inflate the apparent discrimination of the models. Please report the proportion of events in each fracture category in the analysis cohort and provide a sensitivity analysis excluding vertebral fractures, or requiring imaging or encounter confirmation for vertebral fracture events, to assess whether the central C-index gap is robust to outcome definition.","section":"§2.1.1, §4"}],"minor_comments":[{"comment":"The 'No Fracture' count for NYP/WCM is 10,663 in eTable 1 but 10,652 in the text and Table 1; reconcile the discrepancy.","section":"eTable 1"},{"comment":"Section 3.3 appears to be an empty heading; subsequent subsections are numbered 3.4 and 3.4.1. Re-number the results sections.","section":"§3.3"},{"comment":"The significance asterisks in the figures are not fully specified in the captions; state explicitly which comparison each significance marker refers to (e.g., vs FRAX, vs Cox Setting A).","section":"Figures 2 and 3"},{"comment":"The paired t-tests across 20 repeated random splits ignore the non-independence of the test sets within the same cohort; consider reporting participant-level bootstrap confidence intervals or stating that the split-level tests are descriptive.","section":"§2.2.3"},{"comment":"The fixed penalization parameter alpha = 0.01 for the Cox models is prespecified, which is a strength, but no sensitivity analysis is reported for this choice; a brief robustness check would be useful.","section":"§2.2.2"}],"recommendation":"major_revision","confidential_remarks":"The endpoint/horizon mismatch raised in the stress-test is confirmed by my reading and is the central load-bearing issue. It is fixable within the scope of the manuscript: a matched-endpoint and matched-horizon sensitivity analysis, plus a more qualified abstract, would make the claim defensible. The paper is otherwise transparent and the external validation is a genuine strength. I would not reject it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a competent, honest empirical paper. The most useful thing in it is the external validation across two US health systems with paired model-versus-FRAX comparisons on the same patients. The internal Cox C-index of 0.779 versus 0.653 for FRAX, and 0.714 versus 0.590 externally, is a real signal that structured EHR data plus DXA T-scores improve risk ranking in this DXA-tested population. I believe that result.\n\nWhat is actually new: the two-cohort external validation and the consistent value of the minimum central T-score over femoral neck alone. The methods are not novel — penalized Cox, RSF, gradient boosting, XGBoost survival — but they are prespecified and carefully compared with paired tests. The authors are also unusually transparent. They repeatedly flag the FRAX endpoint/horizon mismatch, the FRAX-available restriction in the external cohort, missing-data issues, and the need for calibration and prospective evaluation. That transparency is earned.\n\nThe soft spot is the one the reader flagged, and I think it is a real caveat rather than a fatal flaw. The models are trained on a broader EHR-defined fragility fracture endpoint (including lower leg/ankle/foot and all spine/vertebral codes) and evaluated over a median follow-up around 35 months, while FRAX is a 10-year major osteoporotic fracture probability for a narrower endpoint. Evaluating a shorter-horizon, broader-endpoint model head-to-head against a 10-year score will tend to inflate the apparent advantage. The authors acknowledge this, but the abstract's headline statement is not conditioned on it. To establish \"better than FRAX\" for the same clinical question, they would need to report C-indices on an MOF-matched endpoint or at the same time horizon. Since they don't, the magnitude of the improvement should be read as an upper bound.\n\nTwo smaller issues: the external validation cohort was restricted to patients with FRAX available, which may bias the transportability estimate, and no code or data are released, so the DXA NLP extraction and feature definitions cannot yet be reproduced. Neither is disqualifying, but both limit confidence.\n\nBottom line: this deserves a serious referee and a decision after revision, not desk rejection. The clinical signal is plausible and the limitations are mostly stated. I'd recommend the editor send it out, with a referee asked to request endpoint-matched sensitivity analyses and calibration results.","headline":"Two-cohort external validation is the real contribution; the FRAX comparison is endpoint-mismatched and the claimed gap is likely an upper bound, but this is a solid clinical prediction paper that deserves peer review.","tokens_in":20042,"tokens_out":2479,"would_cite":true,"duration_ms":25240,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62N01"],"pacs":[],"model":"deepseek-v4-flash","headline":"Combining DXA T-scores with routine electronic health record data ranks fragility-fracture risk better than the FRAX score recorded in the DXA report, in two independent US cohorts.","keywords":["fragility fracture","fracture risk prediction","DXA T-score","FRAX","electronic health records","survival analysis","machine learning","external validation"],"falsifier":"Compute both models on the exact endpoint FRAX was designed for—hip, clinical spine, forearm, and humerus—over a 10-year horizon with identical censoring and the same FRAX inputs; if the EHR/DXA model no longer shows a C-index advantage, the central claim would be undercut.","tokens_in":19100,"feed_emoji":"🦴","tokens_out":11290,"duration_ms":89347,"temperature":0.7,"pith_summary":"The paper tries to show that routine electronic health record data, added to the T-scores already in a DXA report, can rank a patient's risk of a fragility fracture more accurately than the FRAX probability printed on that same report. In a development cohort of 11,510 DXA-tested adults and an external cohort of 1,932 from a different health system, an expanded penalized survival-regression model reached concordance indices of 0.779 and 0.714, versus 0.653 and 0.590 for FRAX. A gradient-boosting survival model edged out the regression model in external validation (0.725), but the interpretable regression model was best internally. The practical stake is that better risk ranking from data already in the chart could improve which older adults are offered treatment, though calibration and prospective testing are still needed.","feed_headline":"Outperform FRAX: EHR plus DXA ranks fracture risk better","feed_subtitle":"Adding routine electronic health records to DXA T-scores lifts concordance from 0.65 to 0.78 internally and 0.59 to 0.71 externally","key_machinery":"The carrying mechanism is the expanded (Setting A) penalized survival-regression model, whose risk score is built from the minimum T-score across three DXA sites plus structured EHR predictors aggregated into counts; it is benchmarked against the FRAX probability extracted from the same DXA report. The minimum T-score—the lowest of lumbar spine, femoral neck, and total hip—is the paper's key DXA-derived variable, capturing skeletal-site discordance that femoral-neck-only tools miss. The time-to-event framework with paired internal and external validation allows direct C-index comparison on the same patients.","core_discovery":"The paper claims that a penalized survival-regression model combining DXA-derived T-scores with structured electronic health record (EHR) data—prior fracture, demographics, comorbidity counts, medications with negative skeletal effects, and osteoporosis treatment history—ranks incident fragility-fracture risk better than the FRAX 10-year major osteoporotic fracture probability recorded in the DXA report. In the development cohort this expanded model reached a concordance index (C-index) of 0.779 versus 0.653 for FRAX in internal validation; transported without refitting to an independent health-system cohort it reached 0.714 versus 0.590 for FRAX, with gradient-boosting survival slightly hig","pith_inferences":["Editorial extension: A direct translation of these results is that health systems could deploy an interpretable survival model at the point of DXA reporting, using structured EHR fields already present, while handling missing data with explicit workflows.","Editorial extension: The minimum-T-score finding points to a simple DXA report change—displaying the lowest site T-score—that could improve risk communication even before model deployment.","Editorial extension: A natural next study is prospective collection of both the model score and FRAX at the same visit, with falls and treatment decisions recorded, to test whether the ranking gain reduces undertreatment of high-risk patients."],"forward_implications":["A DXA-tested patient's fracture risk can be ranked more accurately than the FRAX number in the report, using data already in the chart.","The lowest T-score across the three central DXA sites carries more information than femoral neck T-score alone, suggesting femoral-neck-only tools may underuse DXA data.","An interpretable survival-regression model can match or beat machine-learning survival models in development, so the gain does not require black-box methods.","The model retained most of its discrimination when transported to a different health system, indicating the approach is not specific to one site's data.","Shorter-horizon (1- and 2-year) discrimination is especially high, which matters for near-term treatment decisions."],"fun_headline_variants":["EHR plus DXA beats FRAX for fracture prediction","Fracture risk ranking improved by adding EHR data","New model surpasses FRAX in fracture risk scores","Combining DXA T-scores and EHR outdoes FRAX","Machine learning with EHR tops FRAX for fracture risk"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that FRAX, a 10-year probability for a narrower fracture set, is a fair benchmark for models predicting a broader set of fragility fractures over a median follow-up of about three years; if endpoint and horizon differences inflate the gap, the 'better than FRAX' claim is not established.","fun_headline_variants_meta":{"raw":{"variants":["EHR plus DXA beats FRAX for fracture prediction","Fracture risk ranking improved by adding EHR data","New model surpasses FRAX in fracture risk scores","Combining DXA T-scores and EHR outdoes FRAX","Machine learning with EHR tops FRAX for fracture risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000711,"raw_usage":{"total_tokens":3117,"prompt_tokens":901,"completion_tokens":2216,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":2132}},"tokens_in":645,"tokens_out":2216,"duration_ms":14269,"temperature":1.0,"reasoning_tokens":2132,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:42:28.207350+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute both models on the exact endpoint FRAX was designed for—hip, clinical spine, forearm, and humerus—over a 10-year horizon with identical censoring and the same FRAX inputs; if the EHR/DXA model no longer shows a C-index advantage, the central claim would be undercut.","supporting_citations":[],"review_version":1}