{"id":"4ae6862f-d9ec-4b92-81da-5d875f9fd6ed","arxiv_id":"2505.15708","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"In simulated heart failure populations, height-indexed QRS duration is the best-performing, sex-fair ECG criterion for selecting cardiac resynchronization therapy candidates.","lead":"This study used computer simulations of nearly 3,000 hearts to test whether current ECG rules for selecting heart failure patients for cardiac resynchronization therapy unfairly exclude women. It found that adjusting QRS duration for height removes most of the sex difference and may select better candidates for the treatment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Height-indexed QRSd advantage may be an artifact of the uncalibrated model's assumption that sex differences in QRSd arise solely from heart size.","rationale":"The reader's weakest assumption already identifies simulation accuracy and sex-fairness as the key limitation. I partially agree, but I sharpen the concern to a more specific structural issue: the model's design guarantees the qualitative outcome. Because conduction velocity is fixed across sexes and the same fascicular activation is used, sex differences are exclusively due to heart size; height-indexing must therefore attenuate them. The paper does not provide independent evidence that this is true in real patients. The calibration mismatch (simulated normal QRSd ~78 ms vs measured 84–92 ms) and the Figure 2 subset analysis reinforce the concern that the model is not quantitatively aligned with observed QRSd. However, the paper is explicitly an in-silico trial, and its comparative ranking of QRSd metrics may still be useful within the model's assumptions. The reader's CONDITIONAL verdict already captures the need for clinical validation, so my concern does not change the verdict; it strengthens the specific condition that the model's sex-fairness assumption be tested. I recommend UNCHANGED, with the concrete test above as a minimal additional requirement.","tokens_in":12557,"tokens_out":6054,"duration_ms":51251,"concrete_test":"Calibrate the reaction-eikonal model per patient by adjusting conduction velocity (or a global scaling factor) so that simulated normal-activation QRSd matches each patient's measured sinus-rhythm QRSd, then re-run all four scenarios and recompute sex-specific selection rates and ROC AUCs for the three QRSd metrics. If height-indexed QRSd no longer shows the smallest sex disparity or the highest AUC, the reported advantage is an artifact of the uncalibrated fixed-CV model. Alternatively, validate the 0.86 ms/cm threshold in an independent cohort with adjudicated LBBB status, reporting sensitivity/specificity by sex.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that height-indexed QRSd resolves sex disparities in LBBB selection is a direct consequence of the simulation's construction. In the Methods ('Electrophysiological (EP) simulations'), the same conduction velocity (0.65 m/s normal, 0.4 m/s slow) is used for all patients regardless of sex; the fascicular activation model is identical for both sexes. Consequently, simulated QRSd differences between males and females arise entirely from anatomical differences in heart size. Since height is a surrogate for heart size, indexing by height will, by construction, reduce sex differences — the reported 'superior performance' is therefore a model-derived, not empirically constrained, result. The paper does not test whether real sex differences in QRSd are fully explained by heart size; indeed, the simulated normal QRSd (~78 ms) is systematically shorter than measured values (84–92 ms in Table 1), and only 88% of 17 real CRT-indicated patients cross 150 ms even under LBBB+slow conduction (Figure 2). This uncalibrated offset could alter the absolute thresholds and selection rates. If real sex differences include contributions from sex-specific conduction-system or ionic properties, height-indexed thresholds would not resolve them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes an in silico trial framework for evaluating sex-specific QRS duration (QRSd) criteria in cardiac resynchronization therapy (CRT) patient selection. Using 2627 UK Biobank healthy hearts and 359 ischemic heart disease (IHD) patient anatomies, the authors simulate four scenarios (normal or left bundle branch block (LBBB) activation crossed with normal or slowed conduction) with a fixed, sex-independent conduction velocity model, compute QRSd from simulated 12-lead ECGs, and compare conventional thresholds (120/130/150 ms) with QRSd indexed by LVEDV, LV mass, and height for classifying simulated LBBB versus non-LBBB patients, with sex-specific analyses and ROC comparisons. The central claim is that height-indexed QRSd resolves sex disparities in LBBB patient selection while maintaining low non-LBBB selection rates, and achieves the best classification performance among the indexed criteria under the simulated scenarios. The manuscript also reports a validation subset of 17 real CRT-indicated IHD patients to assess how well simulated QRSd reproduces clinical QRSd ≥150 ms.","tokens_in":12842,"tokens_out":10127,"duration_ms":84428,"significance":"The strengths of the paper are its scale and design: 2986 patient-specific ventricular anatomies, four controlled pathological scenarios, evaluation of criteria over the full QRSd distribution rather than only among patients already meeting guideline inclusion criteria, and formal ROC comparisons (DeLong test) with sex-stratified analyses. The finding that LVEDV/LV-mass indexing raises non-LBBB selection rates while height indexing does not is an emergent, internally consistent result, and the proposed framework is transferable to other selection-rule questions. The authors also ship the anatomical model generation pipeline as open source, and the height-indexed thresholds are concrete, falsifiable predictions that could be tested in prospective clinical cohorts. The main caveat is external validity: the simulations assign identical conduction velocities to both sexes, so the resolution of sex differences by height indexing is largely a consequence of the model's construction, and the simulated QRSd are systematically shorter than measured values in the same cohorts.","major_comments":[{"comment":"The Results paragraph on measured QRSd states that in the healthy cohort QRSd were 'similar for both sexes (male: 93.2±23.8 ms vs 95.3±30.8 ms, P=0.7)', which directly contradicts Table 1, where healthy males and females have QRSd of 92.5±12.5 ms and 84±11.3 ms with P=8×10^-90, and also contradicts the Discussion's statement that 'significant sex differences in the baseline QRSds' were observed 'in both cohorts'. Since the paper's premise and its simulation target are sex differences in QRSd, this misreported result must be corrected and reconciled with Table 1.","section":"Results, 'Basic patient characteristics'"},{"comment":"Figure 2 and the accompanying Results text report that under normal conduction scenarios 'none of their simulated QRSds surpassed 150 ms under normal conduction scenarios, regardless of LBBB or normal activation' (LBBB mean 113.2±13.3 ms), but the Discussion states that '6 of 17 cases simulated under LBBB and normal conduction exceeded the 150ms threshold'. A mean of 113.2±13.3 ms makes 6 of 17 values above 150 ms statistically implausible, so one of these passages is erroneous; the discrepancy is load-bearing because it is used to infer that real CRT candidates combine LBBB with slow conduction.","section":"Results, 'Which simulated pathological scenarios will meet the current criteria of CRT' versus Discussion"},{"comment":"The simulated QRSd are systematically shorter than the measured QRSd in the same cohorts: the healthy normal-activation simulation mean is 78.0 ms versus measured values of 84–92.5 ms (Table 1), and in the validation subset of 17 real CRT-indicated patients, the LBBB-plus-slow-conduction scenario brings only 88% (15/17) of patients above 150 ms although all 17 were included on the basis of measured QRSd ≥150 ms. The fixed conduction velocities (0.65 and 0.40 m/s), the subendocardial layer parameters, and the 0.15 spatial-velocity detection threshold are taken from literature medians and are never calibrated to the measured QRSd of these cohorts, and no sensitivity analysis is provided. Since the recommended height-indexed thresholds (e.g., 0.86 ms/cm) are absolute, a systematic bias in simulated QRSd—particularly one that interacts with sex or heart size—would directly shift the recommended thresholds; the authors should quantify this offset and its sex dependence, or temper the threshold-level claims.","section":"Methods, 'Electrophysiological (EP) simulations' and Results validation subset"},{"comment":"Because the same conduction velocities and the same fascicular activation model are applied to both sexes, all simulated sex differences in QRSd arise from anatomical heart-size differences, and height is used as a heart-size proxy. The observation that height-indexed QRSd 'effectively resolved sex differences' is therefore substantially built into the model's construction, and the supporting premise that real sex differences are entirely explained by heart size rests on prior work (ref 5, a medRxiv preprint). The ROC superiority of height indexing over LVEDV/LV-mass indexing is an emergent result, but the 'resolved sex disparities' claim should be reframed as a model prediction, and the robustness to the alternative hypothesis should be tested, for example by simulating a scenario with sex-specific conduction velocity or by directly comparing predicted versus measured sex differences in QRSd within the same cohorts.","section":"Methods, 'Electrophysiological (EP) simulations' and Discussion"},{"comment":"The primary outcome is classification of simulated LBBB versus non-LBBB status, and selection of non-LBBB patients is labelled as a 'false positive' rate throughout. In current guidelines, however, non-LBBB patients with QRSd ≥150 ms are also CRT candidates (e.g., Class IIa in ESC 2021), so selecting non-LBBB patients with slow conduction is not necessarily a clinical error. The comparative claim that height-indexed criteria are 'the most reliable predictor for CRT patient selection' depends on this framing; the authors should either justify it with reference to expected benefit by pathology or soften the conclusion to LBBB-stratification performance.","section":"Results, 'LBBB and non-LBBB Patient stratification'"}],"minor_comments":[{"comment":"The IHD male QRSd entry reads '106.9.2±22.5' and should be '106.9±22.5'.","section":"Table 1"},{"comment":"The sentence describing indexed-criteria selection rates begins '(range: [21.7–99.9%])' with an unmatched bracket, and the units '0.9 ms/m' and '0.86 ms/m' should be 'ms/cm' for consistency with the 0.8 and 0.9 ms/cm cutoffs defined earlier in the same section.","section":"Results, 'LBBB and non-LBBB Patient stratification'"},{"comment":"The sentence on sex-specific optimal thresholds states that male QRSd thresholds differ from female thresholds 'ranging from 7ms for healthy anatomies+normal conduction to 13 ms for healthy anatomies+normal conduction', naming the same scenario twice; one endpoint should refer to a different scenario.","section":"Results, 'Classification performance'"},{"comment":"The IHD cohort contains only 45 females (12.5%), so the sex-stratified selection-rate comparisons and P-values in the heatmaps carry wide confidence intervals; reporting these intervals (or at least noting the small denominator) would strengthen the interpretation.","section":"Table 1 and heatmap analyses"},{"comment":"Reference 5, which underpins the load-bearing assumption that sex differences in QRSd are entirely explained by heart size, is cited as a medRxiv preprint; a peer-reviewed version should be cited if available.","section":"References"},{"comment":"The statement that the modelling approach is 'free from biases inherent in clinical trial recruitment' overstates the case, because the UK Biobank cohort and the single-center IHD cohort carry their own sampling and referral biases; a more careful wording would acknowledge these residual selection effects.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the in silico framework is a useful contribution, but the internal inconsistencies between the Results text and Table 1, and between the Results and Discussion regarding the n=17 validation subset, should be resolved before further review; these look like reporting errors rather than misconduct. I would also encourage the editor to ask the authors to strengthen the evidence for the load-bearing assumption that sex differences in QRSd are entirely explained by heart size, since it currently rests on a preprint (ref 5). The novelty relative to ref 5 is real but modest; the paper's main value is the full-spectrum evaluation of indexed criteria in large virtual cohorts, which could be a reproducible template for similar guideline questions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Rough take: this paper deserves a serious referee, but the headline result — height-indexed QRSd resolves sex disparities and beats structural indexing — is largely a consequence of the model's construction, not an empirical discovery.\n\nWhat's genuinely useful: the systematic comparison of conventional, volume-indexed, mass-indexed, and height-indexed QRSd across four pathological scenarios in ~3000 virtual patients is new. The finding that height-indexing outperforms LVEDV/LV-mass indexing is a solid comparative result within the model. The authors also make a real attempt to validate against the 17 real CRT patients, and their observation that most CRT-indicated patients need both LBBB and slow conduction to reach 150 ms in simulation is thought-provoking.\n\nThe central weakness is that the sex-disparity result is built in. The model uses identical conduction velocities and fascicular timing for both sexes, so simulated sex differences in QRSd arise entirely from heart size. Height is a proxy for heart size. It is therefore a direct consequence of the model's assumptions that height-indexing reduces sex disparities. The ROC comparisons between index schemes are still informative, but the 'resolves sex differences' claim is not an independent test of the hypothesis.\n\nSecond, calibration is off. Simulated normal QRSd (~78 ms) is systematically shorter than measured QRSd (84–92 ms) in the same cohorts. Only 88% of real CRT patients cross 150 ms even in the most favorable simulated scenario. The authors use this to infer that real CRT patients must have slow conduction, which is reasonable, but it also means absolute thresholds and selection rates are not directly transferable to clinics.\n\nThird, the ground truth for LBBB is simulated (RV-only activation), not clinical ECG morphology. The classification performance is partly circular: you're testing whether the model's QRSd can discriminate the model's own LBBB from its normal activation. That's fine for internal comparison, but it doesn't tell you how the criteria behave on real patients.\n\nFinally, the simulation pipeline is not yet available and the electrophysiology solver is proprietary, so the main result is not independently reproducible at this stage.\n\nI'd send this to peer review. The comparative framework is valuable, the limitations are identifiable, and the authors are transparent about model parameters. But a referee should push for: (1) explicit acknowledgment that the sex-disparity result is assumption-driven, (2) a calibration step or sensitivity analysis around the QRSd offset, and (3) a clear statement of what would be needed to validate height-indexed thresholds clinically. Conditional accept, leaning major revision.","headline":"Worth refereeing, but the height-indexed advantage is partly built into the model's assumptions; treat the clinical claims as hypothesis-generating.","tokens_in":13408,"tokens_out":3933,"would_cite":false,"duration_ms":34735,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Height-indexed QRS duration tops CRT selection criteria in virtual trial","keywords":["cardiac resynchronization therapy","QRS duration","sex differences","computer simulation","virtual cohort","heart failure","left bundle branch block","patient selection"],"falsifier":"A prospective observational study in a large mixed-sex heart failure cohort that measures QRSd, height, LVEDV, and CRT outcomes would falsify the claim if height-indexed QRSd did not outperform conventional QRSd in predicting LBBB status (by ECG morphology) or in predicting CRT response, or if it selected more non-responders. Specifically, if the height-indexed threshold's AUC were not at least as high as unindexed QRSd, the central superiority claim would fail.","tokens_in":12382,"feed_emoji":"🫀","tokens_out":5003,"duration_ms":40561,"temperature":0.7,"pith_summary":"This paper uses computational 'virtual cohorts' to test whether current QRS-duration criteria for cardiac resynchronization therapy (CRT) unfairly exclude women. By simulating left bundle branch block and slow conduction in thousands of patient-specific heart models, the authors show that the standard 150 ms threshold selects fewer women with LBBB and more men without LBBB. They then compare alternative criteria that index QRS duration by heart volume, mass, or height. The central claim is that height-indexed QRS duration removes the sex imbalance while keeping the number of non-LBBB patients selected low, making it a more equitable and potentially more accurate basis for CRT guidelines.","feed_headline":"Height-indexed QRS duration tops CRT selection criteria in virtual trial","feed_subtitle":"Simulating 3,000 hearts shows height-indexed QRS duration removes sex bias without over-selecting non-LBBB patients.","key_machinery":"The machinery is a population-based in silico trial: patient-specific biventricular anatomical models (from CMR images) are activated with a fascicular model of His-Purkinje activation, and the reaction-eikonal model in CARP computes activation times and 12-lead ECGs from which QRSd is derived. Four scenarios (normal/LBBB activation × normal/slow conduction) are simulated per heart, and the resulting QRSd distributions are used to evaluate the sensitivity and specificity of conventional and indexed thresholds via ROC analysis. This allows controlled, mechanistic assessment of how sex, heart size, and pathological substrate affect QRSd-based patient selection, free from the selection bias of retrospective clinical cohorts.","core_discovery":"On the paper's own terms, the discovery is that height-indexed QRS duration (QRSd) is the best performing QRSd criterion for identifying LBBB patients who should receive CRT, across both healthy and ischemic heart disease anatomies and with or without slow conduction. Simulated QRSd values were generated for 2627 UK Biobank healthy participants and 359 ischemic heart disease patients under four pathological scenarios. Conventional thresholds (120, 130, 150 ms) under-selected LBBB females and over-selected non-LBBB males; indexing by LVEDV or LV mass reduced sex disparities but inflated false-positive selection; indexing by height (cutoffs 0.8 and 0.9 ms/cm) resolved the sex differences and maintained low non-LBBB selection rates. ROC analysis corroborated this: height-indexed QRSd outperformed volume- and mass-indexed criteria in all scenarios and matched or slightly exceeded conventional QRSd, with sex-specific optimal thresholds differing by only 0.01 ms/cm versus 7–13 ms for unindexed QRSd.","pith_inferences":["A testable extension: the same height-indexed logic might apply to other ECG intervals, such as QRS area or QT corrections, reducing sex bias in other device indications.","The simulation's reliance on fixed conduction velocities rather than patient-calibrated values is a key uncertainty; if true conduction velocity differs by sex beyond heart size, the height thresholds could be miscalibrated, and prospective clinical measurement of QRSd/height would settle this.","The comparison did not include a formal cost-benefit analysis or account for non-ischemic etiologies; height-indexed cutoffs may need adjustment in non-ischemic cardiomyopathy or in populations with different body-size distributions.","If height-indexed criteria were validated clinically, it would shift CRT guidelines from a single QRSd cutoff to a sex-independent one, potentially harmonizing US and European LBBB definitions."],"forward_implications":["If height-indexed QRSd thresholds (e.g., 0.8–0.9 ms/cm) were adopted, more women with LBBB would receive CRT at QRS durations below 150 ms, potentially improving their outcomes.","Because height is already measured routinely, height-indexed criteria could be implemented in clinical practice at negligible cost.","The finding suggests that current CRT guidelines' sex differences in response may partly reflect biased selection rather than biology alone.","The virtual cohort approach could be extended to test other patient-selection criteria (e.g., imaging-based dyssynchrony) before running expensive trials.","The simulated result that most real CRT candidates are consistent with LBBB plus slow conduction could justify targeting both substrates in future device therapy."],"supporting_citations":[{"why":"Supplies the anatomical modelling pipeline and the prior finding that sex differences in QRSd are explained by heart size without intrinsic conduction velocity differences.","marker":"[5]"},{"why":"Provides clinical evidence of sex-specific CRT response and the LVEDV/LV mass cutoffs used for indexed QRSd criteria.","marker":"[4]"},{"why":"Patient-level meta-analysis linking sex, body size, and CRT benefit; the source of the height-indexed cutoffs (0.8 and 0.9 ms/cm).","marker":"[12]"},{"why":"Current ESC guidelines defining the 150 ms QRSd threshold and LBBB criteria that this study evaluates and compares against.","marker":"[8]"},{"why":"HRS/APHRS/LAHRS guideline supplying the additional QRSd thresholds (130 and 120 ms) used in the comparisons.","marker":"[9]"},{"why":"Describes the reaction-eikonal model used to simulate activation times and ECGs, the core simulation engine.","marker":"[20]"},{"why":"Provides the His-Purkinje fascicular activation model that defines the normal and LBBB activation patterns.","marker":"[21]"}],"fun_headline_variants":["Height-indexed QRSd beats sex bias in CRT selection","Virtual hearts show height-indexed QRSd best for CRT","Height indexing removes sex bias in CRT criteria","Simulated hearts favor height-indexed QRSd for CRT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulated QRSd values, generated with fixed conduction velocities and a fascicular activation model, are accurate and sex-fair enough to support threshold-level clinical recommendations, even though the simulations are not calibrated to each patient's measured QRSd.","fun_headline_variants_meta":{"raw":{"variants":["Height-indexed QRSd beats sex bias in CRT selection","Virtual hearts show height-indexed QRSd best for CRT","Height indexing removes sex bias in CRT criteria","Simulated hearts favor height-indexed QRSd for CRT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1339,"prompt_tokens":969,"completion_tokens":370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":303}},"tokens_in":585,"tokens_out":370,"duration_ms":3605,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:12:42.470615+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective observational study in a large mixed-sex heart failure cohort that measures QRSd, height, LVEDV, and CRT outcomes would falsify the claim if height-indexed QRSd did not outperform conventional QRSd in predicting LBBB status (by ECG morphology) or in predicting CRT response, or if it selected more non-responders. Specifically, if the height-indexed threshold's AUC were not at least as high as unindexed QRSd, the central superiority claim would fail.","supporting_citations":[{"cited_title":"medRxiv [Internet] 2023; :2023.12.05.23299435","cited_arxiv_id":null,"evidence_quote":"Supplies the anatomical modelling pipeline and the prior finding that sex differences in QRSd are explained by heart size without intrinsic conduction velocity differences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides clinical evidence of sex-specific CRT response and the LVEDV/LV mass cutoffs used for indexed QRSd criteria."},{"cited_title":"Heart Rhythm Elsevier B.V., 2024","cited_arxiv_id":null,"evidence_quote":"Patient-level meta-analysis linking sex, body size, and CRT benefit; the source of the height-indexed cutoffs (0.8 and 0.9 ms/cm)."},{"cited_title":"Eur Heart J","cited_arxiv_id":null,"evidence_quote":"Current ESC guidelines defining the 150 ms QRSd threshold and LBBB criteria that this study evaluates and compares against."},{"cited_title":"Heart Rhythm Elsevier B.V., 2023; 20:e17–e91","cited_arxiv_id":null,"evidence_quote":"HRS/APHRS/LAHRS guideline supplying the additional QRSd thresholds (130 and 120 ms) used in the comparisons."}],"review_version":1}