{"id":"bdcd39a9-24ce-4bc1-8c4d-6f6de6d05e9d","arxiv_id":"2505.21019","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An automatic pipeline converts cine cardiac MRI into biventricular meshes for over 50,000 UK Biobank participants and releases 1,423 demographic representative meshes with fibers and coordinates.","lead":"Researchers built an open-source pipeline that turns routine cardiac MRI scans from about 55,000 UK Biobank participants into 3D heart meshes, and they release 1,423 representative meshes covering different ages, sexes, and body sizes. If the meshes are reliable, they give the cardiac simulation community a common, large-scale starting point for digital twin studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fixed 3 mm RV free-wall offset makes every biventricular mesh partly synthetic, and unvalidated demographic variation in true RV wall thickness could bias the 1,423 representative meshes.","rationale":"The paper is honest and the scale claim is supported by Table S2: 54,926 participants pass conversion, and 51,358 ED meshes are built. The segmentation Dice scores are solid, and the demographic trends are biologically plausible. Among the reader's conditions, the fixed 3 mm RV wall thickness is the most central to the paper's core promise of patient-specific biventricular meshes. The other objections are real but secondary: missing model downloads are a timing issue, ES frame selection does not affect the ED-based representative cohort, and the LV mass bias is disclosed and could be a metric-integration artifact. The RV offset is different because it injects a non-image-derived geometry into every mesh, and the existing QC cannot catch it. My stress test therefore agrees with the reader's weakest-assumption identification. The recommended verdict is unchanged: conditional acceptance, pending either a sensitivity analysis showing that 3 mm is neutral or external validation of RV wall thickness.","tokens_in":15321,"tokens_out":7981,"duration_ms":97139,"concrete_test":"Conduct a sensitivity sweep with the released pipeline: take a stratified subsample of about 500 participants spanning the sex/age/BMI grid, rebuild RV epicardia with offsets of 2, 3, 4, and 5 mm, recompute the bin-averaged representative meshes, and report RV myocardial mass, RVEDV, and the age/BMI regression slopes at each offset. If the 3-to-4 mm change shifts mean RV mass by more than 10% or changes the sign or significance of any demographic slope, the fixed 3 mm value is load-bearing and must be validated against measured RV wall thickness (e.g., from high-resolution CT or expert manual contours); if the outputs are insensitive across 2-5 mm, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is in the Methods section 'Finite element mesh construction': because the RV myocardium is not segmented, the pipeline creates the RV epicardium by extending the RV endocardium outward by a fixed 3 mm. The paper's Limitations concede this 'can lead to inaccurate R V function estimation.' This is not a minor detail: the central deliverable is a cohort of patient-specific biventricular digital-twin meshes, yet the entire RV free wall of every mesh is synthetic. The 1,423 released representative meshes and their fibers all inherit this same offset, so any simulation involving RV wall stress, mass, or conduction uses geometry that is not derived from the participant's images. The problem is compounded because the QC step (75th percentile plus 1.5 IQR on mesh-versus-segmentation volume differences) cannot detect a uniform offset error: the offset is applied before the mesh phenotypes are computed, and the segmentation-derived phenotypes themselves contain no RV epicardium. If true RV wall thickness varies meaningfully across sex, BMI, and age, which is plausible, the demographic comparisons in Fig 4 and the representative meshes themselves will carry a systematic RV bias. Published normal RV free-wall thickness is often in the 3-5 mm range, so the chosen 3 mm is not self-evidently neutral. The authors do expose the assumption and make it a variable, but no sensitivity analysis or external validation of the 3 mm value is reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an automated pipeline that converts raw short- and long-axis cine CMR images from UK Biobank into biventricular tetrahedral meshes with fibers and universal ventricular coordinates, using nnU-Net segmentation, atlas-based surface reconstruction, and meshtool volume meshing. The pipeline is applied to 54,926 participants, of whom 46,917 pass quality control and contribute to 1,423 demographic-bin representative meshes spanning sex, age, and BMI. The authors validate segmentation Dice scores and derived phenotypes against manual segmentations and previously published values, and report demographic trends (sex differences, age-related decline, BMI-related increase in volumes and mass) that are consistent with the literature.","tokens_in":15575,"tokens_out":4015,"duration_ms":42325,"significance":"The resource is potentially valuable: the code is open source, the pipeline is demonstrated at a scale (>50k) not previously shown for mesh generation, the segmentation validation is technically sound (Dice mostly 0.88–0.98, phenotype errors comparable to prior U-Net work), and the representative meshes with fibers and UVCs would fill a genuine gap in publicly available cardiac digital twin cohorts. The paper is also commendably transparent about its limitations. If the systematic mesh biases and the fixed RV free-wall assumption are addressed or clearly scoped, this could become a standard baseline for population-scale cardiac digital twinning.","major_comments":[{"comment":"The RV epicardium is generated by uniformly offsetting the RV endocardium by 3 mm (Methods, 'Finite element mesh construction'; Limitations). This means every biventricular mesh and all 1,423 representative meshes carry a synthetic RV free wall not derived from the participant's images. Because the QC step compares mesh-derived to segmentation-derived volumes and mass, and the segmentations contain no RV epicardium, a uniform offset error is invisible to QC. The authors cite references [42,43] for the 3 mm value but report no sensitivity analysis or external validation of this parameter, and published normal RV free-wall thickness spans roughly 3–5 mm. For a resource intended for electro-mechanical simulations of RV stress, mass, and conduction, this is a load-bearing limitation; please either add a sensitivity analysis over the physiological range of RV wall thickness, validate against a dataset with RV wall measurements, or revise the 'patient-specific' claim to specify that the LV is image-derived while the RV free wall is a fixed-thickness estimate.","section":"Methods, 'Finite element mesh construction'"},{"comment":"Table 2 shows systematic mesh-versus-segmentation differences: LVEDV −7.7%, LVESV −8.0%, RVEDV −8.0%, RVESV −8.6%, and LV mass +18.6% relative difference (nnUNet–Mesh). The Discussion acknowledges this bias and states it is 'unclear' which phenotype should be preferred, but the bias propagates directly into the representative meshes and the demographic regressions in Fig 4 (e.g., LV mass 133.6±14.8 g for males, above the segmentation-derived values in Table 2). Since the central deliverable is a cohort for quantitative twin studies, the manuscript should either provide a calibration/correction for mesh-derived phenotypes, report which downstream quantities are robust to this bias (e.g., EF appears partly protected), or explicitly flag that representative-mesh phenotypes are not interchangeable with segmentation-derived clinical measurements.","section":"Table 2"},{"comment":"The representative meshes are binned averages with no report of within-bin shape or phenotype variability, and the QC threshold (75th percentile plus 1.5 IQR across three phenotypes) and the minimum bin size of three are arbitrary choices with no sensitivity analysis. For the claim that these 1,423 meshes are 'representative' of demographic groups, please report the within-bin dispersion (e.g., standard deviation or percentiles of volumes and mass) and test whether the demographic trends in Fig 4 persist under alternative QC and binning choices. This would substantiate the representativeness claim and help users understand how much individual variation is lost by averaging.","section":"Methods, 'Representative mesh generation for different sex, age and BMI groups'"}],"minor_comments":[{"comment":"'Steady state free precision' should be 'steady-state free precession'.","section":"Introduction"},{"comment":"The Abstract says pre-trained networks and representative meshes 'will be made available soon', while the Discussion states they 'are made publicly available'; please reconcile this and give a concrete availability date or repository status for the meshes, fibers, and trained networks.","section":"Abstract and Discussion"},{"comment":"Table S2 reports a 'contour quality-control' step that removes 2,377 subjects, but the main text does not describe this QC step; please document it in Methods.","section":"Table S2 and Methods"},{"comment":"The LAX segmentation validation uses only 50 test participants; given the importance of LAX landmarks (valve planes, apex) for mesh construction, adding confidence intervals or a comparison on an independent LAX dataset would strengthen the validation.","section":"Segmentation validation"},{"comment":"The p-values and regression coefficients in Fig 4 are reported without confidence intervals; adding CIs would improve interpretability of the demographic regressions.","section":"Fig 4"},{"comment":"The abbreviation 'UVC' is used without expansion; please define it at first use.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is the largest public biventricular mesh cohort by an order of magnitude, and the pipeline is open, reproducible in principle, and validated against manual segmentations and external phenotypes. The segmentation Dice scores are strong (0.86–0.98), and the demographic trends in the representative meshes match independent UKBB and non-UKBB studies, which gives the cohort external credibility. The 1,423 meshes with fibers and UVCs, once released, will be a useful resource for statistical shape analysis and electro-mechanical simulation.\n\nThe soft spots are real but not fatal. The fixed 3 mm RV free-wall offset is a disclosed assumption, but it does mean every RV epicardium and fiber field in the released cohort is partly synthetic. The authors do not report a sensitivity analysis on that 3 mm value, and if true RV wall thickness varies with sex/BMI/age, the demographic comparisons in Fig 4 could carry a systematic RV bias. That is the paper's main vulnerability, and it deserves a sensitivity analysis or at least a clear statement about which applications are unaffected. Also, the meshes are not actually public yet – “will be made available soon” – so the central artifact is not yet independently usable. The ES frame selection is unquantified, and the mesh vs. segmentation volume/mass biases (LV mass ~20% higher, volumes 4–7% lower) are acknowledged but not resolved. These are conditions, not rejections.\n\nI agree with the reader's CONDITIONAL verdict. The paper deserves a serious referee because the resource is important and the methods are mostly transparent. I would recommend requiring a sensitivity analysis on RV wall thickness and actual release of the meshes before acceptance.","headline":"A genuinely useful large-scale resource from an honest pipeline, with a disclosed but unresolved RV wall-thickness assumption that limits but does not sink it.","tokens_in":16201,"tokens_out":2943,"would_cite":true,"duration_ms":27745,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T13:42:43.072936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}