REVIEW 3 major objections 4 minor 22 references
This paper argues that an apparent ejection-fraction calibration gain in cardiac digital twins is a measurement-convention artifact, not a genuine model improvement.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Phase-conditioned EF model gains on CAMUS are explained by a single-plane vs biplane measurement convention gap, not by improved EF calibration.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection The CAMUS phase-conditioning EF gain is almost certainly a plane-convention artefact, and the paper proves it convincingly on CAMUS itself; the EchoNet replication and one 'indistinguishable' claim need tightening, but this deserves a serious referee. the 3 major comments →
When Measurement Conventions Masquerade as Calibration Gains in Cardiac Digital Twins
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is that the phase-conditioning EF calibration gain observed on CAMUS is not a genuine observation-operator improvement. When EF is computed from ground-truth masks using the same single-plane method-of-disks extractor applied to predictions, the four front-ends (B1, B2, P1, P2) are statistically indistinguishable on extractor-matched error (MAE_gt), and the baseline models are not miscalibrated once the measured convention offset is removed. The apparent gain is instead explained by a +6.30 EF-point systematic difference between single-plane ground-truth EF and CAMUS's biplane clinical EF reference. A plane-aligned replication on EchoNet-Dynamic confirms the explanation
What carries the argument
The measurement-convention offset: the difference between single-plane ground-truth EF computed by the same method-of-disks extractor and the dataset's clinical biplane EF reference. This offset (+6.30 EF points on CAMUS) is the mechanism that makes baseline models look miscalibrated and phase-conditioned models look calibrated. The matched-reference diagnostic (EF MAE_gt, computed from ground-truth masks with the same extractor) is what exposes the illusion, by showing flat segmentation-error across front-ends while clinical-reference error varies.
Load-bearing premise
The EchoNet-Dynamic replication is a valid plane-aligned test: its EF labels, though aligned to the four-chamber plane, may still differ from the extractor in ED/ES frame selection, contour style, preprocessing, and calculation, and those differences could hide a genuine phase-conditioning effect.
What would settle it
A process-matched replication: construct a dataset where four-chamber EF labels are computed with the same ED/ES selection, contour style, and method-of-disks calculation as the extractor, then test whether phase-conditioned front-ends improve extractor-matched EF MAE_gt. If P1/P2 beat B1/B2 on MAE_gt under those conditions, the convention-offset explanation is incomplete or wrong.
If this is right
- EF calibration studies must report extractor-matched error (MAE_gt) alongside clinical-reference error; a flat MAE_gt with varying MAE_clin is a red flag for a convention confound.
- CAMUS baseline EF bias is consistent with zero after subtracting the convention offset, so the U-Net front-ends were never miscalibrated in the single-plane sense.
- Plane-aligned evaluation on EchoNet-Dynamic shows B2 (adjacent-frame-smoothed) is the strongest deployable front-end, with streaming inference, no invalid EF estimates, and the best clinical EF MAE.
- Downstream twin-state errors (stroke volume, cardiac output, contractility) are much smaller when the convention gap is removed; under plane-aligned conditions, haemodynamic errors fall to roughly 3–4%.
- The paper's Convention-Aware EF Audit protocol—compute MAE_gt, characterize the convention offset, check seed consistency, add a safety gate, and replicate under an aligned reference—provides a template for auditing future EF observation operators.
Where Pith is reading between the lines
- The proportional convention gap (slope 0.613) implies that a constant +6.30-point correction leaves structured residual error, especially at high EF; a view-aware or linear correction, validated prospectively, would likely be needed for patient-level EF adjustment.
- The EchoNet replication is strong evidence against phase conditioning, but it remains process-mismatched; a fully process-matched test with identical ED/ES selection, contour style, and calculation on four-chamber labels would directly isolate how much of the reversal is due to the plane alone.
- The same convention-mismatch mechanism may affect other benchmarks that mix single-plane and biplane references, and other physiological ratios computed from segmentation (e.g., stroke volume, regurgitant fractions), where small coherent contour shifts amplify into apparent model differences.
- If the audit protocol were widely adopted, some previously published calibration gains on CAMUS-derived tasks would likely require re-evaluation with matched-reference EF.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper audits the calibration of four U-Net-based ejection-fraction (EF) front-ends (frame-wise B1, smoothed B2, phase-conditioned P1, and P2 with additional consistency) on two echocardiography datasets. On CAMUS, phase-conditioned models appear to reduce bias relative to the dataset's biplane clinical EF. The authors argue this is a measurement-convention artefact: when EF is computed from ground-truth masks with the same single-plane extractor, MAE_gt is claimed to be statistically indistinguishable across front-ends, and a measured +6.30 EF-point offset between single-plane ground-truth EF and CAMUS biplane clinical EF accounts for the baseline clinical bias. A prespecified EchoNet-Dynamic replication, in which the clinical labels are four-chamber (plane-aligned) but not otherwise process-matched, removes the positive baseline bias and reverses the CAMUS ranking. The paper contributes a Convention-Aware EF Audit protocol and reports downstream twin-state, conformal, and EF-stratum analyses.
Significance. If the conclusions hold, the paper makes a useful methodological contribution: calibration gains for image-to-measurement operators must be evaluated under a matched reference convention, and otherwise benchmark artefacts can masquerade as model improvements. The strengths are the clean decomposition of baseline bias into convention and segmentation components, the use of a prespecified external replication, and the release of code/data. The main limitations are the small CAMUS sample (n=49-50) and the acknowledged process mismatch in the EchoNet replication, which leaves the external evidence partly equivocal. The downstream analyses are clearly labeled as simplified.
major comments (3)
- [§3.2, Table 1] The claim that 'EF MAE_gt is statistically indistinguishable across front-ends' is load-bearing (used in the abstract and Section 3.2) but no statistical test is reported. Section 3.1 states that paired Wilcoxon tests, bootstrap CIs, Cohen's d_z, and Holm correction are used, yet none of the corresponding p-values or intervals appear in the text or tables. Table 1 reports mean±std across three seeds, not patient-level dispersion; with P1 MAE_gt=10.62 vs B2=8.08 and Bias_gt=-7.56 vs -1.37, the claim requires the missing patient-level test results. If P1 is significantly worse, the paper should say that; the main conclusion (no genuine improvement) still stands, but 'statistically indistinguishable' would be incorrect and should be revised. Moreover, failure to reject a null hypothesis is not evidence of equivalence; if the tests are underpowered, the phrase should be 'no significant diffe
- [§3.3, Table 2] The statement that the +6.30 EF-point convention offset 'explains nearly all baseline bias' is based on point estimates with wide intervals. The convention offset CI is [+3.69,+8.91], and the residual segmentation component CI is [-4.08,+1.14]. The point estimate of the residual is -1.47, which is about 30% of the B1 clinical bias (+4.83); the CI lower bound allows a residual of -4.08. Given the small sample (n=49), the authors should discuss the uncertainty more explicitly and avoid the phrase 'nearly all' without qualification. In addition, the measured offset is not purely a plane effect: it also includes differences in contour style, ED/ES selection, and calculation between the single-plane extractor and the CAMUS biplane clinical reference. The paper should either call it a 'reference-convention offset' or identify which components are separable.
- [§3.4, Tables 3-4] The EchoNet replication is presented as the decisive external confirmation ('reverses the CAMUS ranking'; 'expected when the dominant plane mismatch is removed'), but the paper concedes that EchoNet is 'plane-aligned, but not process-matched' (Section 3.4 and Limitations: ED/ES selection, contour style, preprocessing, and calculation may all differ). The prespecified criteria in Table 4 assume that a genuine CAMUS effect would require positive baseline bias and phase-model improvement on EchoNet; this only follows if the plane mismatch is the only systematic difference. Without a quantitative bound on the process-mismatch offset, the observed negative baseline bias on EchoNet could be a separate convention offset rather than evidence that the plane mismatch was removed. The authors should either quantify this offset (e.g., apply the extractor to EchoNet expert tracings and compare with E
minor comments (4)
- [Eq. (1)] The volume formula writes `cEF = [EDV− [ESV / [EDV ×100`; the bracket notation is confusing. Please use \widehat{EDV} or \hat{EDV} and add parentheses around (EDV-ESV)/EDV.
- [Fig. 1(c)] The caption reports 'Mean = -6.30 EF pts' for 'biplane clinical - GT single', while Table 2 reports +6.30 for 'single-plane GT - biplane clinical'. Please make the sign convention consistent.
- [§3.2, Table 1] The claim 'Baselines are seed-consistent, whereas phase-conditioned models vary more' is not uniformly supported by the seed std columns. B2 (baseline) has Bias_clin std 1.61 and Bias_gt std 0.93; P2 has Bias_clin std 1.31 and MAE_gt std 1.26; P1 has MAE_gt std 0.48, lower than both baselines. Please qualify this statement.
- [Table 3] The caption uses MAE_gt, MAE_clin, and Biasclin without defining them; refer readers to Section 2.2.
Circularity Check
No significant circularity: the convention-offset decomposition and EchoNet replication are independent evidence.
full rationale
The paper's central claim is that the apparent CAMUS phase-conditioning EF calibration gain is a measurement-convention artefact, not a genuine model improvement. This claim is supported by two independent lines of evidence. First, the CAMUS convention-offset analysis (Section 3.3, Table 2) measures the single-plane-GT-to-biplane-clinical EF gap (+6.30 EF points) directly from ground-truth masks using the same extractor, independent of any front-end model output. The residual segmentation component is then the algebraic difference between the observed B1 clinical bias and this measured offset; it is not a fitted value used to derive the main result. The post-hoc OLS slope and intercept in Table 2 are explicitly labeled 'Post-hoc OLS; not a validated correction' and are not load-bearing for the conclusion. Second, the EchoNet replication (Section 3.4, Table 3) is an external, prespecified test on a different dataset with plane-aligned labels. The paper explicitly acknowledges that EchoNet is 'plane-aligned, but not process-matched' (Introduction and Section 3.4), which is a stated limitation about residual confounds, not a circular reduction. No load-bearing step reduces to its own input by construction, and no self-citation is used to justify the core premise. The derivation chain is self-contained: the convention offset is measured, not fitted, and the replication is a genuine external check. The skeptical concern about EchoNet process mismatch is a validity limitation, not circularity, and is already disclosed in the paper.
Axiom & Free-Parameter Ledger
free parameters (3)
- Single-plane to biplane convention offset (used as correction) =
+6.30 EF points (95% CI 3.69-8.91)
- Post-hoc OLS slope =
0.613 (95% CI 0.47-0.76)
- Post-hoc OLS intercept =
+13.34 EF points
axioms (5)
- domain assumption Single-plane method-of-disks EF (Eq 1) is a valid observation operator for EF from four-chamber clips.
- domain assumption CAMUS biplane clinical EF is the reference convention that causes the apparent gain.
- domain assumption EchoNet-Dynamic clinical EF labels are plane-aligned to the four-chamber extractor (though not process-matched).
- domain assumption Prespecified checks and exclusions were actually fixed before EchoNet evaluation.
- domain assumption Simplified lumped haemodynamic propagation (SV, CO, Ees) captures order-of-magnitude twin-state errors.
Cite this review
Pith. "Pith review of When Measurement Conventions Masquerade as Calibration Gains in Cardiac Digital Twins." pith.science (2026). https://pith.science/paper/MAYRO7VT
@misc{pith2026260801602,
author = {Pith},
title = {Pith review of: When Measurement Conventions Masquerade as Calibration Gains in Cardiac Digital Twins},
year = {2026},
howpublished = {\url{https://pith.science/paper/MAYRO7VT}},
note = {Machine review of arXiv:2608.01602}
}
read the original abstract
Cardiac digital twins convert clinical images into physiological measurements through observation operators, yet calibration studies often assume a fixed reference convention. Across four shared-backbone echocardiographic EF front-ends, phase conditioning appears to remove CAMUS baseline bias. Matched-reference analysis rejects this gain: singleplane ground-truth EF error is statistically indistinguishable across models, while single-plane ground-truth EF exceeds CAMUS biplane clinical EF by +6.30 points, explaining nearly all baseline bias. A prespecified EchoNet-Dynamic replication, with released data and our extractor aligned to the apical four-chamber plane, removes baseline overestimation and reverses the CAMUS ranking. We also quantify haemodynamic effects, conformal residual-width budgets, and EF-stratum changes, yielding a Convention-Aware EF Audit protocol that separates genuine observation operator calibration from measurement artefacts. GitHub: EjectionFraction-Bias-in-Cardiac-Digital-Twin.git
Figures
Reference graph
Works this paper leans on
-
[1]
Foun- dations and Trends in Machine Learning16(4), 494–591 (2023)
Angelopoulos, A.N., Bates, S.: Conformal prediction: A gentle introduction. Foun- dations and Trends in Machine Learning16(4), 494–591 (2023). https://doi.org/ 10.1561/2200000101
-
[2]
The Lancet327(8476), 307–310 (1986)
Bland, J.M., Altman, D.G.: Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet327(8476), 307–310 (1986). https://doi.org/10.1016/S0140-6736(86)90837-8
-
[3]
Medical Image Analysis100, 103361 (2025)
Camps, J., Wang, Z.J., Doste, R., Berg, L.A., Holmes, M., Lawson, B., Tomek, J., Burrage, K., Bueno-Orovio, A., Rodriguez, B.: Harnessing 12-lead ECG and MRI data to personalise repolarisation profiles in cardiac digital twin models for enhanced virtual drug testing. Medical Image Analysis100, 103361 (2025). https: //doi.org/10.1016/j.media.2024.103361
arXiv 2025
-
[4]
European Heart Journal41(48), 4556– 4564 (2020)
Corral-Acero, J., Margara, F., Marciniak, M., Rodero, C., Loncaric, F., Feng, Y., Gilbert, A., Fernandes, J.F., Bukhari, H.A., Wajdan, A., et al.: The ‘digital twin’ to enable the vision of precision cardiology. European Heart Journal41(48), 4556– 4564 (2020). https://doi.org/10.1093/eurheartj/ehaa159
-
[5]
Dezaki, F.T., Dhungel, N., Abdi, A.H., Luong, C., Tsang, T., Jue, J., Gin, K., Hawley, D., Rohling, R., Abolmaesumi, P.: Deep residual recurrent neural networks forcharacterisationofcardiaccyclephasefromechocardiograms.In:DeepLearning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. LectureNotesinComputerScience,vol.10553,p...
-
[6]
Circulation145(18), e895–e1032 (2022)
Heidenreich, P.A., Bozkurt, B., Aguilar, D., Allen, L.A., Byun, J.J., Colvin, M.M., Deswal, A., Drazner, M.H., Dunlay, S.M., Evers, L.R., Fang, J.C., Fed- son, S.E., Fonarow, G.C., Hayek, S.S., Hernandez, A.F., Khazanie, P., Kittle- son, M.M., Lee, C.S., Link, M.S., Milano, C.A., Nnacheta, L.C., Sandhu, A.T., Stevenson, L.W., Vardeny, O., Vest, A.R., Yanc...
2022
-
[7]
Journal of Basic Engineering82(1), 35–45 (1960)
Kalman, R.E.: A new approach to linear filtering and prediction problems. Journal of Basic Engineering82(1), 35–45 (1960). https://doi.org/10.1115/1.3662552
-
[8]
Journal of the American Society of Echocardio- graphy28(1), 1–39.e14 (2015)
Lang, R.M., Badano, L.P., Mor-Avi, V., Afilalo, J., Armstrong, A., Ernande, L., Flachskampf, F.A., Foster, E., Goldstein, S.A., Kuznetsova, T., et al.: Recom- mendations for cardiac chamber quantification by echocardiography in adults: An update from the American Society of Echocardiography and the European Associ- ation of Cardiovascular Imaging. Journal...
-
[9]
IEEE Transactions on Medical Imaging 38(9), 2198–2210 (2019)
Leclerc, S., Smistad, E., Pedrosa, J., Østvik, A., Cervenansky, F., Espinosa, F., Espeland, T., Berg, E.A.R., Jodoin, P.M., Grenier, T., Lartizien, C., D’hooge, J., Lovstakken, L., Bernard, O.: Deep learning for segmentation using an open large-scale dataset in 2D echocardiography. IEEE Transactions on Medical Imaging 38(9), 2198–2210 (2019). https://doi....
-
[10]
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019), https://openreview.net/forum? id=Bkg6RiCqY7
work page 2019
-
[11]
Nature ComputationalScience1(5), 313–320(2021)
Niederer, S.A., Sacks, M.S., Girolami, M., Willcox, K.: Scaling digital twins from the artisanal tothe industrial. Nature ComputationalScience1(5), 313–320(2021). https://doi.org/10.1038/s43588-021-00072-5 10 D. P. M. Cao and H. Pham
-
[12]
Nature580(7802), 252–256 (2020)
Ouyang, D., He, B., Ghorbani, A., Yuan, N., Ebinger, J., Langlotz, C.P., Heidenre- ich, P.A., Harrington, R.A., Liang, D.H., Ashley, E.A., Zou, J.Y.: Video-based AI for beat-to-beat assessment of cardiac function. Nature580(7802), 252–256 (2020). https://doi.org/10.1038/s41586-020-2145-8
-
[13]
Biomechanics and Modeling in Mechanobiology20(3), 803–831 (2021)
Peirlinck, M., Sahli Costabal, F., Yao, J., Guccione, J.M., Tripathy, S., Wang, Y., Ozturk, D., Segars, P., Morrison, T.M., Levine, S., Kuhl, E.: Precision medicine in human heart modeling: Perspectives, challenges, and opportunities. Biomechanics and Modeling in Mechanobiology20(3), 803–831 (2021). https://doi.org/10.1007/ s10237-021-01421-z
work page 2021
-
[14]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Perez, E., Strub, F., de Vries, H., Dumoulin, V., Courville, A.: FiLM: Visual rea- soning with a general conditioning layer. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 32, pp. 3942–3951 (2018). https://doi.org/10.1609/ aaai.v32i1.11671
work page 2018
-
[15]
Nature Cardiovascular Research4(5), 624–636 (2025)
Qian, S., Ugurlu, D., Fairweather, E., Dal Toso, L., Deng, Y., Strocchi, M., Ci- cci, L., Jones, R.E., Zaidi, H., Prasad, S., Halliday, B.P., Hammersley, D., Liu, X., Plank, G., Vigmond, E., Razavi, R., Young, A., Lamata, P., Bishop, M., Niederer, S.: Developing cardiac digital twin populations powered by machine learning provides electrophysiological ins...
work page 2025
-
[16]
In: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015
Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional networks for biomed- ical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Lecture Notes in Computer Science, vol. 9351, pp. 234–241. Springer (2015). https://doi.org/10.1007/978-3-319-24574-4_28
-
[17]
Circulation Research35(1), 117–126 (1974)
Suga, H., Sagawa, K.: Instantaneous pressure-volume relationships and their ratio in the excised, supported canine left ventricle. Circulation Research35(1), 117–126 (1974). https://doi.org/10.1161/01.RES.35.1.117
-
[18]
In: Advances in Neural Information Processing Systems
Tarvainen, A., Valpola, H.: Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In: Advances in Neural Information Processing Systems. vol. 30, pp. 1195–1204 (2017)
work page 2017
-
[19]
Vovk, V., Gammerman, A., Shafer, G.: Algorithmic Learning in a Random World. Springer (2005). https://doi.org/10.1007/b106715
doi:10.1007/b106715 2005
-
[20]
In: Medical Image Computing and Computer-Assisted Intervention – MICCAI
Wei, H., Cao, H., Cao, Y., Zhou, Y., Xue, W., Ni, D., Li, S.: Temporal-consistent segmentation of echocardiography with co-learning from appearance and shape. In: Medical Image Computing and Computer-Assisted Intervention – MICCAI
-
[21]
Echocardiography31(1), 87–100 (2014)
Wood, P.W., Choy, J.B., Nanda, N.C., Becher, H.: Left ventricular ejection fraction and volumes: It depends on the imaging method. Echocardiography31(1), 87–100 (2014). https://doi.org/10.1111/echo.12331
-
[2020]
Lecture Notes in Computer Science, vol. 12264, pp. 623–632. Springer (2020). https://doi.org/10.1007/978-3-030-59713-9_60
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.