{"id":"298f0041-6426-4202-8c38-785159a1c9d0","arxiv_id":"2608.02135","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of cardiovascular digital twin modeling paradigms that concludes hybrid physics-informed and graph-based methods are the most promising direction for clinical deployment.","lead":"This paper reviews how computer models of the human heart and blood vessels, called cardiovascular digital twins, are built. It compares physics-based and data-driven approaches and argues that combining them is the most promising route to patient-specific medicine.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's central claim that hybrid architectures are a unifying framework rests on a self-declared non-systematic literature selection; if the selection is unrepresentative, the conclusion is unsupported.","rationale":"The paper is a structured narrative review, so its central claim is a synthesis rather than a novel empirical result. The only way to support the claim that hybrid/multi-paradigm architectures are 'unifying' is to base it on a representative and unbiased sample of the literature. The authors' own methodology statement in §1.3 undercuts this: no PRISMA protocol, subjective selection for 'balanced coverage.' This is the same weakness the reader identified, and I agree it is the weakest link. The concern is not that the authors are dishonest; it is that a subjective selection cannot by itself establish a field-level generalization, especially when the conclusion is a research-priority recommendation that could steer future work. The proposed test—a systematic replication of the search and comparison of the included study set to the full eligible set—would directly settle representativeness. If the included references are a biased sample, the central claim is not verified; if they are representative, the conclusion stands. The secondary Eq. 1 error (missing ∂u/∂t in the Navier-Stokes residual) reinforces the need for careful verification of the paper's technical description of PINNs, but it is not the primary load-bearing issue. I therefore keep the reader's UNVERDICTED verdict: the claim is plausible but not established by the evidence presented.","tokens_in":19419,"tokens_out":5852,"duration_ms":48883,"concrete_test":"Systematically replicate the search using the exact terms in §1.3 across PubMed, Scopus, Web of Science, and Google Scholar; apply a PRISMA-style screening protocol; and code each eligible study by paradigm (mechanistic, data-driven, PINN, GNN, hybrid) and by whether it reports positive, negative, or neutral evidence for hybrid approaches. Then compare the distribution and the proportion of hybrid-favoring studies against the 80 references included in this review. If the included set is significantly enriched for hybrid/PINN/GNN success papers relative to the full eligible corpus, the central claim should be downgraded. As a secondary check, correct Eq. 1 by adding the ρ∂u/∂t term and re-examine whether §5.2's claims about recovering time-dependent hemodynamic fields still follow from the cited applications.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that hybrid and multi-paradigm architectures 'constitute a unifying framework'—is a generalization about the state of the field. The only evidence offered is a narrative synthesis. In Section 1.3 the authors explicitly state that 'No formal PRISMA screening protocol was applied' and that studies were 'selected to achieve balanced coverage across mechanistic, data-driven, and hybrid paradigms.' Because the inclusion criteria are not reproducible, the synthesis is vulnerable to confirmation bias: papers reporting successful PINN/GNN/hybrid demonstrations may be over-represented, while negative results or head-to-head comparisons favoring pure mechanistic or pure data-driven approaches may be under-represented. Table 1's qualitative ratings are said to be 'based on representative published studies,' but no rule for choosing those representatives is given. If the selected set is not a faithful sample of the published evidence, the conclusion that hybrids unify physical consistency, efficiency, and adaptivity is not established. A secondary technical signal of the same fragility: Eq. 1 defines the PINN Navier-Stokes residual without the time-derivative term, so the paper's formal description of its central hybrid paradigm is inaccurate for the unsteady cardiovascular flows it claims to support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative review of computational modelling paradigms for cardiovascular digital twins, spanning mechanistic biophysical models (0D/1D/3D and electromechanical), data-driven machine learning, physics-informed neural networks (PINNs), graph neural networks (GNNs), and hybrid/multi-paradigm frameworks. It discusses definitions and architecture of digital twins, clinical data sources, personalisation and uncertainty quantification, verification and validation, clinical applications, and ethical/regulatory considerations. The central claim, stated in the Conclusions, is that hybrid and multi-paradigm architectures constitute a unifying framework combining physical consistency, computational efficiency, and adaptive learning. The review is structured as a targeted but explicitly non-systematic literature synthesis.","tokens_in":19740,"tokens_out":3205,"duration_ms":29922,"significance":"If its central claim is accepted, the paper provides a useful methodological map for an interdisciplinary audience. Its strengths include a clear taxonomy of paradigms, a comparative summary in Table 1, an honest discussion of PINN training pathologies and GNN conservation-law limitations, and an explicit treatment of validation standards (including ASME V&V 40) and the distinction between retrospective personalisation and prospective validation. The paper is also commendable for disclosing its methodological limitations rather than presenting the synthesis as a systematic review. However, because the evidence base is a self-selected narrative sample and the quantitative formalism contains a technical error, the field-level conclusion about hybrid architectures should be treated as a plausible research direction rather than an established finding.","major_comments":[{"comment":"The Navier-Stokes residual is stated as L_F = (1/N_c) Σ [ || ρ(u·∇)u + ∇p − μ∇²u ||² + ||∇·u||² ], but this omits the temporal term ρ ∂u/∂t. For the unsteady, pulsatile flows that characterise cardiovascular haemodynamics, the residual as written is not the incompressible Navier-Stokes residual and would incorrectly admit steady-flow solutions. This is load-bearing because the section uses Eq. (1) to define the 'physics' that PINNs enforce. The equation should include ρ ∂u/∂t inside the momentum residual. The surrounding prose also refers to 'governing physical laws' and 'conservation of mass and momentum', so the correction is necessary for formal accuracy.","section":"§5.1, Eq. (1)"},{"comment":"The central claim that hybrid and multi-paradigm architectures 'constitute a unifying framework' is a generalization about the state of the field, but it rests on a non-systematic literature selection. The authors state that no PRISMA protocol was applied and that studies were 'selected to achieve balanced coverage'; no inclusion/exclusion criteria or list of screened studies is provided. A narrative review can be valuable without PRISMA, but the strength of the conclusion should match the evidence. As written, the synthesis is vulnerable to selection bias, particularly if positive hybrid/PINN/GNN demonstrations are over-represented relative to negative results or head-to-head comparisons. I recommend either softening the conclusion to a research priority or hypothesis, or adding a transparent evidence table that documents how representative studies were chosen.","section":"§1.3, Conclusions"},{"comment":"The comparative ratings for 'UQ Maturity' and 'Clinical Readiness' are presented as summary findings, but no rubric, scoring rule, or per-cell citation is given. The caption says ratings are 'qualitative and based on representative published studies', yet the reader cannot verify which studies support each rating or how categories such as 'Moderate' vs 'Moderate–High' were distinguished. Because Table 1 is one of the main comparative outputs of the review and is used to orient the paradigm-by-paradigm discussion, the ratings should either be accompanied by a documented rubric or explicitly labelled as author judgment rather than literature-derived evidence.","section":"Table 1"}],"minor_comments":[{"comment":"The typesetting of the residual norm is broken: the norm bars appear as 'h ... i' rather than as \\(\\|\\cdot\\|\\). This should be fixed for readability.","section":"§5.1, Eq. (1)"},{"comment":"Given the structured-narrative design, a short appendix listing the included studies and search dates would improve reproducibility, even without a full PRISMA flow diagram.","section":"§1.3"},{"comment":"The aleatory/epistemic uncertainty distinction and the discussion of identifiability appear twice, in nearly identical wording. Consolidating these passages would reduce redundancy.","section":"§3.3 and §8"},{"comment":"The caption credits 'WHO and partners' but no specific report or data source is cited in the reference list. A formal citation would help readers verify the mortality figure.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent narrative review with honest limitations, but its main field-level claim is stronger than the non-systematic evidence base supports. The Eq. (1) error is a concrete technical blemish in the central PINN formalism. I believe the paper can be made publishable by correcting the equation, adding transparency about study selection, and tempering the conclusions accordingly. The current version does not meet the bar for acceptance because at least one load-bearing point—the evidence for the 'unifying framework' claim—needs substantial revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a review, not a research contribution, so don't look for a new result. What it does well: it gives a genuinely useful map of the field—mechanistic, data-driven, PINN, GNN, and hybrid—with careful attention to validation and uncertainty quantification, and it is honest about limitations. The section on retrospective vs. prospective validation is particularly sound, and the discussion of practical vs. structural identifiability is better than most reviews at this level.\n\nTwo things bother me. First, Eq. (1) states the PINN residual as ρ(u·∇)u + ∇p − μ∇²u, omitting the temporal term ∂u/∂t. That is formally wrong for unsteady flows, which the paper explicitly claims to target. It is the kind of error that looks like a typo, but it sits in the central equation describing the paper's key hybrid paradigm, so it needs fixing before this can be trusted as a reference. Second, the stress-test concern is on target: Section 1.3 discloses that no PRISMA protocol was applied and that studies were selected to achieve 'balanced coverage.' That makes the synthesis non-reproducible. It does not invalidate the conclusions—the hybrid recommendation is broadly consistent with the surrounding literature—but it does mean the review is an informed opinion with a good bibliography, not an evidence map. The reader's UNVERDICTED verdict seems right.\n\nNo circularity, no invented entities, no hidden parameter fitting. The citation pattern is normal for a review; the authors are not leaning on their own work. The writing is clear and the thinking is serious, even with the equation lapse.\n\nWho is this for? A graduate student or clinician-engineer wanting a single readable overview of the landscape, or a researcher looking for a quick orientation before diving into specific methods. It deserves a serious referee because a good review is useful and this one is mostly good. I would send it out, but flag the equation and ask for a more transparent literature selection—either a PRISMA-style flow diagram or an explicit inclusion rubric. With those changes it would be a solid contribution. I would probably cite it in a background section, though I would not hinge a methods choice on it.","headline":"A competent, clearly structured review of cardiovascular digital twin modelling paradigms, with a real formal error in the PINN Navier-Stokes residual and a disclosed but non-reproducible literature selection; worth refereeing with revision.","tokens_in":20102,"tokens_out":1408,"would_cite":true,"duration_ms":12947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-04T13:48:39.356132+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}