{"id":"bdf48c22-6195-4b5b-9691-5cd616c965f4","arxiv_id":"1908.04875","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A multilevel multifidelity Monte Carlo workflow with 3D, 1D, and 0D cardiovascular models estimates hemodynamic uncertainties with 10 to 100 times lower computational cost than single-fidelity Monte Carlo.","lead":"This paper applies a multilevel multifidelity Monte Carlo method to cardiovascular blood flow simulations, combining cheap 0D and 1D models with expensive 3D models to estimate uncertainties at lower computational cost. It reports 10 to 100 times cost savings compared to standard Monte Carlo for four patient-based anatomies, including healthy and diseased coronary and aortic models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Extrapolated cost savings rely on variances and correlations from 25 pilot samples per level; validation in §5.2 tests only ≤8× variance reductions for one QoI per model, leaving the 10–100× improvement claims sensitive to pilot sampling error.","rationale":"The paper is a solid application paper: it integrates Dakota with SimVascular, provides a public interface, performs mesh convergence, and validates one MLMF extrapolation with actual runs. I see no internal inconsistency or mathematical flaw that would warrant rejection. The load-bearing issue is the epistemic status of the headline numbers: those are point predictions from a single small pilot, and the only direct empirical check is narrow. Since the reader already assigned CONDITIONAL on closely related grounds, my stress-test does not change the verdict; it sharpens the condition. A bootstrap of the pilot data is cheap, uses only data already collected, and would either confirm the claimed ranges or show that the lower bound of the improvement is much smaller than \"one to two orders.\" If the bootstrap CIs are wide (e.g., MLMC/MLMF ratio 5th percentile <10), the paper should weaken the headline or add more pilot samples to tighten estimates. If the CIs are narrow, the conditional concern is resolved. This is the single most load-bearing check because the entire contribution rests on the extrapolation accuracy.","tokens_in":37535,"tokens_out":10887,"duration_ms":112553,"concrete_test":"Bootstrap the pilot data. For each model and each of the four QoI categories, resample with replacement the 25 per-level HF/LF pairs (and the 3D level pairs for MLMC) 1000 times. For each resample, re-estimate V[Y_l], ρ_l, and per-level costs from Table 2, then recompute the extrapolated nCI=0.01 equivalent costs for MC, MLMC, MLMF 3D-1D, and MLMF 3D-1D-0D using Eqs. (16), (21), and (23). Report the 5th–95th percentile of the cost ratios MLMC/MLMF and MC/MLMF. If the lower 5th percentile for any QoI category falls below the claimed 10× (MLMC) or 100× (MC), the headline improvement is not robust to pilot sampling variability. As a complementary empirical check, run the MLMF 3D-1D-0D scheme to an actual nCI of 0.02 for one local TAWSS QoI and compare actual to extrapolated cost; this tests extrapolation in the regime where the advantage is smallest.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Tables 5a/5b claim 10–100× cost reductions for MLMF over MLMC/MC at nCI=0.01. These are extrapolated, not measured, costs: the pilot (Section 5.1.2) supplies 25 paired samples per level, from which Eqs. (16), (21), and (23) are fed estimates of V[Y_l], ρ_l, and cost ratios. With 25 samples, V[Y_l] has a relative standard error of roughly 29%, and ρ_l has wide uncertainty (for nominal ρ=0.8, the 95% interval is about 0.59–0.91; for ρ=0.99, about 0.98–0.995). Because the variance-reduction factor 1−Λ(r*) depends nonlinearly on ρ² and w, these sampling errors propagate into the extrapolated costs and can shift ratios by factors of 2–4. The paper does not report uncertainties on any Table 5 entry. The direct validation in Section 5.2 (Fig. 10) uses a smaller pilot and ε=0.5, 0.25, 0.125—at most an 8× variance reduction—and covers one pressure QoI per healthy model, not the local TAWSS QoIs where the paper still claims 10–40× gains. Thus the order-of-magnitude headline is not yet robust to pilot sampling error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper integrates Sandia Dakota's multilevel multifidelity Monte Carlo (MLMF) estimators with the SimVascular modeling pipeline to perform uncertainty quantification for cardiovascular hemodynamics. Three fidelities are used (3D, 1D, and 0D) with multiple spatial resolution levels, and the framework is demonstrated on healthy and diseased aorto-femoral and coronary anatomies with eight stochastic inputs and 132 or 148 quantities of interest. The central claim is that MLMF estimators achieve one to two orders of magnitude cost reduction over multilevel Monte Carlo and one to three orders over Monte Carlo when extrapolating to a normalized confidence interval of nCI=0.01, with extrapolated costs reported in Tables 5a and 5b. The paper also compares 3D-1D and 3D-1D-0D schemes, healthy versus diseased models, and global versus local QoIs, and it validates the extrapolation procedure for one representative pressure QoI per healthy model in Section 5.2.","tokens_in":37792,"tokens_out":4178,"duration_ms":45130,"significance":"If the reported cost reductions hold, the paper is a valuable demonstration that multilevel multifidelity estimators can make UQ feasible for expensive patient-specific hemodynamic simulations. The work is strengthened by a public workflow repository with a DOI, careful construction and validation of 0D/1D models against 3D results, a wide range of QoIs, and consistent comparisons across methods and anatomies. The main weakness is that the headline cost numbers are extrapolated from small pilot samples, and the validation covers only one pressure QoI per healthy model at variance reductions far smaller than those claimed for the nCI=0.01 target. The manuscript therefore presents a sound methodological application whose quantitative headline needs additional support before the order-of-magnitude claims can be considered robust.","major_comments":[{"comment":"The extrapolated cost comparisons in Tables 5a and 5b are computed from pilot estimates of level variances, correlations, and cost ratios using only 25 paired samples per level. With n=25, a sample variance has a relative standard error of roughly 28%, and a correlation estimate near 0.8 has a 95% confidence interval of approximately 0.59–0.91; since the control-variate variance reduction factor 1−Λ(r*) depends nonlinearly on ρ² and the cost ratio w, these sampling errors can plausibly change the extrapolated cost ratios by factors of 2–4. No uncertainty or sensitivity analysis is reported for any entry in Tables 5a and 5b. The direct validation in Section 5.2 uses a smaller pilot and targets only variance improvements of ε=0.5, 0.25, and 0.125, which correspond to at most an 8× variance reduction, and it covers one pressure QoI per healthy model rather than the local TAWSS QoIs where the paper still claims 10–100× gains. I request either bootstrap confidence intervals on the extrapolated costs or an independent validation at a variance reduction much closer to the nCI=0.01 target, for several representative QoIs including local WSS.","section":"§5.1.2, Tables 5a/5b, Eqs. (16), (21), (23)"},{"comment":"Equation (28) defines the validation target as Vtarget[Q̂_MLMF] = ε Vpilot[Q̂_ML], where Q̂_ML is stated to be the MLMC estimator. If the pilot variance on the right-hand side is not the variance of the MLMF estimator being validated, then the extrapolation exercise does not directly test the MLMF cost curves shown in Figure 10 and Tables 5a/5b. The text should clarify why the MLMC pilot variance is the reference, and if the equation is correct, it should be shown that this target corresponds to the same nCI improvement used in Section 5.1.2. As written, the validation protocol appears to test a different variance-reduction path than the one used for the central cost claims.","section":"§5.2, Eq. (28)"},{"comment":"The manuscript explicitly targets only the variance contribution to the mean squared error and assumes that the 3D model's discretization bias is already satisfactory, with the underlying convergence studies stated to be 'not reported here for brevity'. This is an acceptable choice for a relative cost comparison of estimators, but it means the reported nCI values quantify uncertainty around the 3D model prediction, not around the true clinical quantity. Because the abstract and introduction describe improved accuracy of hemodynamic quantities of interest, the manuscript should state this limitation more prominently and either provide the mesh-convergence data in supplementary material or give a reference where it can be found.","section":"§3.1 and §6 (Discussion)"}],"minor_comments":[{"comment":"In the Aorto-Femoral Healthy Model TAWSS row, the Monte Carlo effective cost is printed as '23 94.0', which appears to be a typo for '23 940.0'; please correct this.","section":"Table 5a"},{"comment":"The caption says the values shown in green pentagons and triangles are the projected and actual extrapolated costs, respectively, but the legend text is not fully explicit about which symbol corresponds to which quantity; please clarify the symbol mapping.","section":"Figure 10 caption"},{"comment":"The validation section uses a pilot with 15, 10, and 5 samples on the three discrepancy levels, whereas the main pilot uses 25 samples per level; the text should state explicitly why this smaller pilot was chosen and how it affects the comparison with the main extrapolation.","section":"§5.2"},{"comment":"Some numbers in Table 2 use inconsistent digit-grouping separators (e.g., '1 219 672', '1 036 483', '1 026,675'); please standardize the formatting.","section":"Table 2"},{"comment":"The phrase 'a simpliﬁed LPN was constructed from only resistor elements to serve the purpose of a coarse model representation' is clear, but the later statement that the zero-dimensional models count as 'resolution levels' could be made more explicit about how a lack of spatial resolution is handled in the MLMF hierarchy.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid application-oriented contribution with a reproducible workflow and a clear demonstration of the MLMF concept for cardiovascular UQ. My main reservation is that the order-of-magnitude cost claims are extrapolated from very small pilot samples and are validated only for one pressure QoI per healthy model at much smaller variance reductions; this is fixable with additional uncertainty quantification on the cost estimates or a more demanding validation experiment. If the revision addresses this, I would view the paper as acceptable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is my read on Fleeter et al. The paper does something useful: it takes the MLMF estimator (from Geraci et al. and others) and actually runs it on four cardiovascular models with three fidelities, comparing healthy/diseased, local/global QoIs, and two MLMF configurations. The workflow coupling Dakota and SimVascular is real, the repository is available, and the low-fidelity models are validated against 3D in a sensible way. That is a solid engineering contribution.\n\nWhere it gets soft: the central claim—10 to 100 times cheaper than MLMC/MC—is not measured. It's extrapolated from a pilot of 25 samples per level. With n=25, the variance estimates have ~30% standard error and correlations are quite uncertain. The formulas (16)-(23) are nonlinear in these quantities, so the extrapolated cost ratio can shift by factors of 2-4. The paper validates the extrapolation in Figure 10, but only for one pressure QoI per healthy model and only for variance reductions up to 8x, not the 10-100x headline. They do not report uncertainty on any Table 5 entry. This is not fatal—the method is sound and the qualitative conclusion (MLMF beats MLMC by a lot) will likely survive—but the exact improvement factors should be treated as provisional.\n\nAlso minor: the high-fidelity bias is assumed acceptable based on mesh convergence, not quantified; the pilot is admittedly larger than needed for some global QoIs (they say so in the discussion); and the 3D-1D vs 3D-1D-0D comparison is based on pilot nCIs, so it inherits the same sampling noise. None of these invalidates the paper.\n\nWho is this for? People doing UQ for patient-specific hemodynamics who want a practical recipe and realistic cost estimates. The method itself is not new, but the systematic evaluation across anatomies and QoI types is a genuine addition. The citation pattern is appropriate—MLMF methodology cites the method papers, and the application results are new.\n\nRecommendation: send it out. A good referee should ask for a sensitivity analysis on the pilot size (e.g., bootstrap the pilot estimates or report confidence intervals on the extrapolated costs) and for validation on a local QoI, but the paper is worth the referee's time. I would not desk reject. It might be more honest to soften the '10 to 100 times' language in the abstract until the numbers carry error bars.","headline":"Solid application of MLMF to cardiovascular UQ with real cost gains, but the headline savings are extrapolated from small pilots and need uncertainty bounds before I'd trust the factor-of-10-100 numbers.","tokens_in":38341,"tokens_out":1929,"would_cite":true,"duration_ms":19979,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C05","92C35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Cheap 0D/1D blood-flow models cut uncertainty-quantification cost by up to 1000x","keywords":["multilevel multifidelity Monte Carlo","cardiovascular hemodynamics","uncertainty quantification","reduced-order models","control variates","wall shear stress","normalized confidence interval","patient-specific modeling"],"falsifier":"Run the 3D-1D-0D workflow to an actual normalized confidence interval of 0.01 for all quantities of interest on a new anatomy, with independent pilot and convergence phases, and compare realized cost with the extrapolated cost; if realized costs systematically exceed the extrapolated costs beyond the spread seen in the paper's single validation case, the central efficiency claim fails in that setting. A cheaper check is to resample the pilot many times with different 25-sample seeds and see whether the extrapolated cost varies by more than an order of magnitude.","tokens_in":37326,"feed_emoji":"🩺","tokens_out":8401,"duration_ms":72628,"temperature":0.7,"pith_summary":"The paper tries to establish that uncertainty quantification for patient-specific cardiovascular hemodynamics can be made practical by pooling three model fidelities—full 3D simulations plus 1D and 0D reduced models—inside a single multilevel multifidelity Monte Carlo estimator. The claim is that for the same target confidence interval, this estimator needs one to two orders of magnitude less compute than multilevel Monte Carlo on 3D meshes alone, and one to three orders less than plain Monte Carlo, across healthy and diseased aortic and coronary anatomies. If true, clinical users with constrained computational budgets could report error bars on simulated pressures, flows, and wall shear stress rather than only point predictions.","feed_headline":"Cheap models cut blood-flow uncertainty cost by 1000x","feed_subtitle":"Multilevel multifidelity Monte Carlo reaches target confidence using 1-3 orders less compute than standard simulation-based UQ.","key_machinery":"The load-bearing object is the multilevel multifidelity (MLMF) estimator, which writes the target expectation as a sum over levels $\\ell$ of discrepancies $Y_\\ell$ between successive 3D mesh resolutions and, at each level, augments the high-fidelity discrepancy estimator with a control variate $\\alpha_\\ell(\\hat{Y}^{\\mathrm{LF}}_\\ell-\\mathbb{E}[Y^{\\mathrm{LF}}_\\ell])$ built from a low-fidelity model. The variance reduction is governed by the per-level Pearson correlation $\\rho_\\ell$ between high- and low-fidelity discrepancies and the cost ratio $w_\\ell$; these determine the optimal sample allocation $N_\\ell^{\\mathrm{HF}}$ and the low-fidelity oversampling factor $r^\\star_\\ell$. The paper documents high correlations (mostly above 0.99 for global quantities, and above 0.6 for wall shear stress), which is what lets the cheap models carry most of the sampling burden.","core_discovery":"The central discovery is that the variance-reduction machinery of multilevel Monte Carlo—a telescoping sum of discrepancies between successive mesh resolutions—can be fused with a control-variate correction from cheap low-fidelity models at every level, and for cardiovascular hemodynamics this fusion yields estimators whose normalized confidence interval ($\\mathrm{nCI}=6\\sigma/\\mu$) reaches 0.01 at a fraction of the cost of single-fidelity estimators. With two low-fidelity families (1D and 0D, the latter in full and resistor-only forms) plus three 3D mesh levels, the paper reports extrapolated costs as low as tens of equivalent 3D-fine runs for global quantities, versus thousands for Monte Carlo and hundreds for multilevel Monte Carlo, with local wall-shear-stress quantities showing smaller but still substantial gains. The low-fidelity models need not be unbiased: because each level is corrected by a control variate, bias in the 0D or 1D model is automatically compensated as long as the correlation with the high-fidelity quantity stays high.","pith_inferences":["If the pilot-based variance and correlation estimates generalize across patient cohorts, the workflow could be wrapped in an adaptive pilot that spends fewer 3D runs per level (e.g., 5-10 instead of 25) and still reaches the target confidence interval, trading expensive high-fidelity samples for many more cheap low-fidelity samples.","The method's reliance on correlation suggests a cheap screening test for a new anatomy: measure the per-level correlation between 3D and 1D/0D discrepancies for the target quantity before launching a full campaign; if correlations fall well below the values reported here, the MLMF advantage over multilevel Monte Carlo will shrink accordingly.","Time-resolved estimators (full pressure and flow waveforms) would likely preserve the same ordering of gains because the per-time-step variance reduction depends on the same correlations, although the pilot cost would increase by the number of time points reported."],"forward_implications":["Global quantities of interest (outlet flow, outlet pressure, model-averaged pressure) can be converged to a normalized confidence interval of 0.01 with tens of equivalent 3D-fine runs in the best cases, making routine uncertainty quantification feasible within typical clinical compute budgets.","Local quantities such as spatially averaged wall shear stress remain the hard case: even the best MLMF costs are hundreds to thousands of equivalent 3D-fine runs, and diseased geometries widen the gap relative to healthy ones.","Adding very cheap 0D models as an extra coarse level improves the 3D-1D-0D scheme over the 3D-1D scheme, with better post-pilot confidence intervals and lower extrapolated costs, most visibly for local quantities.","Under a fixed high-fidelity budget equal to 50 fine 3D runs, the MLMF schemes deliver smaller (better) confidence intervals for healthy than for diseased anatomies, and for global than for local quantities."],"supporting_citations":[{"why":"Establishes the multilevel Monte Carlo telescoping-sum formulation and the optimal per-level sample allocation that the MLMF estimator extends and that serves as the MLMC baseline.","marker":"[43]"},{"why":"Provides optimal model management for multifidelity Monte Carlo estimation, underpinning the control-variate cost analysis.","marker":"[44]"},{"why":"Introduces the multilevel Monte Carlo with control variate formulation that the MLMF estimator builds on.","marker":"[45]"},{"why":"Defines the normalized confidence interval used as the convergence target and supplies the multifidelity control-variate estimator foundations.","marker":"[46]"},{"why":"Supplies the MLMF algorithm variant used here, including the correlation and cost-ratio formulas for optimal sampling and the validation approach.","marker":"[47]"},{"why":"Provides the cardiovascular modeling and simulation pipeline used to build and solve the 3D and 1D models and to extract the quantities of interest.","marker":"[49]"},{"why":"Provides the uncertainty-quantification driver with MLMF estimators, pilot-run management, and automated convergence.","marker":"[50]"},{"why":"Supplies the coupled momentum method used for deformable-wall fluid-structure interaction in the 3D simulations.","marker":"[59]"}],"fun_headline_variants":["Multilevel multifidelity UQ: 1000x cheaper blood-flow models","Biased low-fidelity models still cut cardiovascular UQ cost","Multilevel multifidelity Monte Carlo: 1000x faster UQ for blood flow","Cardiovascular UQ at 1/1000th compute with multilevel multifidelity","Low-fidelity models with bias: key to affordable heart UQ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported cost reductions are extrapolated from variances and correlations estimated in a single pilot run of 25 samples per level, so the savings assume those pilot estimates represent the true sampling behavior of the quantities of interest.","fun_headline_variants_meta":{"raw":{"variants":["Multilevel multifidelity UQ: 1000x cheaper blood-flow models","Biased low-fidelity models still cut cardiovascular UQ cost","Multilevel multifidelity Monte Carlo: 1000x faster UQ for blood flow","Cardiovascular UQ at 1/1000th compute with multilevel multifidelity","Low-fidelity models with bias: key to affordable heart UQ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":4057,"prompt_tokens":1028,"completion_tokens":3029,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":2928}},"tokens_in":644,"tokens_out":3029,"duration_ms":23161,"temperature":1.0,"reasoning_tokens":2928,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:29:29.204338+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the 3D-1D-0D workflow to an actual normalized confidence interval of 0.01 for all quantities of interest on a new anatomy, with independent pilot and convergence phases, and compare realized cost with the extrapolated cost; if realized costs systematically exceed the extrapolated costs beyond the spread seen in the paper's single validation case, the central efficiency claim fails in that setting. A cheaper check is to resample the pilot many times with different 25-sample seeds and see whether the extrapolated cost varies by more than an order of magnitude.","supporting_citations":[{"cited_title":"Giles, Multilevel Monte Carlo methods., Acta Numer 24 (2015) 259–328","cited_arxiv_id":null,"evidence_quote":"Establishes the multilevel Monte Carlo telescoping-sum formulation and the optimal per-level sample allocation that the MLMF estimator extends and that serves as the MLMC baseline."},{"cited_title":"Peherstorfer, K","cited_arxiv_id":null,"evidence_quote":"Provides optimal model management for multifidelity Monte Carlo estimation, underpinning the control-variate cost analysis."},{"cited_title":"Nobile, F","cited_arxiv_id":null,"evidence_quote":"Introduces the multilevel Monte Carlo with control variate formulation that the MLMF estimator builds on."},{"cited_title":"Geraci, M","cited_arxiv_id":null,"evidence_quote":"Defines the normalized confidence interval used as the convergence target and supplies the multifidelity control-variate estimator foundations."},{"cited_title":"Geraci, M","cited_arxiv_id":null,"evidence_quote":"Supplies the MLMF algorithm variant used here, including the correlation and cost-ratio formulas for optimal sampling and the validation approach."},{"cited_title":"Updegrove, N","cited_arxiv_id":null,"evidence_quote":"Provides the cardiovascular modeling and simulation pipeline used to build and solve the 3D and 1D models and to extract the quantities of interest."},{"cited_title":"Adams, M","cited_arxiv_id":null,"evidence_quote":"Provides the uncertainty-quantification driver with MLMF estimators, pilot-run management, and automated convergence."},{"cited_title":"Figueroa, I","cited_arxiv_id":null,"evidence_quote":"Supplies the coupled momentum method used for deformable-wall fluid-structure interaction in the 3D simulations."}],"review_version":1}