{"id":"57bb6827-7ef9-4905-b5f0-7feb7d87426a","arxiv_id":"2505.01666","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Multi-fidelity Gaussian process regression fusing sparse experimental damage indices with simulated ones estimates damage state more accurately than single-fidelity GPs alone.","lead":"This paper combines a small set of experimental guided-wave measurements with cheaper simulated signals to estimate structural damage through multi-fidelity Gaussian process regression. The aim is to reduce the amount of costly real-world data needed for damage state quantification.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test Case 2's 'simulated' DIs are generated by a load-compensation model calibrated on the same experimental load states, so the multi-fidelity advantage there may be leakage rather than genuine fusion of an independent low-fidelity source.","rationale":"The reader's weakest assumption is the constant-rho autoregressive link in Eq. 2. I agree that is a real modeling risk, and the paper's own Fig. 8 showing initial RMSE increases is consistent with misspecification. But a single-rho violation would not necessarily sink the empirical claim: with enough experimental data the discrepancy GP can correct a state-dependent bias, and the paper repeatedly shows final RMSE decreases. The calibration leakage in Test Case 2 is more load-bearing because it attacks the interpretation of the strongest reported improvement. The 'simulated' signals are produced by a model whose constants are estimated from experimental signals at 'various loading conditions' (Sec. 4.2), and the paper's experimental corpus for this test case has only five load states. Using only 0 and 20 kN for GP training does not make the 5, 10, and 15 kN states independent, because those signals appear to have been used to calibrate the generator. Thus the RMSE gains in Figs. 14-15 could be an artifact of leakage rather than evidence that independent low-fidelity data help. This concern does not impugn Test Case 1, where FEM simulation is plausibly independent, and it does not suggest misconduct. The concrete holdout-calibration check would settle the issue: if the advantage persists under calibration on only the 0 and 20 kN training signals, the claim is supported; if it vanishes, the paper's cross-case generalization claim should be restricted to genuinely independent simulators. The reader's CONDITIONAL verdict remains appropriate; I would keep CONDITIONAL and add this explicit condition.","tokens_in":18628,"tokens_out":10037,"duration_ms":106957,"concrete_test":"Re-run Test Case 2 with the compensation model constants (Eqs. 19 and 20) estimated using only the 0 and 20 kN experimental signals, i.e., the same two states available to the multi-fidelity GP, and regenerate the simulated DI curves at intermediate loads. Compare RMSE and R2 against the standard GPRM baselines in Table 3 and the multi-fidelity results in Figs. 14-17, using at least 10 random train/test splits. If the advantage over standard GPRM largely disappears, the reported gain is due to calibration leakage; if it persists, the leakage is not the operative mechanism.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing weakness is in the second test case (Sec. 4). The low-fidelity dataset is not an independent simulation: the physics-based load-compensation model of Sec. 4.2 has constants A, B (Eq. 19) and K_phase (Eq. 20) that are 'calculated from experimental signals collected at various loading conditions.' Since the entire experimental corpus in this test case consists of only five loads (0, 5, 10, 15, 20 kN) and simulated DIs are generated at 0.5 kN increments, those constants are effectively fit to the same intermediate states (5, 10, 15 kN) that Task 1 and Task 2 then treat as target states. The multi-fidelity GPRM is therefore not combining independent experimental and simulated information; it is feeding a physics-based interpolation of the experimental response through the low-fidelity channel. This makes the large RMSE/R2 gains in Sec. 4.3.1 (e.g., 0.0075 to 0.0025, Fig. 14) uninterpretable as evidence that cheap simulations can substitute for experiments. The constant-rho assumption in Eq. 2 is a genuine modeling risk, but this calibration leakage is a more direct threat to the central claim because it concerns whether the comparison is fair at all.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies multi-fidelity Gaussian process regression (MF-GPRM), based on the autoregressive co-kriging model of Kennedy and O'Hagan, to damage state estimation in two aluminum plate test cases. Damage indices (DIs) extracted from experimental guided-wave signals are treated as high-fidelity data, and DIs from FEM simulations (test case 1) or a physics-based load-compensation model (test case 2) as low-fidelity data. The authors compare MF-GPRM with standard GPRM trained on the same experimental data, reporting lower RMSE and higher R^2 for MF-GPRM across several tasks, and additionally investigate active-learning acquisition rules for selecting simulated data points. The central claim is that multi-fidelity fusion of cheap simulated data with a small experimental set improves probabilistic damage-state estimation.","tokens_in":18848,"tokens_out":8014,"duration_ms":76778,"significance":"If the claims hold, the paper offers a practical route to reducing experimental data collection costs in guided-wave SHM by incorporating lower-fidelity simulated data into a probabilistic regression framework. The main strengths are the use of two different data-source types (FEM and physics-based reconstruction), the integration of active learning with multiple acquisition criteria, and the availability of a first test case that relies on an independent FEM simulation. However, the second test case's low-fidelity data are not independent of the experimental data, which substantially weakens the generalizability claim. Overall, the paper is an incremental but potentially useful application of an established method rather than a methodological breakthrough.","major_comments":[{"comment":"The low-fidelity dataset in the second test case is not an independent simulation source. The constants A, B in Eq. (19) and K_phase in Eq. (20) are computed from experimental signals collected at the five loads 0, 5, 10, 15, 20 kN, and the simulated signals are then generated at 0.5 kN increments. The simulated DIs are therefore effectively a dense interpolation of the same experimental response, rather than predictions from an independently calibrated physics model. As a result, the large RMSE/R^2 gains in Figs. 14 and 15 and the numbers in Table 4 cannot be interpreted as evidence that inexpensive independent simulations can substitute for experiments; at most they demonstrate that a physics-based interpolation of the experimental response can be exploited. This compromises the universal-superiority conclusion in Sec. 5. The authors should either redo the second test case with a genuinely independent simulation (e.g., an FEM model of the loaded plate) or explicitly reframe the test case as physics-based data augmentation and correspondingly limit the claims in the abstract and conclusion.","section":"Sec. 4.2, Sec. 4.3"},{"comment":"The variance lower-bound constraint is a free parameter whose value appears to be selected based on the results. Section 3.3.1 states that \"three times the largest variance of the experimental data was chosen as the lower constraint\" for sigma_1^2 and sigma_2^2, while Appendix A.2 shows outputs for multipliers 1, 5, 10, and 15 but gives no criterion for choosing 3. The sensitivity of the main RMSE/R^2 comparisons to this multiplier is not reported. Because the uncertainty bounds are a central claimed advantage of the method, the choice should be justified by a principled rule such as cross-validation, or a sensitivity analysis over the multiplier range should be reported for at least the main comparisons.","section":"Sec. 3.3.1, Appendix A.2"},{"comment":"The experimental training sets are formed by randomly drawing 15 of 20 realizations, but no repeated-split statistics are reported. RMSE and R^2 values are presented for a single split per configuration, so the observed differences between standard GPRM and multi-fidelity GPRM may be within the sampling variability of the experimental realizations. The authors should repeat the random splitting (e.g., 10 or more seeds) and report mean and standard error of the performance metrics to support the claim that the improvements are statistically meaningful.","section":"Sec. 3.3, Figs. 8 and 11"},{"comment":"The model assumes a single state-independent correlation coefficient rho between the low- and high-fidelity maps, and the paper acknowledges that \"the actual value can change at different locations\" without providing any diagnostic for this assumption. The empirical observation in Sec. 3.3.1 and Fig. 8 that RMSE initially increases when simulated points are added could be a symptom of rho being misestimated in some regions of the state domain. Please provide evidence for the adequacy of the constant-rho assumption, for example by estimating rho over a moving window of the damage-state domain or by comparing against a model with a spatially varying rho; at minimum, report the fitted values of rho for the studied paths and discuss whether they are consistent across the domain.","section":"Sec. 2.1.1, Eq. (2)"}],"minor_comments":[{"comment":"The text states that RMSE decreased from 0.032593 to 0.030838, \"approximately halving the original value from GPRM,\" but this is a reduction of about 5%, not a halving. Please correct the description to reflect the actual magnitude of the improvement.","section":"Sec. 3.3.1, Fig. 7(d)"},{"comment":"The captions for these figures list panel labels incorrectly, repeating \"(a)\" and \"(b)\" instead of \"(a)\" through \"(d)\". Please correct the panel designations so that the text and figures are unambiguous.","section":"Appendix A.2, Figs. A.4 and A.5"},{"comment":"The text says \"an increment of 0.5N, which is one-tenth of the experimental state increment\"; since the experimental increments are in kN, this should be \"0.5 kN.\"","section":"Sec. 4.2"},{"comment":"There are several typos: \"the experimental dat\" should be \"the experimental data\" in Sec. 3.3.1, and \"demonstrat\" should be \"demonstrate\" in the introductory paragraph of Sec. 4. Please proofread the final text.","section":"Sec. 3.3.1, Sec. 4"},{"comment":"The active-learning RMSE curves appear to be single-run results, whereas the random-selection comparison uses 10 seeds. If the active-learning criteria were not repeated across multiple trials, please state this explicitly and discuss the potential variability of the active-learning results.","section":"Sec. 4.3.3, Fig. 17(b)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable application of a standard multi-fidelity GP method to SHM, but the second test case's low-fidelity data are derived from the same experimental measurements, which is a substantial validation flaw. The first test case provides some independent grounding, but the lack of repeated-split statistics and the post-hoc choice of the variance multiplier are additional concerns. I would encourage the editors to request a revision that addresses the independence issue directly, either by replacing the second test case with an independent simulation or by transparently reframing it as physics-based data augmentation, and to require stronger statistical evidence for the reported gains. The paper may be of interest to the SHM community, but in its current form the claims exceed what the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: the paper does something genuinely useful—it applies an established multi-fidelity GP framework to guided-wave damage indices for state estimation, shows that simulated points can help when experiments are scarce, and couples it with active learning. The first test case, where FEM simulations are independent of the experiments, gives real support to the central claim. The method itself is standard auto-regressive co-kriging, so the novelty is the application and the DI-based feature integration, not the machinery.\n\nWhere the paper earns credit: the experimental work is careful (multiple realizations, multiple paths, two damage scenarios), the figures are informative, and the authors are transparent about the uncertainty floor constraint. The citation pattern is fine—Kennedy and O'Hagan, Perdikaris, and the other standard references are there, and the self-citation to their EWSHM 2023 paper is legitimate as the direct predecessor.\n\nThe main soft spot is test case 2. The low-fidelity data are generated by a physics-based load-compensation model whose constants are fit to the same experimental signals at the five measured loads. Generating simulated points at 0.5 kN increments then amounts to interpolating the experimental trend. The large RMSE drops in Figures 14–15 (0.0075 to 0.0025) are not clean evidence that an independent simulation can substitute for experiments; they may reflect calibration leakage. That's a fairness problem, not misconduct, but it needs to be fixed—either by reframing test case 2 as a calibration study or by using a truly independent low-fidelity source.\n\nMinor issues: the variance lower-bound multiplier (3x) is chosen after trying 1x, 5x, 10x, and 15x in the appendix—treat it as tuned and say so. The 'approximately halving' in Sec. 3.3.1 is inconsistent with the reported numbers (0.0326 to 0.0308 is a 5% drop). And there are no repeated-split statistics, so we don't know how stable the improvements are across random train/test splits.\n\nThe central qualitative claim survives because test case 1 is independent. But the paper needs major revision before it's publishable: fix the test case 2 leakage, correct the numerical inconsistency, and report variance across splits. I'd send it to peer review—the core is sound, the application is practical, and the flaws are addressable. Would I cite it? Not in my own work, but I'd want to read the revised version.","headline":"A solid applied demonstration that multi-fidelity GPRM helps when experimental data are scarce, but the second test case's 'simulated' data are calibrated on the same experiments, so that part of the evidence is weaker than it looks.","tokens_in":19399,"tokens_out":2704,"would_cite":false,"duration_ms":28006,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulated guided-wave data can replace costly experiments in structural damage estimation when fused through a two-level Gaussian process.","keywords":["multi-fidelity Gaussian process regression","structural health monitoring","damage index","guided waves","active learning","co-kriging","damage state estimation","uncertainty quantification"],"falsifier":"Take a dataset where the experimental DI trend is known at many states and the simulated DI deviates from it by a bias that changes sign along the state axis. Train the multi-fidelity model on two experimental states plus simulated points located only inside the sign-change region, and check whether test RMSE on the held-out experimental states stays below the standard-Gaussian-process baseline; if it does not, the constant-scaling premise is the cause.","tokens_in":18360,"feed_emoji":"📡","tokens_out":6576,"duration_ms":66801,"temperature":0.7,"pith_summary":"This paper aims to show that structural damage states can be estimated accurately even when real experimental data are scarce, by training a two-level Gaussian process on both experimental and simulated guided-wave features. Across two test cases—crack size on an aluminum coupon and applied load on another coupon—the multi-fidelity model repeatedly achieves lower RMSE and higher $R^2$ than a standard Gaussian process regression model trained on the same experimental data alone. The practical promise is that cheap simulation can substitute for costly experiments without sacrificing accuracy, and that active learning makes the data-efficient procedure stable.","feed_headline":"Cheap simulated data sharpens structural damage estimates","feed_subtitle":"A multi-fidelity Gaussian process trained on a few experiments plus simulated waves beats experimental-only models.","key_machinery":"The load-bearing object is the two-level auto-regressive co-kriging identity $f_2(x)=\\rho f_1(x)+\\delta(x)$, in which the experimental mapping $f_2$ is written as a scaled version of the simulated mapping $f_1$ plus a Gaussian-process discrepancy term $\\delta$. This turns the problem of fusing experiment and simulation into one joint Gaussian process with a squared-exponential kernel, a cross-correlation parameter $\\rho$, and a variance lower-bound constraint that keeps predicted uncertainty no smaller than the scatter in the experimental data. The same Gaussian process posterior then supplies both the mean damage-state estimate and the acquisition function values (L2 loss, max variance, upper confidence bound, expected improvement) used for active learning.","core_discovery":"The central claim is that auto-regressive co-kriging, applied to damage indices extracted from guided waves, lets a small set of experimental DI values (as few as two or three states) be augmented with simulated DI values to produce mean predictions and confidence bounds that follow the experimental trend better than a standard Gaussian process using only the experimental points. In the fixed-high-fidelity task, adding simulated points to a fixed experimental set lowers RMSE and raises $R^2$ once enough points are included, although the first added points can initially increase error. In the constant-total-states task, replacing experimental states with simulated states degrades accuracy more slowly than removing the experimental data outright. In the load case, combining the multi-fidelity model with active learning gives faster and more stable RMSE convergence than randomly selected simulated points, and can outperform a standard Gaussian process trained on more experimental states.","pith_inferences":["The same co-kriging structure should transfer to other damage-sensitive features, such as spectral or time-frequency features, as long as the simulation-to-experiment relationship is approximately affine; a head-to-head comparison of DI types under identical experimental budgets would test this.","The reported initial RMSE rise suggests a practical diagnostic: monitoring the derivative of RMSE with respect to added simulated points could let an operator stop augmentation before simulation bias dominates the estimate.","Because the paper applies a hard variance lower bound, a fully Bayesian treatment that treats the bound as a prior on noise levels would reveal how much of the gain comes from the constraint versus the multi-fidelity fusion itself.","For deployment on real structures, the main unknown is whether the cross-correlation parameter learned on one sensor path or coupon generalizes to another; a cross-validation across sensor paths would settle that."],"forward_implications":["Simulated guided-wave data can substitute for a large fraction of experimental states: in the load test, models with two experimental sets plus simulated points beat a standard Gaussian process with three or four experimental sets.","The multi-fidelity model can be used to fill data-sparse regions of the damage-state axis, where experiments are impractical or prohibitively costly.","An optimal amount of low-fidelity data exists: RMSE first rises and then falls as simulated points are added, so a sweep of added points can identify the best low-fidelity budget.","Active-learning acquisition functions, especially upper confidence bound and expected improvement, make convergence faster and less dependent on the random order of data addition."],"supporting_citations":[{"why":"Supplies one of the two damage-index formulas used to convert guided-wave signals into regression targets.","marker":"[8]"},{"why":"Introduces the auto-regressive co-kriging formulation that the paper adapts to two fidelity levels.","marker":"[34]"},{"why":"Decouples the co-kriging problem into independent kriging steps, the computational basis for training.","marker":"[35]"},{"why":"Provides the multi-fidelity Gaussian process framework the paper implements and extends with experimental data.","marker":"[36]"},{"why":"Combines experimental data and simulations in a Gaussian process calibration setting, the precedent for fusing the two data sources.","marker":"[37]"},{"why":"Reduces the dimensionality of the calibration framework, cited for handling highly multivariate simulator outputs.","marker":"[38]"},{"why":"Cited as the comprehensive treatment of the two-level fidelity formulation used in the paper.","marker":"[40]"},{"why":"Supplies the physics-based load compensation model that generates the simulated guided-wave signals in the second test case.","marker":"[45]"},{"why":"Demonstrates Gaussian process regression for active-sensing structural health monitoring and is used to generate baseline signals at different loads.","marker":"[47]"}],"fun_headline_variants":["Multi-fidelity GP sharpens damage estimates from sparse data","Simulated waves + few experiments improve damage state prediction","Data fusion in GP yields better structural damage estimates","Less experimental data, same damage accuracy with GP fusion","Multi-fidelity Gaussian process beats experimental-only models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the difference between simulated and experimental damage indices is a fixed scaling factor plus a stable discrepancy; if the simulation's bias changes across damage states, the extra simulated points can pull the estimate away from the real trend—the paper itself reports RMSE rising before it falls.","fun_headline_variants_meta":{"raw":{"variants":["Multi-fidelity GP sharpens damage estimates from sparse data","Simulated waves + few experiments improve damage state prediction","Data fusion in GP yields better structural damage estimates","Less experimental data, same damage accuracy with GP fusion","Multi-fidelity Gaussian process beats experimental-only models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000733,"raw_usage":{"total_tokens":3296,"prompt_tokens":980,"completion_tokens":2316,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":2241}},"tokens_in":596,"tokens_out":2316,"duration_ms":17805,"temperature":1.0,"reasoning_tokens":2241,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:13:21.165166+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a dataset where the experimental DI trend is known at many states and the simulated DI deviates from it by a bias that changes sign along the state axis. Train the multi-fidelity model on two experimental states plus simulated points located only inside the sign-change region, and check whether test RMSE on the held-out experimental states stays below the standard-Gaussian-process baseline; if it does not, the constant-scaling premise is the cause.","supporting_citations":[{"cited_title":"Inferring solutions of differential equations using noisy multi-fidelity data,","cited_arxiv_id":null,"evidence_quote":"Cited as the comprehensive treatment of the two-level fidelity formulation used in the paper."},{"cited_title":"Damage detection sen- sitivity characterization of acousto-ultrasound-based structural health monitoring techniques,","cited_arxiv_id":null,"evidence_quote":"Supplies one of the two damage-index formulas used to convert guided-wave signals into regression targets."},{"cited_title":"Multi-fidelity optimization via surrogate modelling,","cited_arxiv_id":null,"evidence_quote":"Introduces the auto-regressive co-kriging formulation that the paper adapts to two fidelity levels."},{"cited_title":"Recursive co-kriging model for design of computer experiments with multiple levels of fidelity,","cited_arxiv_id":null,"evidence_quote":"Decouples the co-kriging problem into independent kriging steps, the computational basis for training."},{"cited_title":"Multi-fidelity modelling via recursive co-kriging and Gaussian–Markov random fields,","cited_arxiv_id":null,"evidence_quote":"Provides the multi-fidelity Gaussian process framework the paper implements and extends with experimental data."},{"cited_title":"Com- bining experimental data and computer simulations, with an application to flyer plate experi- ments,","cited_arxiv_id":null,"evidence_quote":"Combines experimental data and simulations in a Gaussian process calibration setting, the precedent for fusing the two data sources."},{"cited_title":"Computer model calibration using high-dimensional output,","cited_arxiv_id":null,"evidence_quote":"Reduces the dimensionality of the calibration framework, cited for handling highly multivariate simulator outputs."},{"cited_title":"Loadmonitoringandcompensationstrategiesforguided- waves based structural health monitoring using piezoelectric transducers,","cited_arxiv_id":null,"evidence_quote":"Supplies the physics-based load compensation model that generates the simulated guided-wave signals in the second test case."},{"cited_title":"Gaussian Process Regression for Active Sensing Probabilistic Structural Health Monitoring: Experimental Assessment Across Multiple Damage and Loading Scenarios","cited_arxiv_id":"2106.14841","evidence_quote":"Demonstrates Gaussian process regression for active-sensing structural health monitoring and is used to generate baseline signals at different loads."}],"review_version":1}