{"id":"3246481b-efe5-4fda-bb15-5116c6188588","arxiv_id":"2608.12603","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A scalable hierarchical Bayesian calibration framework using the Bayesian Committee Machine achieves large runtime speedups and is demonstrated on benchmark and accelerator simulation data.","lead":"This paper combines two existing statistical tools, hierarchical Bayesian calibration and the Bayesian Committee Machine, to speed up computer model calibration on large simulation datasets. It reports large runtime savings while recovering target parameters on benchmark functions and a particle accelerator case study.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 2.7 factorizes the experimental likelihood into a product of univariate GP densities for each zexp,i, while Eq.","rationale":"The reader and I identify the same load-bearing assumption: the factorized experimental likelihood in Eq. 2.7 is contradicted by the joint GP density in Eq. 2.11. This is not a minor notational slip; it changes the target posterior whenever experimental outputs are correlated under the GP, which is generically the case in Kennedy-O'Hagan calibration. The paper never flags Eq. 2.7 as an approximation, and the BCM discussion in Section 2.4 does not repair the missing joint experimental density. The Table 6 discrepancies independently confirm that the reported uncertainties are unreliable in the accelerator example. A computational speedup may remain, but the paper's central claim is conjunctive: substantial computational gains while preserving inferential accuracy. The accuracy component is not supported as stated, so rejection is appropriate unless the likelihood is corrected and all benchmark and accelerator results are re-derived.","tokens_in":21395,"tokens_out":8800,"duration_ms":95712,"concrete_test":"Re-run the Park and cantilever-beam calibrations replacing the Eq. 2.7 product with the exact joint log-likelihood from Eq. 2.11, including Kδ, σ²I, and all cross-covariances among the Nexp experimental points, while holding fixed the random seeds, priors, chain lengths, and BCM splits. Compare posterior means and 68% intervals for θ_i, µpost, and σpost, and the coverage of the true Dexp values. If the differences exceed the Monte Carlo standard error, the factorized likelihood is demonstrated to be misspecified and the reported UQ is invalid; if they do not, Eq. 2.7 should be restated as an approximation and its scope justified.","verdict_should_be":"REJECT","load_bearing_attack":"Section 2.2 writes the posterior as a product over experiments of univariate GP densities pGPz(zexp,i | xexp,i, θ_i, Dsim, ϕδ, ϕsim), together with a separate simulation-data term. Equation 2.11, presented as 'so that the probability density is', gives a joint multivariate normal for (ysim, zexp) with covariance Kh(θ, ϕ). These expressions are not equivalent. Under the KOH model, experimental outputs are correlated through the shared GP emulator: after integrating out η and δ, Cov(zexp,i, zexp,j) = kη((x_i, θ_i), (x_j, θ_j)) + kδ(x_i, x_j), with independent noise only on the diagonal. The product in Eq. 2.7 ignores all off-diagonal experimental covariances and is never identified as an approximation. This is an internal inconsistency, not a disagreement with an external consensus. If the implementation uses Eq. 2.7, the reported posterior intervals and σpost are computed from a misspecified likelihood; if it uses Eq. 2.11, then Eq. 2.7 misstates the model. Either way, the central claim that the framework 'quantifies uncertainty reliably' is unsupported. Table 6 provides independent corroboration: several Dexp values are many posterior standard deviations from E[θ_i] (e.g., row 5: 57.58 vs 53.83 with Σ = 0.028), contradicting the text's claim that true values lie within the chain standard deviation. The computational-scaling results may survive, but the accuracy half of the conjunctive central claim does not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a hierarchical extension of Kennedy-O'Hagan Bayesian calibration in which each experiment has its own calibration parameter theta_i with a hierarchical prior, and uses the Bayesian Committee Machine to split the simulation dataset for scalable Gaussian-process likelihood evaluation inside NUTS-based MCMC. The method is tested on the Park and cantilever beam benchmarks and on a synthetic Argonne Wakefield Accelerator example. The central claim is that the BCM integration yields substantial runtime savings while preserving inferential accuracy.","tokens_in":21719,"tokens_out":6630,"duration_ms":59863,"significance":"The combination of hierarchical calibration with BCM is a natural and potentially useful contribution for simulation-heavy calibration problems, and the paper provides a concrete Julia implementation, runtime scaling experiments, and an honest discussion of BCM limitations. However, the manuscript's accuracy claims are undermined by an internal inconsistency in the likelihood definition and by a direct contradiction in the reported application results, so the significance as presented is not established.","major_comments":[{"comment":"Equation (2.7) writes the posterior as a product over experiments of univariate GP densities p_GPz(z_exp,i | ...), while Equation (2.11) gives the joint multivariate normal density with covariance K_h(theta, phi). These are not equivalent: after integrating out the shared emulator eta and discrepancy delta, the experimental outputs are correlated through the off-diagonal terms of K_h, and the product in Eq. (2.7) ignores those correlations. The manuscript never identifies Eq. (2.7) as an approximation. If the implementation follows Eq. (2.7), the sampled posterior is based on a misspecified likelihood; if it follows Eq. (2.11), the paper misstates the model. Either way, the claim in §5 that the framework 'quantify[ies] uncertainty reliably' is not supported.","section":"§2.2, Eq. (2.7) and Eq. (2.11)"},{"comment":"The text states that 'the true value lies within the standard deviation of the corresponding NUTS-MC chain.' This is contradicted by the table in many rows; for example, row 5 has D_exp = 57.5829, E[theta] = 53.8301, and Sigma = 0.0284, a discrepancy of roughly 132 posterior standard deviations, and rows 2, 6, 9, 11, 12, 13, and 14 show similar departures. These results indicate that the individual calibration parameters are not recovered and that the reported posterior uncertainties are far too small, consistent with the misspecification in Eq. (2.7). The accuracy half of the central claim is therefore empirically contradicted by the paper's own application.","section":"§3.6, Table 6"},{"comment":"Section 2.4.1 concedes that there are no convergence guarantees for BCM when the number of query points (here, the number of experiments) is below the effective degrees of freedom, and §4 acknowledges that the approximation error has no known bounds. Nevertheless, §4 concludes that 'the BCM approximation is accurate and samples the same target distribution.' That conclusion rests on two benchmarks with particular data ratios and does not cover the AWA regime with N_exp = 15; together with the likelihood inconsistency above, the general claim of 'preserving inferential accuracy' is not established.","section":"§2.4.1 and §4"}],"minor_comments":[{"comment":"Equation (2.2) contains rendering artifacts ('radicaltp/radicalvertex') and should be typeset properly.","section":"Eq. (2.2)"},{"comment":"The discussion after Eq. (2.14) is confusing: it says D_exp is 'not conditionally independent from itself,' but the equation concerns splitting D_sim; please clarify the conditional-independence assumptions in the BCM combination.","section":"§2.4, after Eq. (2.14)"},{"comment":"Section 2.5 states that emulator hyperparameters Phi_eta are optimised beforehand, whereas the posterior in Eq. (2.5) includes all hyperparameters; the change in inferential target should be stated explicitly when the model is introduced.","section":"§2.5"},{"comment":"Table 6 would be clearer if the columns used sigma instead of Sigma for the chain standard deviation, and if the number of committee members used in the AWA run were reported.","section":"Table 6"},{"comment":"The statement that chain convergence was 'not explicitly evaluated' should be addressed with standard diagnostics such as R-hat, especially since the AWA chains are short.","section":"§3.5"}],"recommendation":"reject","confidential_remarks":"The manuscript is not ready for publication. The internal inconsistency between Eq. (2.7) and Eq. (2.11) is a core methodological problem, and Table 6 directly contradicts the text's accuracy claims. Even if the likelihood issue were fixed, the reported individual-parameter recovery in the accelerator example would need substantial re-analysis. I would not consider acceptance without a complete reworking of the model specification and the experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one. It's the first combination of hierarchical Kennedy-O'Hagan calibration with the Bayesian Committee Machine, applied to likelihood evaluation rather than point prediction. That is genuinely new, and the scaling results are worth taking seriously: at |Dsim|=150, eight committee members give about 16x speedup, and the error degrades only once each committee gets fewer than ~20 points. The hierarchical recovery of a ground-truth Gaussian on the Park and cantilever benchmarks looks solid, and the Julia/Turing implementation is a sensible way to keep the MCMC tractable. Credit where due: the paper is clearly written and the BCM background is handled honestly, including its lack of convergence guarantees for few query points.\n\nThe soft spot is load-bearing. Equation 2.7 writes the posterior as a product over experiments of univariate GP densities for zexp,i. Equation 2.11, presented as \"so that the probability density is,\" gives the joint multivariate normal with covariance Kh(θ,ϕ). Those are not equivalent. Under the KOH model, the experimental outputs are correlated through the shared emulator and discrepancy processes, so the joint density has off-diagonal covariances that the product ignores. The paper never flags this as an approximation. If the implementation uses Eq 2.7, the posterior intervals are from a misspecified likelihood; if it uses Eq 2.11, Eq 2.7 misstates the model. Either way the central claim that the framework \"quantifies uncertainty reliably\" is unsupported.\n\nTable 6 makes the problem concrete. Row 5: true Vgun = 57.58, posterior mean 53.83, chain sigma 0.028. That is over a hundred standard deviations off, and several other rows are similar. The text claims \"the true value lies within the standard deviation of the corresponding NUTS-MC chain.\" It does not. That is an internal contradiction, not a disagreement with a prior.\n\nThere are smaller issues: no MCMC convergence diagnostics, code/data only on request, and the runtime study uses chain length 100, which is short. Those are fixable. The likelihood misspecification is not.\n\nMy take: the computational-scaling half of the paper may survive a correction, but the accuracy half doesn't as written. The right reader is an applied statistician or domain scientist who needs fast, parallelizable calibration on large emulator data sets. This deserves a serious referee, not a desk reject—the combination is new and the scaling evidence is real. Send it back with a clear request to fix or justify the factorization, re-run the benchmarks and the AWA example under the corrected likelihood, address Table 6, and release code/data. If that comes back clean, it's a useful applied-statistics contribution.","headline":"Useful engineering combination with real scaling gains, but the experimental likelihood is factorized incorrectly and the accelerator table contradicts the text—send back for major revision.","tokens_in":22224,"tokens_out":3882,"would_cite":false,"duration_ms":38918,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","60G15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that splitting a hierarchical Kennedy–O'Hagan calibration's simulation data across a Bayesian committee of Gaussian processes cuts runtime by over an order of magnitude while preserving the inferred distribution of the…","keywords":["Bayesian calibration","hierarchical calibration","Bayesian Committee Machine","Gaussian process emulator","Kennedy–O'Hagan model","No-U-Turn Sampler","uncertainty quantification","particle accelerator"],"falsifier":"Re-run the Park-function or electron-gun calibration with the exact joint multivariate likelihood of Equation 2.11 in place of the factorized product of Equation 2.7, holding priors and data fixed, and compare the posterior of $\\sigma_\\theta$ and the per-experiment credible intervals; a detectable shift would confirm the factorization is load-bearing. A cheaper probe is to raise the number of experiments from 10 toward 50 while keeping the per-experiment design fixed: if the factorized posterior still matches the generating distribution, the independence assumption is innocuous, and if coverage or the recovered variance degrades, the missing cross-experiment correlations are the cause.","tokens_in":21193,"feed_emoji":"⚡","tokens_out":19730,"duration_ms":147883,"temperature":0.7,"pith_summary":"The paper sets out to make Kennedy–O'Hagan-style Bayesian calibration practical in the regime that breaks it: many experiments whose calibration parameters—such as an accelerator's beam-injection amplitude—drift between runs, plus a simulation dataset numbering in the thousands, where every Gaussian-process likelihood evaluation costs $O(n^3)$. It extends the standard model with a hierarchical prior $\\theta_i \\sim \\mathcal{N}(\\mu_\\theta, \\sigma_\\theta)$ that pools information across experiments, and it accelerates the likelihood with the Bayesian Committee Machine, which splits the simulation data into subsets, fits a Gaussian process to each, and multiplies the per-subset likelihoods into one corrected product. On the Park function, the cantilever beam, and a simulated electron-gun calibration, the framework recovers the true distribution of the per-experiment parameters while cutting runtime by over an order of magnitude—a 15.86-fold speed-up at 150 simulation points with eight committee members, with the runtime scaling falling from $O(n^{2.31})$ to $O(n^{1.24})$. The paper concludes that the BCM preserves inferential accuracy once each committee member holds enough points, making MCMC calibration with large simulation sets feasible. If the claim holds, calibration that otherwise takes hours moves inside the experimental operator's window.","feed_headline":"Hierarchical calibration runs 16x faster via a Bayesian committee","feed_subtitle":"Splitting simulation data across committee Gaussian processes keeps accuracy, cutting calibration from hours to minutes.","key_machinery":"The load-bearing object is the Bayesian Committee Machine product formula (Equation 2.13): $$p(y_*|D,x_*,\\theta_*) \\approx \\frac{\\prod_{i=1}^{N_{\\mathrm{MBC}}} p_{\\mathrm{GP}}(y_*|$D^{{(i)}}$,x_*,\\theta_*)}{p(y_*|x_*,\\theta_*)^{N_{\\mathrm{MBC}}-1}},$$ which converts one large Gaussian-process likelihood into a product of small ones, lowering each likelihood evaluation from $O(n^3)$ to $O((n/N_{\\mathrm{MBC}})^3)$ while still using every training point, and which parallelizes exactly across the committee members. The second mechanism is the hierarchical prior $\\theta_i \\sim \\mathcal{N}(\\mu_\\theta, \\sigma_\\theta)$ over experiment-specific calibration parameters, combined with the joint covariance matrix $K_h$ that adds the simulator-emulator kernel, the model-inadequacy kernel, and observation noise (Equation 2.11); this is what lets the sampler borrow strength across experiments and output a distribution of parameters instead of a single value, with a vine-based sampled correlation matrix for multidimensional parameters. The No-U-Turn Sampler—an adaptive Hamiltonian Monte Carlo algorithm that sets its own trajectory length—explores the resulting posterior, and because its gradients come from automatic differentiation, the statistical model can be specified freely without analytic derivative work.","core_discovery":"The central claim, argued through benchmark functions and a realistic accelerator case, is that the Bayesian Committee Machine's product-of-experts likelihood can be embedded inside hierarchical Kennedy–O'Hagan calibration without changing the distribution that MCMC samples. The simulation data is split into $N_{\\mathrm{MBC}}$ subsets; each subset is combined with the experimental data to form a Gaussian-process likelihood for the trial parameters, and the subset likelihoods are multiplied together and divided by the $(N_{\\mathrm{MBC}}-1)$-th power of the Gaussian-process prior so the variance is not underestimated (Equation 2.13). The hierarchical layer draws each experiment's parameter $\\theta_i$ from $\\mathcal{N}(\\mu_\\theta, \\sigma_\\theta)$, so the posterior yields both per-experiment values and the ensemble distribution; the paper shows that the recovered distribution matches the true generating distribution's moments and covariance on all three test problems, including a 2,400-run simulation set for an accelerator electron gun whose field amplitude must be inferred separately for each experiment. From this the paper concludes that the BCM delivers substantial computational gains while preserving inferential accuracy, and that because it partitions data rather than fitting an inducing-point or low-rank approximation, it avoids biases such approximations can introduce.","pith_inferences":["The factorization in Equation 2.7 (a product over experiments of univariate Gaussian-process densities) is never flagged in the paper as an approximation, yet it omits the cross-experiment correlations present in the joint covariance of Equation 2.11; if those correlations are informative, the per-experiment uncertainties and the recovered $\\sigma_\\theta$ could be biased, and an exact-likelihood r","The paper itself concedes (Section 2.4.1) that the BCM has no proven convergence guarantees outside a small-query regime and that dependence between committee members can distort uncertainty estimates; since the accuracy results are empirical, a conservative transfer rule for new applications is to validate against the full-GP calibration on a subset of the data and to keep each committee member a","The same product-of-experts construction should transfer to any inverse problem that repeatedly evaluates a Gaussian-process likelihood against a large, splittable simulation set, such as other accelerator components or physics simulations with expensive forward models, because nothing in the derivation is specific to beam physics.","Replacing the random split with a domain-informed partition and replacing the fixed Gaussian hierarchical density with a heavier-tailed or mixture family would directly address the two acknowledged sources of bias—committee dependence and parametric misspecification of the ensemble distribution—that the paper flags in Sections 2.4.2 and 4."],"forward_implications":["With eight committee members instead of one, the reported runtime speed-up reaches 15.86× at 150 simulation points and the scaling exponent drops from $O(n^{2.31})$ to $O(n^{1.24})$, so doubling the committee roughly allows doubling the simulation set at the same expected cost.","The hierarchical layer recovers the ensemble distribution of the calibration parameters rather than a pooled mean: the inferred $\\mu_\\theta$ and $\\sigma_\\theta$ match the generating distribution's moments on the Park and accelerator examples, and the two-dimensional cantilever case additionally recovers the covariance structure between parameter dimensions.","Accuracy is preserved only when each committee member sees enough data: the paper finds the error on the inferred moments grows when subsets fall below about 20 simulation points, setting a practical floor on how finely to split.","Calibration runtime drops from roughly an hour to about five minutes in the demonstrated regime, which is the difference between an offline analysis and an update an accelerator operator can run between measurements.","The BCM partition approach keeps all simulation points in use rather than selecting a subset, so larger training sets reduce the emulator error without a proportional runtime penalty once the committee structure absorbs the extra points."],"supporting_citations":[{"why":"Supplies the Kennedy–O'Hagan calibration formulation—simulator emulator plus model-inadequacy term—that the paper extends with a hierarchical layer.","marker":"[2]"},{"why":"Provides the MCMC-based implementation of BayesCal and the hyperparameter priors that the framework adopts unchanged.","marker":"[3]"},{"why":"Introduces the Bayesian Committee Machine and the product-of-experts formula (Equation 2.13) on which the entire speed-up rests.","marker":"[20]"},{"why":"Generalized Robust BCM; its analysis of variance underestimation motivates the denominator correction in the combination formula.","marker":"[31]"},{"why":"The hierarchical calibration reference from which the latent distribution over experiment-specific parameters is adopted.","marker":"[11]"},{"why":"Justifies the uniform prior on the hierarchical variance, shaping how $\\sigma_\\theta$ is estimated.","marker":"[22]"},{"why":"Defines the No-U-Turn Sampler used to explore the hierarchical posterior.","marker":"[24]"},{"why":"Provides the probabilistic programming framework in which the model is specified and sampled.","marker":"[25]"},{"why":"The accelerator simulation code that generated the synthetic training and observation data for the electron-gun application.","marker":"[36]"}],"fun_headline_variants":["Bayesian committee machine runs hierarchical calibration 16x faster","Hierarchical calibration with Bayesian committee: 16x speedup, same accuracy","Divide-and-conquer Gaussian processes accelerate accelerator calibration","Parallel emulator committee makes hierarchical Bayesian calibration practical","Committee machine cuts hierarchical calibration time without losing accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The posterior in Equation 2.7 treats each experiment's measured output as its own independent Gaussian-process likelihood factor, discarding the correlations among experimental outputs that the joint covariance matrix in Equation 2.11 encodes; if those cross-experiment correlations carry information, the sampled posterior is misspecified and the reported per-experiment uncertainties are not strictly valid.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian committee machine runs hierarchical calibration 16x faster","Hierarchical calibration with Bayesian committee: 16x speedup, same accuracy","Divide-and-conquer Gaussian processes accelerate accelerator calibration","Parallel emulator committee makes hierarchical Bayesian calibration practical","Committee machine cuts hierarchical calibration time without losing accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3227,"prompt_tokens":1039,"completion_tokens":2188,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":655,"completion_tokens_details":{"reasoning_tokens":2123}},"tokens_in":655,"tokens_out":2188,"duration_ms":13926,"temperature":1.0,"reasoning_tokens":2123,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:04:38.696074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Park-function or electron-gun calibration with the exact joint multivariate likelihood of Equation 2.11 in place of the factorized product of Equation 2.7, holding priors and data fixed, and compare the posterior of $\\sigma_\\theta$ and the per-experiment credible intervals; a detectable shift would confirm the factorization is load-bearing. A cheaper probe is to raise the number of experiments from 10 toward 50 while keeping the per-experiment design fixed: if the factorized posterior still matches the generating distribution, the independence assumption is innocuous, and if coverage or the recovered variance degrades, the missing cross-experiment correlations are the cause.","supporting_citations":[{"cited_title":"Turing: A Language for Flexible Probabilistic Inference","cited_arxiv_id":null,"evidence_quote":"Provides the probabilistic programming framework in which the model is specified and sampled."},{"cited_title":"Combining Field Data and Computer Simulations for Calibration and Predic- tion","cited_arxiv_id":null,"evidence_quote":"Provides the MCMC-based implementation of BayesCal and the hyperparameter priors that the framework adopts unchanged."},{"cited_title":"A Bayesian Committee Machine","cited_arxiv_id":null,"evidence_quote":"Introduces the Bayesian Committee Machine and the product-of-experts formula (Equation 2.13) on which the entire speed-up rests."},{"cited_title":"Generalized Robust Bayesian Committee Machine for Large-scale Gaussian Process Regression","cited_arxiv_id":null,"evidence_quote":"Generalized Robust BCM; its analysis of variance underestimation motivates the denominator correction in the combination formula."},{"cited_title":"A unified framework for multilevel uncertainty quantification in Bayesian inverse problems","cited_arxiv_id":null,"evidence_quote":"The hierarchical calibration reference from which the latent distribution over experiment-specific parameters is adopted."},{"cited_title":"Prior distributions for variance parameters in hierarchical models (comment on article by Browne and Draper)","cited_arxiv_id":null,"evidence_quote":"Justifies the uniform prior on the hierarchical variance, shaping how $\\sigma_\\theta$ is estimated."},{"cited_title":"The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo","cited_arxiv_id":null,"evidence_quote":"Defines the No-U-Turn Sampler used to explore the hierarchical posterior."}],"review_version":1}