{"id":"c96a045a-fa0e-4ce8-b85e-fc15a0ebb1ac","arxiv_id":"2507.18683","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A Bayesian deep Gaussian process with correlated functional noise estimates infinite-volume matter power spectra and, via PCA and GP emulation, predicts spectra at unobserved cosmologies.","lead":"This paper introduces a Bayesian deep Gaussian process model that estimates the underlying matter power spectrum of the universe from correlated simulation outputs at different fidelities. It also builds an emulator that predicts complete power spectra for new cosmological parameters, and reports competitive results against the existing Cosmic Emu benchmark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Mira-Titan emulator comparison cannot support 'competes favorably': α=1.95 is tuned on the six held-out cosmologies, and both methods are scored against their own in-sample posterior means rather than a common target.","rationale":"The reader's verdict is CONDITIONAL with high confidence, and the rationale already identifies the test-data leakage and the self-referential in-sample benchmark as compromising the Mira-Titan comparison. My stress-test agrees that these are the load-bearing weaknesses for the central comparative claim. However, the reader's formal 'weakest_assumption' field points to a different issue: the sufficiency and covariance specification of the weighted average Ȳ in Sections 3.1 and 3.5. That is a legitimate concern about the estimation stage, but it is not the assumption that most directly underpins the paper's headline 'competes favorably with Cosmic Emu' claim; the emulator comparison does. The covariance specification is also partially mitigated by the CAMB validation, where the DGP.FCO posterior mean is checked against a true infinite-resolution spectrum, and by the synthetic study with known truth. By contrast, the Mira-Titan prediction comparison has no common ground truth at all, and the α tuning leakage is an unambiguous methodological flaw. Thus I regard the benchmarking protocol as the single most load-bearing concern, and the reader's weakest_assumption as related but not identical. My recommended verdict is therefore unchanged from the reader's CONDITIONAL: the estimation contribution can stand, but the comparative emulation claim should not be stated without either fixing the protocol or reframing the claim as a demonstration of competitive performance conditional on a chosen in-sample benchmark.","tokens_in":16553,"tokens_out":4043,"duration_ms":47329,"concrete_test":"Re-run the Mira-Titan prediction exercise with α selected by leave-one-out cross-validation on the 111 training cosmologies only (or fixed a priori), and score both DGP.FCO+PC and Cosmic Emu against the same common reference, e.g., the high-resolution spectrum P_h(k) on 0.04<k<5 Mpc^-1, where it is approximately unbiased. Recompute the k-by-k MSE comparison and the '88% of k values' statistic under this corrected protocol. If DGP.FCO+PC still attains lower MSE at ≥88% of k values against the common reference with an honestly chosen α, the 'competes favorably' claim is supported; otherwise the claim should be softened or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that DGP.FCO 'competes favorably' with Cosmic Emu rests on the Mira-Titan prediction exercise in Section 4.3. Two protocol choices undermine it. First, Section 4.1 fixes α=1.95 'by minimizing MSE across the six held-out Mira-Titan cosmologies over a grid of candidate values.' This is direct test-data leakage: the emulator's kernel shape is tuned on the same six cosmologies later used for evaluation, so the reported out-of-sample MSE is not genuinely out-of-sample. Since Cosmic Emu's settings are not re-tuned on these hold-outs, the comparison is biased in favor of DGP.FCO+PC. Second, Section 4.3 scores predictions against each method's own in-sample posterior mean: DGP.FCO+PC predictions are centered by the in-sample DGP.FCO posterior mean, and Cosmic Emu predictions by the in-sample DPC posterior mean. MSE relative to two different, method-specific references is not a comparison of accuracy to the underlying infinite-volume spectrum. A smoother in-sample estimate will automatically appear closer to a smooth out-of-sample prediction, so the '88% of k values' statistic in Figure 13 does not establish that DGP.FCO is a better predictor of P∞(k). The paper itself acknowledges these in-sample spectra 'are not formal truths,' but the headline claim is nevertheless drawn from this comparison. The core estimation contribution (Sections 3.3-3.5, CAMB validation) is not affected by this concern; only the comparative emulation claim is unsupported as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian deep Gaussian process model for correlated functional data, called DGP.FCO, and applies it to cosmological matter power spectra. For a given cosmology, the model synthesizes perturbation-theory spectra, sixteen low-resolution simulations, and one high-resolution simulation into an estimate of the infinite-volume power spectrum P∞(k) with quantified uncertainty. The estimation is validated on synthetic data and on CAMB, where a true infinite-resolution spectrum is available. The paper then uses the estimated posterior means for training cosmologies to build a principal-component Gaussian process emulator that predicts spectra at unobserved cosmologies, comparing its Mira-Titan predictions with the Cosmic Emu emulator. The paper claims improved estimation and uncertainty quantification over a weighted average and a standard GP, and states that DGP.FCO competes favorably with Cosmic Emu on held-out cosmologies.","tokens_in":16885,"tokens_out":5241,"duration_ms":52279,"significance":"If the central claims hold, the paper makes a useful contribution to nonstationary functional emulation with uncertainty quantification for multi-fidelity cosmological outputs. The manuscript has notable strengths: the closed-form posterior derivations in Appendices A and B are sound; the code, data, and R package are publicly available; and the CAMB experiment provides an external ground truth, showing that DGP.FCO reduces mean MSE from 6.7e-5 (weighted average) to 4.3e-5 (standard GP with correlated functional outputs) to 3.0e-5 (DGP.FCO) across 32 cosmologies. However, the headline comparative claim against Cosmic Emu is not supported as presented because the comparison in Section 4.3 uses test-data leakage in the tuning of the kernel exponent and scores predictions against method-specific in-sample means rather than a common target. The core estimation contribution, validated on CAMB, is not affected by these concerns.","major_comments":[{"comment":"The kernel exponent α=1.95 is selected by minimizing MSE across the six held-out Mira-Titan cosmologies, and those same six cosmologies are used in Section 4.3 to claim favorable comparison with Cosmic Emu. Since Cosmic Emu's kernel settings are not re-tuned on these hold-outs, this leaks test information into DGP.FCO+PC and biases the comparison in its favor. This directly undermines the claim that DGP.FCO 'competes favorably with the state-of-the-art Cosmic Emu model' (Section 4.3, final paragraph). The comparison should be rerun with α chosen without touching the hold-outs, for example by cross-validation within the 111 training cosmologies or by fixing α to a standard value such as 2.","section":"Section 4.1"},{"comment":"The Mira-Titan out-of-sample predictions are evaluated against each method's own in-sample posterior mean: DGP.FCO+PC predictions are centered by the in-sample DGP.FCO posterior mean, and Cosmic Emu predictions by the in-sample DPC posterior mean. MSE relative to two different, method-specific references does not measure accuracy for the common target P∞(k), and a smoother in-sample estimate will automatically appear closer to a smooth out-of-sample prediction. The paper itself acknowledges in this section that these in-sample spectra 'are not formal truths.' Consequently, the '88% of k values' statistic in Figure 13 does not establish that DGP.FCO predicts the infinite-volume spectrum better than Cosmic Emu. To support the comparative claim, both methods should be scored against a common target, such as the high-resolution run on the wavenumber range where it is unbiased, or the CAMB external-validation protocol should be used as the primary evidence.","section":"Section 4.3, Figures 12-13"},{"comment":"The Mira-Titan uncertainty quantification relies on Σε = (Λp + Σℓ^{-1} + Λh)^{-1}, where Σℓ is a Matérn covariance estimated separately from the low-resolution runs and referenced to the Walsh (2023) thesis. Unlike the CAMB experiment in Section 3.4, where Σε is diagonal and an external truth exists, the correlated-noise case in Mira-Titan has no ground truth to check calibration, and the hyperparameters of Σε are fixed rather than inferred within the Bayesian model. Section 5 itself lists estimation of Σε hyperparameters within the sampling scheme as 'an area for further investigation.' The credible intervals for Mira-Titan should therefore be described as conditional on an externally estimated Σε, and the UQ claims for the dense-covariance setting should not be stated as validated by the CAMB experiment, which uses a different (diagonal) noise structure.","section":"Section 3.5"}],"minor_comments":[{"comment":"In the integrated-likelihood derivation, the displayed exponential contains 'YT i Σ−1 ε Y' and 'Σ−1 ε Y' where the subscript i on Y is missing in two places; this should be corrected to Yi for clarity.","section":"Appendix A"},{"comment":"The reference for Higdon et al. (2010) includes a stray leading number '749' in the title field; it should be removed.","section":"Reference list"},{"comment":"The caption reads 'MSE across each of the 6 hold out cosmologies'; it should be 'hold-out cosmologies'.","section":"Figure 13 caption"},{"comment":"The choice of Σℓ and the statement that a Matérn covariance trained on low-resolution runs 'performs the best' is delegated to Walsh (2023); since this choice is load-bearing for the Mira-Titan estimation, the manuscript should at least summarize the candidate set and selection criterion used to reach that conclusion.","section":"Section 3.5"},{"comment":"The phrase 'with theGPfitR-package' should be formatted as 'with the GPfit R package' for readability.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The core statistical contribution—the DGP.FCO hierarchical model with correlated functional outputs and the closed-form conditional posterior—is sound and is well supported by the synthetic and CAMB experiments. The problem is localized to the comparative Mira-Titan emulation claims in Section 4.3, which are presented as headline results but are undermined by leakage and method-specific scoring. If the authors reframe these claims or fix the protocol, the manuscript could be suitable for publication. I would also encourage the authors to provide more detail on the external estimation of Σℓ rather than deferring entirely to the Walsh thesis, as the reproducibility of the Mira-Titan UQ depends on it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere’s the takeaway: the paper’s first contribution—estimating a single infinite-volume power spectrum from correlated, multi-fidelity simulation curves via a deep GP with dense noise covariance—is real and mostly holds up. The second contribution, the claim that DGP.FCO “competes favorably” with Cosmic Emu on held-out Mira-Titan cosmologies, does not survive scrutiny as currently written. The fix is straightforward, and the first half is worth engaging with regardless.\n\nWhat is new: previous Bayesian DGPs handle scalar responses; adding a dense observation-noise covariance Σε to the outer GP likelihood and conditioning on a precision-weighted average of curves is a sensible extension. The closed-form posterior for S|Ȳ is correctly derived in Appendices A and B. The CAMB experiment has a true infinite-resolution spectrum as ground truth and shows a clean MSE reduction: 6.7e-5 for the weighted average, 4.3e-5 for a standard GP.FCO, 3.0e-5 for DGP.FCO, better on all 32 held-out cosmologies. Code, data, and an R package are public. That is solid empirical validation of the estimation model.\n\nThe soft spot is the Mira-Titan emulator comparison. Section 4.1 fixes α=1.95 by minimizing MSE on the same six held-out cosmologies later used for evaluation. That is test-data leakage; the “out-of-sample” numbers are not out-of-sample for that hyperparameter. To the authors’ credit, they state this openly, but it still biases the comparison in their favor. Section 4.3 then scores each method against its own in-sample posterior mean—DGP.FCO+PC against the DGP.FCO fit, Cosmic Emu against the DPC fit. With different references, the MSE comparison is not between two predictors of the same target. The paper admits these in-sample spectra “are not formal truths,” but the headline claim sits on exactly this comparison. So “competes favorably” is not supported as presented.\n\nAlso worth a look: Σε = (Λp + Σℓ^{-1} + Λh)^{-1} is a constructed noise covariance, not the covariance of the weighted average Ȳ one would get from the component precisions. For UQ claims, it would help to state explicitly that this is a modeling choice and to check coverage of the credible intervals against the CAMB truth, not just MSE.\n\nFixes: refit α by cross-validation or a separate validation set, and compare both emulators to a common reference—say the high-resolution run on its valid k range, or one fixed in-sample estimator. The core estimation contribution should survive those changes.\n\nThis paper deserves a serious referee. I would send it to review with a request to fix the emulator comparison and clarify the Σε construction, and I would cite the DGP.FCO estimation part in my own work.","headline":"Solid DGP extension for correlated functional outputs with a credible CAMB validation, but the headline Mira-Titan comparison with Cosmic Emu is undermined by tuning α on the held-out cosmologies and comparing against method-specific in-sample means.","tokens_in":17536,"tokens_out":4917,"would_cite":true,"duration_ms":46825,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian deep Gaussian process combines perturbation theory and simulation outputs to estimate the universe's matter power spectrum with quantified uncertainty, and its emulator matches the leading Cosmic Emu benchmark.","keywords":["deep Gaussian process","functional data","matter power spectrum","uncertainty quantification","emulator","Mira-Titan","Cosmic Emu","principal components analysis"],"falsifier":"Take a cosmology where an independent, larger simulation gives the true infinite-volume spectrum, generate many independent batches of the low- and high-resolution runs at that cosmology, and count how often the DGP.FCO 95% credible intervals contain the truth at each wavenumber; if the rate is far from 95%, the assumed error structure is wrong.","tokens_in":16284,"feed_emoji":"🔭","tokens_out":6837,"duration_ms":59400,"temperature":0.7,"pith_summary":"The paper tries to establish that a deep Gaussian process can synthesize three sources of cosmological information—linear perturbation theory, sixteen low-resolution simulation runs, and one high-resolution run—into an estimate of the infinite-volume matter power spectrum with valid uncertainty. It argues that accounting for smooth dependence across wavenumbers, not just pointwise precision, is essential for this estimate, and that the resulting posterior mean is accurate enough to serve as training data for an emulator at unobserved cosmologies. The payoff, if right, is a practical recipe for nonstationary functional emulation with uncertainty quantification in a setting where simulations are expensive and no single curve is trustworthy over the whole wavenumber range.","feed_headline":"Bayesian deep GP rivals Cosmic Emu on matter power spectra","feed_subtitle":"A new model fuses theory and simulations into one spectrum estimate with quantified uncertainty, then predicts unseen cosmologies.","key_machinery":"The central object is the dense correlated-noise covariance $\\Sigma_\\varepsilon = (\\Lambda_p + \\Sigma_\\ell^{-1} + \\Lambda_h)^{-1}$, where $\\Lambda_p$ and $\\Lambda_h$ are diagonal precision matrices encoding which wavenumbers each data source can be trusted and $\\Sigma_\\ell$ is a Matérn covariance estimated from the low-resolution runs. This matrix lets the model treat the sixteen low-resolution curves and the high-resolution curve as smooth correlated functional realizations rather than independent points. It is combined with a deep Gaussian process prior on the spectrum: a latent monotonic warp layer $W$ warps the wavenumber inputs, so the outer Matérn covariance $\\Sigma_S(W)$ is evaluated at warped locations and can stretch or compress correlation structure across $k$. The closing identity is the integrated likelihood $\\bar{Y}|W \\sim \\mathcal{GP}(0,\\Sigma_S(W)+\\Sigma_\\varepsilon)$, which permits elliptical slice sampling for $W$ and yields a closed-form conditional posterior for $S$.","core_discovery":"For each cosmology, the paper models the precision-weighted average of the 18 spectra as a Gaussian process centered on the unknown infinite-volume spectrum $S$. The noise covariance is dense: it inverts the sum of a Matérn covariance estimated from the low-resolution runs and diagonal precisions encoding which wavenumbers perturbation theory and the high-resolution run can be trusted. $S$ itself is given a deep Gaussian process prior through a latent monotonic warp, letting the reconstructed spectrum be smooth where dynamics are linear and more variable where baryonic acoustic oscillations appear. Conditioning on the data yields a Gaussian posterior for $S$ in closed form, so posterior means and 95% credible intervals are obtained by direct sampling. On CAMB data, where the truth is known, this posterior mean cuts MSE roughly in half relative to the weighted average and adds a further 30% reduction over a shallow GP; on Mira-Titan held-out cosmologies, the principal-component emulator built on these posterior means has lower MSE than Cosmic Emu at 88% of wavenumbers.","pith_inferences":["A direct testable extension would be to run DGP.FCO on many independent simulation seeds at a single cosmology and check empirical coverage of the credible intervals; the paper's validation is mostly against synthetic and CAMB truth, not repeated Mira-Titan runs.","Because the covariance $\\Sigma_\\varepsilon$ is fixed before sampling rather than learned jointly, there may be room to estimate its Matérn parameters inside the MCMC, which would fold uncertainty about the correlation structure into the final intervals.","The per-cosmology fitting strategy ignores borrowing of strength across cosmologies at the spectrum-estimation stage; a hierarchical version that shares warp parameters or covariance parameters across cosmologies could improve predictions at held-out points with few neighbors.","The PC emulator uses only the posterior means, discarding the posterior covariance of each fitted spectrum; feeding posterior draws through the PC decomposition would produce a full predictive distribution for unobserved cosmologies rather than a point prediction."],"forward_implications":["With the dense correlated-noise covariance, the posterior mean of the spectrum outperforms the weighted average and a shallow GP on CAMB truth, so both the smoothing step and the deep layer earn their place in the pipeline.","The PC-GP emulator trained on DGP.FCO posterior means predicts held-out Mira-Titan cosmologies with lower MSE than Cosmic Emu at 88% of wavenumbers, and matches it elsewhere.","The same fitted spectra can be reused as training data for any emulator that maps cosmological parameters to functional outputs, not only this paper's principal-component model.","Within each cosmology the model produces a closed-form Gaussian posterior for the spectrum, so credible intervals and posterior draws are cheap to obtain after the MCMC for the warp is done.","Nonstationarity is handled by a latent monotonic warp rather than by changing the kernel family, which generalizes to other smooth functional responses with scale-dependent variability."],"supporting_citations":[{"why":"Introduces deep Gaussian processes, the compositional prior the model extends to correlated functional outputs.","marker":"Damianou and Lawrence (2013)"},{"why":"Supplies the MCMC sampling scheme and prior defaults for the deep GP layers.","marker":"Sauer et al. (2023b)"},{"why":"Provides the Mira-Titan dataset, the precision regression estimates, and the Cosmic Emu benchmark.","marker":"Moran et al. (2023)"},{"why":"Supplies the earlier comparison that selects the Matérn covariance for the low-resolution runs used in $\\Sigma_\\varepsilon$.","marker":"Walsh (2023)"},{"why":"Describes the reduction of simulation density fields to matter power spectra that defines the data object.","marker":"Heitmann et al. (2010)"},{"why":"The PCA basis-decomposition approach the emulator uses to turn functional outputs into scalar weight regressions.","marker":"Higdon et al. (2008)"},{"why":"Provides the monotonic warp layer that gives the deep GP its nonstationary flexibility.","marker":"Barnett et al. (2025)"},{"why":"Supplies CAMB, the code whose infinite-resolution spectra give the paper an external truth for validation.","marker":"Lewis and Challinor (2011)"}],"fun_headline_variants":["Bayesian deep GP cuts spectrum error by half","Deep GP emulator surpasses Cosmic Emu on power spectra","Deep GP for correlated functional data: sharper spectra","New Bayesian DGP beats Cosmic Emu on 88% of modes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole calculation rests on treating one combined average of the 18 simulation curves, together with a pre-chosen model of how errors in those curves relate to one another, as containing all the useful information about the true spectrum; if that combination discards information or gets the error relationships wrong, the uncertainty intervals will not mean what they claim.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian deep GP cuts spectrum error by half","Deep GP emulator surpasses Cosmic Emu on power spectra","Deep GP for correlated functional data: sharper spectra","New Bayesian DGP beats Cosmic Emu on 88% of modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00085,"raw_usage":{"total_tokens":3713,"prompt_tokens":975,"completion_tokens":2738,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":2671}},"tokens_in":591,"tokens_out":2738,"duration_ms":22587,"temperature":1.0,"reasoning_tokens":2671,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:10:16.952562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a cosmology where an independent, larger simulation gives the true infinite-volume spectrum, generate many independent batches of the low- and high-resolution runs at that cosmology, and count how often the DGP.FCO 95% credible intervals contain the truth at each wavenumber; if the rate is far from 95%, the assumed error structure is wrong.","supporting_citations":[{"cited_title":"The coyote universe. I. Precision determination of the nonlinear matter power spectrum","cited_arxiv_id":null,"evidence_quote":"Describes the reduction of simulation density fields to matter power spectra that defines the data object."},{"cited_title":"Bayesian \"Deep\" Process Convolutions: An Application in Cosmology","cited_arxiv_id":"2411.14747","evidence_quote":"Supplies CAMB, the code whose infinite-resolution spectra give the paper an external truth for validation."}],"review_version":2}