{"id":"302abe51-5d19-42fa-ac31-2f393fce1545","arxiv_id":"2412.07306","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A review chapter that surveys Gaussian process models and sequential sampling strategies for stochastic simulators with input-dependent or non-Gaussian noise.","lead":"This chapter reviews methods for building Gaussian process surrogates for computer simulators that produce noisy outputs. It is a useful entry point for practitioners choosing among constant-noise, heteroscedastic, quantile, and deep-GP models, and for deciding when to replicate runs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's practical recommendations are carried by a single 1D SIR illustration with no quantitative metrics or repeated-run variability; if that example is atypical, the chapter's normative claims lose support.","rationale":"I read the chapter in good faith as a survey. Its central claim—that GP surrogates can be adapted to stochastic simulators through heteroscedastic modeling and replication-aware sequential design—is supported by a broad and mostly appropriate citation list, and the displayed prediction equations in Section 2 are algebraically consistent. No internal contradiction or mathematical error rose to the level of a fatal flaw. The most load-bearing weakness is the chapter's own empirical component: the qualitative statements that give the survey its practical orientation rest on a single 1D SIR example, with two datasets, no error bars, no repeated-seed analysis, and no quantitative metrics in Sections 2.3, 3.2, and 4.4. The reader's weakest-assumption field identifies this same point. Because the chapter explicitly labels these as illustrations, the absence of rigorous validation does not invalidate the survey's existence claims; it does mean that any normative reading ('this is how one should gear GP modeling toward noisy simulators') is under-supported. The appropriate disposition remains UNVERDICTED: the chapter can be evaluated as a literature review rather than a research contribution, but its illustrative evidence cannot support strong practical conclusions. Thus I do not change the reader's verdict.","tokens_in":15093,"tokens_out":8007,"duration_ms":74525,"concrete_test":"Re-run the illustrations in §2.3, §3.2, and §4.4 on the same SIR simulator for at least 20 independent simulator-seed realizations (same designs) and report mean ± standard error of: RMSE of predicted mean and variance, 90% predictive-interval coverage, and terminal IMSPE/SUR values, for homoscedastic GP, hetGP, stochastic kriging, and quantile GP. Repeat on a second simulator with nearly constant noise. If rankings reverse across seeds or vanish on the constant-noise simulator, the chapter's practical recommendations must be qualified as example-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The chapter's central claim is largely existential: methods for heteroscedastic GP modeling, quantile and deep GPs, and sequential design with replication exist and have been adapted. The equations in §2 are correctly stated, and the cited literature provides independent support for existence. The load-bearing weakness is the original evaluative evidence. Section 2.3 fits homoscedastic and heteroscedastic models to one SIR data set (and one no-replication variant) and concludes, visually, that the heteroscedastic model 'better represents the black-box at hand'; Section 3.2 asserts that the latent quantile model 'manages to fit the actual quantiles well, except at the origin' and that deep GP predictions are 'not significantly different' across data sets, with no statistical test or error bar; Section 4.4 states that IMSPE and contour-SUR strategies 'both succeed in improving the initial estimate towards the goal' without reporting a quantitative metric. Because the chapter is a review, these illustrations are not the sole support for existence of the methods, but they are the only support for the chapter's practical orientation—that these adaptations are worthwhile. A single 1D SIR simulator with strongly input-dependent variance may be unrepresentative of stochastic simulators whose noise is nearly homoscedastic, heavy-tailed, or only mildly input-dependent. The absence of repeated-seed variability means the reader cannot distinguish genuine differences from artifacts of one realization. Thus the weakest assumption is not a mathematical error but an empirical one: the illustrative example is assumed to be representative enough to carry qualitative recommendations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This chapter surveys Gaussian process (GP) surrogates for stochastic simulators with complex, input-dependent noise. It reviews homoscedastic and heteroscedastic GP regression, stochastic kriging with replication, quantile GPs, deep GPs, and adaptations of sequential design criteria (IMSPE, SUR, EI/UCB variants, replication allocation) to noisy settings. A running 1D SIR example with two data sets, one with replication and one without, illustrates the models and design strategies. The paper's central assertion is that each of these modeling and design frameworks has been developed in the literature and can be deployed depending on data availability and replication structure.","tokens_in":15336,"tokens_out":7237,"duration_ms":70973,"significance":"If taken as a survey, the chapter fills a useful niche by organizing a scattered literature around the replication/no-replication distinction and by presenting predictive equations in a unified notation. Its strengths are breadth, accurate attributions (I found no mis-attributed equations), and practical pointers to implementations such as hetGP, deepgp, and GPyTorch. The chapter does not introduce new methodology, and its original comparative evidence is anecdotal, relying on visual inspection of a single simulator without quantitative metrics or repeated runs. Thus the paper is a competent review whose practical recommendations would need more evidence to be considered established.","major_comments":[{"comment":"The original illustrative experiments are the only place where the paper moves beyond cataloguing existing methods, yet all performance claims are made from visual inspection of a single 1D SIR example. In §2.3 the conclusion that the heteroscedastic model \"better represents the black-box at hand\" is not supported by any quantitative metric (e.g., RMSE, coverage, log predictive density) or by variability across random seeds; in §3.2 the statement that the latent quantile model \"manages to fit the actual quantiles well, except at the origin\" and the claim that deep GP predictions are \"not significantly different\" across data sets lack statistical or numerical support; in §4.4 \"both succeed in improving the initial estimate towards the goal\" is not accompanied by a reported improvement. These claims should either be backed by quantitative results or be explicitly framed as anecdotal; as written, the paper's practical recommendations rest on an unvalidated example.","section":"§2.3, §3.2, §4.4"},{"comment":"The sequential design illustration reports only absolute outcomes (≈120 unique designs for IMSPE after 900 evaluations, ≈30 iterations for SUR to reach 1000 evaluations) and does not compare against simple baselines such as a space-filling design without replication, a homoscedastic GP, or a constant-replication strategy. Without such comparisons, the reader cannot judge whether replication-aware strategies actually provide the claimed benefit in this example. Adding a baseline, even for one example, would substantially strengthen the section's conclusions.","section":"§4.4"}],"minor_comments":[{"comment":"There is a typo: \"heteroscedasic\" should be \"heteroscedastic\".","section":"§2.3"},{"comment":"The phrase \"it as been proposed\" should be \"it has been proposed\".","section":"§3.1.1"},{"comment":"The sentence \"dedicated criteria have also been obtained in closed form, for multi- van der Herten et al. (2016)\" is truncated; it should read \"for multi-objective optimization\".","section":"§4.2"},{"comment":"The phrase \"distinguish between aleatoric uncertainty – arising from noise in observations, from epistemic uncertainty\" reads awkwardly; consider inserting \"and\" before \"from epistemic uncertainty\" or restructuring the sentence.","section":"§1"},{"comment":"The sentence \"This even shows for the reference on the variance and skewness estimation\" is unclear; clarify what \"this\" refers to.","section":"§2.3"},{"comment":"Equations (1)–(2) would be easier to follow if c(x) and C_N were explicitly defined as the correlation vector and correlation matrix, since the notation c(x,x) as a variance function may be confusing when paired with the outer noise term r(x).","section":"§2.1"},{"comment":"The label \"IMPSE infill criterion value\" should be \"IMSPE infill criterion value\" in the figure and caption.","section":"Figure 4"}],"recommendation":"minor_revision","confidential_remarks":"This manuscript reads as a book chapter rather than a standard research article. If the intended venue is a journal, the novelty is limited to the organization and the illustrative example, so the editor should ensure the scope matches the journal's expectations. The review appears careful and the citations, while naturally favoring the authors' own hetGP line of work, are appropriate for a survey. The main revision concern is the illustrative evidence: either strengthen it with quantitative metrics and baselines or soften the claims so that they are clearly anecdotal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a survey/tutorial chapter, not a research contribution, and it is a solid one. The equations and attributions check out, and the organization is genuinely useful. The main weakness is the illustrative example—a single 1D SIR simulator, visual comparisons only, no error bars or repeated-seed variability—so the practical recommendations are directionally sensible but not empirically established.\n\nThe chapter does a good job consolidating the literature on GP surrogates for stochastic simulators: heteroscedastic GPs, stochastic kriging, quantile/expectile models, deep GPs, and the corresponding sequential design criteria (IMSPE, SUR, EI/UCB variants). I found the discussion of replication versus latent-variable approaches helpful, and the trade-offs in Section 5 are honest. It also points to the original sources carefully, including the authors' own work where appropriate; the citations are not self-serving.\n\nThe soft spots are where the stress-test note lands. Section 2.3 fits homoscedastic and heteroscedastic models to one dataset and concludes visually that the heteroscedastic model 'better represents the black-box.' Section 3.2 says the deep GP predictions are 'not significantly different' between datasets without any statistical test. Section 4.4 reports that IMSPE and contour-SUR 'both succeed' without quantitative metrics. That language implies more evidentiary weight than a single illustrative run can carry. It would be straightforward to add repeated-seed variability, quantitative comparisons (e.g., coverage, RMSE, budget used), or at least hedge the claims. The absence of those doesn't break the survey, because the existence of these methods does not depend on the example, but it does weaken the authors' practical orientation.\n\nThe paper is best read as an entry point for practitioners and graduate students who want a concise map of the noisy-GP landscape. It is not a research paper and should not be judged as one. I would accept it for peer review as a review article, with the recommendation that the empirical illustration be strengthened or explicitly labeled as illustrative only.","headline":"A competent, useful survey of noisy-GP modeling and sequential design; the only real weakness is the thin illustrative example.","tokens_in":15897,"tokens_out":2676,"would_cite":true,"duration_ms":27348,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62-02","62K05","62M30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Gaussian process surrogates extend to stochastic simulators by modeling input-dependent noise and adapting sequential design criteria to replication.","keywords":["Gaussian process surrogates","stochastic simulators","heteroscedastic noise","replication","sequential design","quantile Gaussian processes","Bayesian optimization","stochastic kriging"],"falsifier":"Run a systematic benchmark on several stochastic simulators with different signal-to-noise ratios, replication costs, and non-Gaussian error structures, and compare quantitative mean-squared error and budget-accuracy curves; the chapter's visual illustrations alone do not settle whether the claimed ranking of methods holds.","tokens_in":14897,"feed_emoji":"🎲","tokens_out":10039,"duration_ms":98995,"temperature":0.7,"pith_summary":"This chapter argues that Gaussian process surrogates remain practical and analytically tractable for stochastic simulators when the noise is treated as input-dependent rather than constant. It maps the main modeling options, from heteroscedastic GPs with a latent log-variance process to stochastic kriging, quantile GPs, and deep GPs, and shows what each requires in terms of replication and inference. It then shows how sequential design criteria such as IMSPE, SUR, EI, and UCB can be adapted to decide where to sample and how often to replicate. The reason this matters is that stochastic simulators are common in epidemiology, operations research, and hyperparameter tuning, and deciding whether to replicate or explore is a real budget question.","feed_headline":"Heteroscedastic GP surrogates tame noise in stochastic simulators","feed_subtitle":"A review maps modeling options — replication, latent variance, quantile GPs — and the sequential designs that match them.","key_machinery":"The central object is the heteroscedastic Gaussian process with predictive equations $m_N(\\mathbf{x}) = \\mathbf{c}(\\mathbf{x})^\\top (C_N + \\Lambda_N)^{-1} \\mathbf{y}$ and $s_N^2(\\mathbf{x}) = r(\\mathbf{x}) + \\sigma^2 ( c(\\mathbf{x},\\mathbf{x}) - \\mathbf{c}(\\mathbf{x})^\\top (C_N + \\Lambda_N)^{-1} \\mathbf{c}(\\mathbf{x}) )$, where $r(\\mathbf{x})$ is the input-dependent noise variance. The key structural identity is that when replication is available, these equations reduce to a form depending only on the $n$ unique designs: $m_n(\\mathbf{x}) = \\mathbf{c}(\\mathbf{x})^\\top(C_n + \\Lambda_n A_n^{-1})^{-1} \\bar{\\mathbf{y}}$, with $A_n$ the diagonal replication counts, so computational cost drops from $N$ to $n$. Around this core, the chapter organizes two inference families: latent log-variance GPs to learn $r(\\mathbf{x})$ when replicates are absent, and stochastic kriging to estimate $r(\\mathbf{x})$ from empirical variances when replicates are present. For sequential design, the central mechanism is the adaptation of criteria such as IMSPE and SUR to account for $r(\\mathbf{x})$ and replication.","core_discovery":"On its own terms, the chapter claims that the standard GP surrogate, including its closed-form predictive equations, remains the right backbone for stochastic simulators once the constant-noise assumption is relaxed to a variance function $r(\\mathbf{x})$ that changes with the input. It reviews three main routes to this relaxation: known noise from Monte Carlo error or tunable fidelity, latent log-variance GPs estimated by MCMC, EM, variational, or maximum-likelihood schemes, and replication-based empirical variance estimates in stochastic kriging. It further claims that non-Gaussian noise can be handled either by heavy-tailed likelihoods and approximate inference, by modeling conditional quantiles directly with a GP, or by deep (warped) GPs that add flexibility without explicit noise control. For sequential design, it argues that criteria built for deterministic surrogates can be adapted to the noisy setting, with replication itself becoming a design choice.","pith_inferences":["A natural codification of the chapter's trade-offs is a selection rule: latent-variable heteroscedastic GPs when replicates per site are scarce, stochastic kriging when replicates are abundant, and quantile GPs when tails or quantiles are the target; the chapter itself stops short of stating such a rule.","The replication-versus-exploration framing suggests that look-ahead IMSPE can be benchmarked against optimal computing budget allocation on discrete design sets, since both answer the same question of where to spend the next simulator calls.","The chapter's arguments imply that model choice should depend not only on data availability but on the smoothness of the noise variance function, a quantity that could be estimated from replicated data and used to forecast when heteroscedastic models will pay off."],"forward_implications":["With replicated observations, heteroscedastic GP inference can be carried out at a cost driven by the number of unique input locations rather than total simulator calls, enabling large replication budgets.","Because IMSPE depends only on predictive variance, sequential design can look ahead and decide between replicating an existing location and exploring a new one without needing observed outputs.","Quantile GPs and latent-variable models extend noisy surrogate modeling beyond Gaussian error, but require approximate inference and more data.","Plugging predictive means or quantile improvements into EI gives workable Bayesian optimization of the mean response under low signal-to-noise ratios.","For level-set estimation, contour SUR criteria can be adapted to noisy simulators and will place replicates on both sides of the crossing."],"supporting_citations":[{"why":"Introduces stochastic kriging, using replication-based empirical variances to model input-dependent noise and giving the predictive equations the chapter builds on.","marker":"Ankenman et al. (2010)"},{"why":"Introduces the latent log-variance Gaussian process that underpins heteroscedastic GP inference when noise variances are unobserved.","marker":"Goldberg et al. (1998)"},{"why":"Supplies practical maximum-likelihood inference for heteroscedastic GPs and shows that replicates reduce the predictive equations to unique designs.","marker":"Binois et al. (2018)"},{"why":"Formulates the replication-or-exploration sequential design question and provides look-ahead strategies the chapter recommends.","marker":"Binois et al. (2019)"},{"why":"Proposes quantile kriging, the basis for the chapter's treatment of non-Gaussian noise via empirical quantiles.","marker":"Plumlee and Tuo (2014)"},{"why":"Provides latent quantile GPs and the inference framework used in the chapter's quantile modeling illustration.","marker":"Picheny et al. (2022)"},{"why":"Compares kriging-based infill criteria under heterogeneous noise and underlies the chapter's sequential design recommendations.","marker":"Jalali et al. (2017)"},{"why":"Introduces the contour stepwise uncertainty reduction criterion used for level-set estimation in the illustration.","marker":"Lyu et al. (2021)"},{"why":"Establishes deep Gaussian processes and their variational inference, the basis for the chapter's deep GP discussion.","marker":"Damianou and Lawrence (2013)"}],"fun_headline_variants":["GP noise modeling meets stochastic simulation design","Heteroscedastic GPs: a guide to stochastic emulation","Replication-aware Gaussian processes for noisy simulators","From homoscedastic to input-varying noise in GP surrogates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The chapter's qualitative guidance is illustrated with a single one-dimensional SIR simulator; if that example is atypical of stochastic simulators, the suggested ordering of methods may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["GP noise modeling meets stochastic simulation design","Heteroscedastic GPs: a guide to stochastic emulation","Replication-aware Gaussian processes for noisy simulators","From homoscedastic to input-varying noise in GP surrogates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1350,"prompt_tokens":809,"completion_tokens":541,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":474}},"tokens_in":425,"tokens_out":541,"duration_ms":49914,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:55:56.249776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic benchmark on several stochastic simulators with different signal-to-noise ratios, replication costs, and non-Gaussian error structures, and compare quantitative mean-squared error and budget-accuracy curves; the chapter's visual illustrations alone do not settle whether the claimed ranking of methods holds.","supporting_citations":[{"cited_title":"A portfolio approach to massively parallel Bayesian optimization","cited_arxiv_id":"2110.09334","evidence_quote":"Introduces stochastic kriging, using replication-based empirical variances to model input-dependent noise and giving the predictive equations the chapter builds on."},{"cited_title":"W., Williams, C","cited_arxiv_id":null,"evidence_quote":"Introduces the latent log-variance Gaussian process that underpins heteroscedastic GP inference when noise variances are unobserved."}],"review_version":1}