{"id":"2625be62-d1f0-4774-902c-c03ac8720bd2","arxiv_id":"2509.04203","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Gibbs posterior over linear pool weights, built by exponentiating a CRPS-based risk, gives forecast ensemble weights with uncertainty estimates and often improves on BMA, AVS, and equal weighting.","lead":"This paper introduces a new way to combine probabilistic forecasts: a Gibbs posterior over ensemble weights, optimized on the continuous ranked probability score, with a prior that can pull weights toward equality. Tests in simulations and the 2023-24 CDC FluSight competition suggest it often beats Bayesian model averaging, adaptive variable selection, and equal weighting.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FluSight real-data claim rests on an unvalidated quantile-to-distribution reconstruction; if the qGP conversion distorts the submitted forecasts, the SGP weights and CRPS rankings in Table 1 are not measuring the actual ensemble forecasts.","rationale":"The reader's weakest assumption is the same as the one I find most load-bearing: the FluSight analysis is the only real-data demonstration, and it is built on a two-stage procedure in which the submitted quantiles are first converted to continuous distributions by a qGP model and then stacked by CRPS. The paper provides no evidence that the qGP predictive distribution preserves the information in the original quantiles. This is not an external or consensus-based objection; it is an internal validity issue: the metric being optimized (CRPS of reconstructed draws) may not equal the metric that defines the forecasts. If the conversion is biased, the SGP weights, the credible intervals in Fig. 7, and the Table 1 rankings are all measuring an approximation. I considered the alternative concern that SGP is never compared to non-Bayesian point-optimal linear stacking, and the learning-rate/prior choices are ad hoc. Those are real limitations, but they affect the interpretation and scope of the contribution rather than the truth of the specific comparative claims against AVS/BMA/EQW. The theory in Theorem 1 is for iid fixed components, and the authors explicitly acknowledge that dynamic settings are not covered, so I do not rest the objection on the missing dynamic theory. The simulations are internally coherent, though the SGP's advantage in the SIR study comes mainly from the SGP50 variant, which is a post-hoc regularization choice. The proposed concrete test is feasible with the public FluSight data: recompute CRPS directly from quantiles and see whether Table 1's rankings survive. If they do, my concern is answered and the conditional accept is justified; if they do not, the real-data claim needs substantial qualification.","tokens_in":19618,"tokens_out":10149,"duration_ms":109353,"concrete_test":"Recompute the SGP weight-selection pipeline of Section 5 on the same FluSight data, but compute the CRPS of each component forecast and ensemble forecast directly from the submitted 23 quantiles using a piecewise-linear/trapezoidal approximation to equation (12) rather than from the qGP posterior predictive draws. Then re-derive the regional and weekly CRPS rankings in Table 1. If the rankings shift materially, or if the SGP no longer beats BMA/EQW in a majority of regions, the real-data claim is an artifact of the qGP reconstruction. As a confirmatory check, compare the PIT of observed hospitalizations under the qGP draws with the PIT under the original quantile forecasts; large systematic deviations indicate the conversion is not faithful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central real-data claim (Section 5, Fig. 7, Table 1) is that the SGP ensemble outperforms AVS/BMA/EQW on the 2023-24 FluSight forecasts. But the SGP is not fit to the submitted quantile forecasts: it is fit to 50,000 posterior predictive draws from a quantile Gaussian process (qGP) matching model (Wadsworth and Niemi, 2025b), and the CRPS values are Monte Carlo estimates of expectations under that reconstructed distribution. Every conclusion about FluSight therefore presupposes that the qGP predictive distribution faithfully represents each team's submitted quantile forecast. The paper provides no fidelity check for this conversion: no PIT comparison between the qGP draws and the original quantile forecasts, and no comparison of the CRPS computed from qGP draws to a CRPS/WIS computed directly from the 23 submitted quantiles. If the qGP is misspecified, over- or underdispersed, or dominated by prior uncertainty, then the 'true' forecasts are distorted; the SGP weights are optimal for the qGP approximation rather than for the forecasts as submitted, and the regional/weekly CRPS rankings in Table 1 measure the approximation, not the actual FluSight forecasts. This is load-bearing because without the FluSight analysis the empirical claim of outperformance rests only on simulations. The missing stacking baseline and the ad hoc eta/prior choices are important but secondary.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 'stacked Gibbs posterior' (SGP) for combining probabilistic forecasts in a linear pool. The SGP is defined in Eq. (4) as a Gibbs posterior over the simplex of combination weights, with empirical risk based on the CRPS and a prior distribution on the weights. The authors give a proposition expressing the CRPS of a mixture of continuous forecast distributions as a quadratic form in the weights, a consistency theorem for i.i.d. data and fixed component models, three simulation studies (i.i.d. mixture, dynamic mixture, and SIR epidemic forecasts), and an application to the 2023-24 CDC FluSight hospitalization forecasts, comparing SGP with adaptive variable selection (AVS), Bayesian model averaging (BMA), and equal weighting (EQW). The abstract claims that SGP 'often outperforms' the other methods in both simulations and the FluSight analysis.","tokens_in":20031,"tokens_out":8177,"duration_ms":80639,"significance":"If the empirical claims hold, the SGP is a useful addition to the forecast combination toolkit: it gives a principled Bayesian-style posterior over ensemble weights, permits regularization toward equal weighting through the prior, and provides uncertainty quantification for the weights. The consistency theorem and the continuous-mixture CRPS formula (Prop. 1) are useful, and the paper is honest about the scope of the theory, explicitly limiting Theorem 1 to the i.i.d., fixed-model setting. The main value is the clear construction and the Monte Carlo CRPS evaluation for mixtures. However, the central real-data evidence is currently weakened by the unvalidated quantile-to-distribution reconstruction used in the FluSight analysis, and some simulation claims are stronger than the plots support. The authors do not provide code or data, and they do not compare against the most closely related stacking method (Yao et al., 2018).","major_comments":[{"comment":"The SGP weights and all CRPS scores in the FluSight analysis are computed from 50,000 posterior predictive draws of the quantile Gaussian process (qGP) matching model, not directly from the quantile forecasts submitted by the teams. The manuscript provides no validation that the qGP predictive distribution faithfully represents each team's submitted quantile forecast. If the qGP is misspecified, over- or underdispersed, or dominated by prior uncertainty, then the CRPS rankings in Table 1 and the SGP weights measure the reconstructed distributions, not the actual FluSight forecasts. Because this analysis is the primary real-data evidence for the abstract's outperformance claim, please add a fidelity check (e.g., PIT of the qGP draws against the original quantiles, or a comparison of CRPS computed from qGP draws with WIS computed directly from the 23 submitted quantiles) or substantially s","section":"Section 5, Fig. 7, Table 1"},{"comment":"The proof of Theorem 1 contains a garbled line: 'Now Ω\\B_δ ⊂ Ω\\B_δ′, so G(Ω\\B_δ′)⊂ G(Ω\\B_δ′)' — the set on the left should be G(Ω\\B_δ). The correction is straightforward: Ω\\B_δ ⊂ Ω\\B_δ′ implies G(Ω\\B_δ) ⊂ G(Ω\\B_δ′), hence inf_{Ω\\B_δ′} G ≤ inf_{Ω\\B_δ} G; since ω* ∉ Ω\\B_δ′ and ω* is the unique minimizer, G(ω*) < inf_{Ω\\B_δ′} G, giving (20). In addition, the theorem statement should include a finite first-moment condition on the component models (E|X_c| < ∞ for each c), because the domination bound in the proof uses m(y) = E|X| + |y|; the present statement only assumes E|Y| < ∞. Without this, the hypotheses do not guarantee that the CRPS of the mixture is finite.","section":"Section 8.2, proof of Theorem 1"},{"comment":"The statement that 'The SGP largely outperformed AVS, BMA, and EQW ensembles' is not supported by the SIR simulation for the default SGP. Figure 6 shows that the plain SGP does not perform as well as the competitors, and the competitive variant SGP50 requires a specially strengthened Dirichlet prior (the text appears to say the prior parameter vector is '504', presumably 50/4) that heavily regularizes toward equal weights. The paper should separate the two claims: a default SGP is not uniformly better; a strongly regularized SGP is. This matters because the FluSight analysis uses the default uninformative Dirichlet prior, so the real-data outperformance of the default method is not explained by this simulation.","section":"Section 4.3 and Section 4 summary"}],"minor_comments":[{"comment":"Editorial notes remain in the text (e.g., 'Jarad: Include lower is better' near Fig. 4), and there are typos such as 'foreacsts' in the abstract, 'Similary' in §8.1, 'peprformed' near Fig. 7, and 'a alternative' in §3.1.","section":"Throughout"},{"comment":"The text says 'the empirical risk function from (5) is used'; this should be Eq. (9). Eq. (5) defines the risk minimizer, not the empirical risk.","section":"Section 4.1"},{"comment":"The prior parameter values are written as '14' and '504'; these presumably mean 1/4 and 50/4. Please write them unambiguously, e.g., Dir(1/4, ..., 1/4) and Dir(50/4, ..., 50/4).","section":"Section 4.3"},{"comment":"The text states that the left panel shows 90% credible intervals for Rhode Island, while the caption says the figure is for the national level; please clarify which is displayed. Also, define the number of included teams/models in the text (the caption mentions 10) and state the number of regions explicitly.","section":"Section 5, Fig. 7"},{"comment":"The Monte Carlo mean formula in the text should be \\bar w_{c,t} = M^{-1} \\sum_{m=1}^M w^{(m)}_{c,t}; as written it is missing the summation over posterior draws.","section":"Section 5"},{"comment":"No comparison is made to the original stacking method of Yao et al. (2018) or to a point-estimate CRPS weight optimizer, despite the paper's emphasis on stacking. Adding such a baseline would help calibrate the 'often outperform' claim and is important for positioning the contribution.","section":"Empirical comparisons"}],"recommendation":"major_revision","confidential_remarks":"The key issue is the unvalidated qGP reconstruction in the FluSight analysis; this should be addressed with a concrete fidelity check before the real-data claims can be accepted. The methodological core is sound and the paper is within scope for stat.ME, but the current real-data evidence is conditional on a model that is not validated in the manuscript. No concerns about citation patterns or novelty disclosure beyond the missing comparison to existing stacking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's core idea is sound and genuinely new: put a Gibbs posterior over the simplex of linear-pool weights with a CRPS-based risk, which gives you both optimal weights and posterior uncertainty, plus a prior that can regularize toward equal weighting. Proposition 1 (CRPS of a continuous mixture as a weighted sum of expected absolute differences) is simple but useful, and the consistency theorem, while an application of Martin and Syring (2022), is correctly specialized.\n\nThe simulations are fairly thorough: i.i.d., dynamic, and SIR outbreak data, with comparisons to BMA, AVS, and equal weighting. The authors are honest that the theory only covers the i.i.d. fixed-model case, and they don't oversell the dynamic results. The FluSight application is also a nice real-data stress test.\n\nThe soft spots are real but addressable. First, there is no comparison to the original stacking method (Yao et al. 2018), which the paper explicitly extends. That is a notable omission, since the whole point is to show the Bayesian version is competitive with or better than point-estimate stacking.\n\nSecond, and more serious, the FluSight analysis does not evaluate the submitted quantile forecasts directly. The SGP is fit to 50,000 posterior predictive draws from a quantile Gaussian process reconstruction, and the CRPS values in Table 1 are Monte Carlo estimates under that reconstructed distribution. There is no fidelity check--no PIT comparison between qGP draws and the original quantiles, and no direct CRPS/WIS comparison on the submitted quantiles. If the reconstruction is misspecified, the reported rankings measure the approximation, not the actual FluSight forecasts. This is load-bearing, because without FluSight the empirical claim rests on simulations where the method does not always win (e.g., uninformative SGP in the SIR study).\n\nMinor issues: the learning rate eta is chosen ad hoc (15 for i.i.d., 1 elsewhere), and the SGP50 prior in the SIR study is tuned to performance; sensitivity is not explored. A garbled line in the theorem proof (section 8.2) suggests a proofreading error. No code is provided.\n\nOverall, this is a useful paper for people working on forecast combination. It deserves peer review; the revisions should add the stacking baseline, validate the qGP conversion, discuss eta/prior sensitivity, and release code.","headline":"New and useful Gibbs-posterior stacking method for linear-pool weights; FluSight evidence rests on an unvalidated quantile reconstruction.","tokens_in":20447,"tokens_out":3139,"would_cite":true,"duration_ms":30610,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62F12","62M20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces the stacked Gibbs posterior, a Bayesian distribution over linear-pool weights that optimizes a proper scoring rule, and reports it often beats model averaging and equal weighting in simulations and in the 2023-24 FluSig","keywords":["stacked Gibbs posterior","optimal linear pooling","probabilistic forecasting","proper scoring rules","continuous ranked probability score","ensemble forecasting","Bayesian model averaging","forecast combination"],"falsifier":"Take the FluSight data and re-estimate the ensemble weights twice: once using the quantile Gaussian process draws (as in the paper) and once computing CRPS directly from a piecewise-linear CDF on the submitted quantile grid, then compare the SGP, AVS, BMA, and EQW rankings. If the SGP no longer ranks first in 31 of 53 regions, the reported advantage depends on the reconstruction. Separately, in the i.i.d. simulation, check whether the SGP posterior mean tracks the exact CRPS-minimizing weights as n grows; if not, the consistency result is not operative.","tokens_in":19555,"feed_emoji":"📊","tokens_out":13572,"duration_ms":123529,"temperature":0.7,"pith_summary":"The paper's central claim is that a Gibbs posterior over linear-pool weights—probability mass proportional to exp(−η n times the average CRPS) times a Dirichlet prior—is a practical way to do stacking for probabilistic forecasts. Unlike point-optimized stacking or model-averaging posteriors, this approach returns a full distribution over weights, so each component's influence carries uncertainty, and the prior can pull the ensemble toward equal weighting. The authors report that in three simulation studies the stacked Gibbs posterior (SGP) frequently beats Bayesian model averaging, adaptive variable selection, and equal weighting on proper-score and calibration measures, and in the 2023-24 FluSight flu-hospitalization forecasts it ranked first in more regions and weeks than the other methods. The appeal, if the claim holds, is that forecast hubs get a principled way to optimize the exact scoring rule they are judged by without losing the safety of equal-weight pooling.","feed_headline":"31 of 53 flu regions won by Gibbs posterior weights","feed_subtitle":"A Bayesian stacking method that optimizes CRPS beats model averaging and equal-weight pools in simulations and flu forecasts.","key_machinery":"The central object is the stacked Gibbs posterior (SGP): a distribution over the weight simplex, π_n(η)(ω) ∝ exp{−η n S_n(ω)} π(ω), where S_n is the average CRPS of the mixture forecast over past outcomes (optionally discounted) and π(ω) is a Dirichlet prior. The prior's full-simplex support satisfies the consistency condition and lets the user regularize toward equal weights by increasing concentration. The load-bearing identity is Proposition 1, CRPS(P̄, y) = Σ w_c E|X_c−y| − ½ ΣΣ w_c w_d E|X_c−X_d|, which turns CRPS evaluation of a mixture into expected absolute differences between components and observations, estimable from draws. The Gibbs-posterior formulation converts weight selection","core_discovery":"The core discovery is the stacked Gibbs posterior (SGP), defined as π_n(η)(ω) ∝ exp{−η n S_n(ω)} π(ω), with S_n the empirical CRPS of the linear-pool forecast (optionally discounted for dynamic settings) and π(ω) a Dirichlet prior. Theorem 1 shows that for i.i.d. data and fixed component models the posterior concentrates on the weight vector minimizing expected CRPS. Proposition 1 extends the known normal-mixture CRPS formula to arbitrary continuous mixture components: CRPS(P̄, y) = Σ_c w_c E|X_c−y| − ½ Σ_c Σ_d w_c w_d E|X_c−X_d|, which is what makes the method usable when components are Monte Carlo draws. In the paper's analysis, SGP ensembles score well on CRPS/LogS and PIT uniformity; in","pith_inferences":["A same-family extension is to replace CRPS with the weighted interval score (WIS) that FluSight actually uses; because WIS approaches CRPS as the interval set grows, an SGP-like posterior over WIS would directly optimize the competition's headline metric.","The fixed Dirichlet prior makes the posterior weight distribution static; a state-space prior on ω would let SGP track regime changes, a direction the paper itself notes is not covered by its i.i.d. theory.","The FluSight analysis keeps only teams with complete weekly submissions, so SGP weights are estimated on a survivor set; extending the Gibbs posterior to missing forecasts is a natural way to recover value from dropped teams.","Treating the Dirichlet concentration (1 vs 50 in the paper's SGP50) as a tuned shrinkage parameter selected by rolling-origin cross-validation is a natural extension consistent with the paper's learning-rate tuning philosophy."],"forward_implications":["Forecast hubs gain a way to publish posterior intervals for ensemble weights, making the combination method auditable rather than a single point estimate.","A stronger Dirichlet prior centered at equal weights becomes a built-in shrinkage knob; in the SIR simulation the regularized SGP50 variant produced the lowest LogS and CRPS among the compared methods.","Proposition 1 lets SGP accept component forecasts as samples, not only closed-form CDFs, so hubs that receive quantiles or posterior draws can still build CRPS-optimal linear pools.","Theorem 1 gives static settings a guarantee: with i.i.d. data and fixed models, the posterior provably concentrates on the CRPS-minimizing weights as n grows, so the method does not rely on ad hoc asymptotics in that regime.","In the 2023-24 FluSight reevaluation, SGP placed first by mean CRPS in 31 of 53 regions and 14 of 29 weeks, consistent with the claim that it outperforms the equal-weight default often enough to matter in practice."],"supporting_citations":[{"why":"Supplies the Gibbs posterior framework, consistency conditions, and learning-rate guidance that the SGP builds on.","marker":"Martin and Syring (2022)"},{"why":"Defines proper scoring rules and the CRPS representation used for the empirical risk and Proposition 1.","marker":"Gneiting and Raftery (2007)"},{"why":"Establishes stacking of Bayesian predictive distributions in the M-open setting; SGP is the Bayesian extension with a prior over weights.","marker":"Yao et al. (2018)"},{"why":"Introduces optimal prediction pools, the basis for optimizing linear pool weights by proper scoring rules.","marker":"Geweke and Amisano (2011)"},{"why":"Provides the normal-mixture CRPS formula that Proposition 1 generalizes to arbitrary continuous components.","marker":"Li et al. (2019)"},{"why":"AVS is a main comparison method and supplies the discount factor α=0.98 used in the dynamic risk function.","marker":"Lavine et al. (2021)"},{"why":"Quantile Gaussian process matching model converts FluSight quantile forecasts into draws so CRPS can be approximated.","marker":"Wadsworth and Niemi (2025b)"},{"why":"Uniform law of large numbers lemma used in the proof of Theorem 1's consistency.","marker":"Newey and McFadden (1994)"}],"fun_headline_variants":["Gibbs posterior stacking wins most flu regions","Bayesian stacking beats equal-weight pools","Stacked Gibbs posterior optimizes CRPS scores","Gibbs posterior gives weighted edge in flu forecasts"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The FluSight result, the primary real-data evidence, assumes the quantile Gaussian process matching model converts submitted quantile forecasts into continuous distributions accurately enough that Monte Carlo CRPS values computed from its posterior draws faithfully represent the original submissions; if that conversion is inaccurate, the reported weight estimates and CRPS rankings are not measuring what they claim.","fun_headline_variants_meta":{"raw":{"variants":["Gibbs posterior stacking wins most flu regions","Bayesian stacking beats equal-weight pools","Stacked Gibbs posterior optimizes CRPS scores","Gibbs posterior gives weighted edge in flu forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3096,"prompt_tokens":820,"completion_tokens":2276,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":2220}},"tokens_in":564,"tokens_out":2276,"duration_ms":16181,"temperature":1.0,"reasoning_tokens":2220,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:18:20.089604+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the FluSight data and re-estimate the ensemble weights twice: once using the quantile Gaussian process draws (as in the paper) and once computing CRPS directly from a piecewise-linear CDF on the submitted quantile grid, then compare the SGP, AVS, BMA, and EQW rankings. If the SGP no longer ranks first in 31 of 53 regions, the reported advantage depends on the reconstruction. Separately, in the i.i.d. simulation, check whether the SGP posterior mean tracks the exact CRPS-minimizing weights as n grows; if not, the consistency result is not operative.","supporting_citations":[{"cited_title":"and Syring, N","cited_arxiv_id":null,"evidence_quote":"Supplies the Gibbs posterior framework, consistency conditions, and learning-rate guidance that the SGP builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes stacking of Bayesian predictive distributions in the M-open setting; SGP is the Bayesian extension with a prior over weights."},{"cited_title":"and Amisano, G","cited_arxiv_id":null,"evidence_quote":"Introduces optimal prediction pools, the basis for optimizing linear pool weights by proper scoring rules."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"AVS is a main comparison method and supplies the discount factor α=0.98 used in the dynamic risk function."}],"review_version":1}