{"id":"392e3311-c767-4d6d-aee4-daa49057954e","arxiv_id":"2506.15216","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A sleeping-expert framework, activated by gradient boosted trees, makes an online forecast aggregation more reactive without raising its average error, slightly lowering the 95th percentile of absolute temperature errors.","lead":"This paper combines a method for merging weather forecasts with a machine learning model that switches on forecasters normally too biased to use, but useful in extreme weather. The new method keeps the same average error while slightly reducing the largest errors, with a clear improvement during a cold spell in Chamonix.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GBRT activation skill (ESS 0.11) is too weak to support the Q95 gain; permutation test against random awakenings is needed.","rationale":"I considered the theoretical regret bound as a candidate for the most load-bearing concern. On inspection, the SEF regret definition (eq. 9) includes the indicator 1_{i_t∈E_t}, but because a sleeping expert is forced to predict the aggregation's prediction (eq. 2), its loss equals the aggregation's loss (eq. 5), making the regret contribution zero when the expert is asleep. Hence the equality (10) and the subsequent bound (16) hold for any comparator, and the paper's explicit assumption that 'Et contains the best expert' is an interpretational convenience rather than a mathematical necessity. The second-order bound (8) is a valid bound for the SEF with any activation sets. So the theoretical contribution is internally sound.\n\nThe real vulnerability is the empirical central claim. The paper's own headline improvement (Q95 2.52 vs 2.53°C) is a 0.01°C difference, and the mechanism advanced to explain it — GBRT-based activation — has near-zero skill (ESS 0.11, hit rates 0.15/0.12). The authors acknowledge that 'there is not much information detected in the features signal' and that the SEF sometimes degrades performance. The Chamonix event is a single case selected from many station–lead-time pairs. Without a statistical test against a random-activation control, the reported gain could be a chance artifact of the SEF's occasional use of biased experts rather than evidence that the GBRT 'knows when to wake up' the right experts. This is load-bearing because the paper's title, abstract, and conclusion all present the GBRT-driven activation as the novel mechanism enabling the improvement. If random activation performs equally well, the method reduces to a stochastic perturbation of BOA, with no demonstrated value from the learned component.\n\nThe recommended verdict is CONDITIONAL (unchanged from the reader): the authors should run the permutation test described above, and report error bars or a significance test for the Q95 difference. This is a concrete, low-cost check using the data already in hand.","tokens_in":25580,"tokens_out":16356,"duration_ms":157564,"concrete_test":"Permutation test: for each (station, lead time), take the GBRT-produced activation sequence (the binary indicators of when Q10/Q30 or Q70/Q90 are awakened) and randomly permute it across the 1253 iterations, preserving the exact total number of awakenings per expert and the initial 100 iterations without activation. Re-run BOAs with each permuted activation sequence and recompute the global Q95|e| and the Chamonix lead-48 Q95|e|. Repeat 1000 times to build a null distribution. If the observed global improvement (2.53−2.52 = 0.01°C) and the Chamonix improvement (3.53−3.06 = 0.47°C) lie outside the 95% central interval of the null distribution, the GBRT contributes genuine predictive skill. Otherwise, the central claim that learned activation reduces large errors is not supported, and the method should be presented as a heuristic with no demonstrated advantage over random awakening.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that SEF combined with GBRT reduces large errors (Q95|e|: 2.52°C vs 2.53°C) while maintaining RMSE, and that the mechanism is the GBRT's ability to wake specialized experts (Q10/Q30/Q70/Q90) when BOA is about to make an error beyond ±2.5°C. The paper's own diagnostic numbers undermine this mechanism: across all stations and lead times, ESS = 0.11 (near the random/climatology baseline), with hit rates of only 0.15 for positive large errors and 0.12 for negative ones. The GBRT thus misses roughly 85% of the large-error events it is supposed to detect, and most of its alarms are false. With such a weak signal, the observed global Q95 reduction of 0.01°C over ~740,000 pooled predictions is well within what chance selection across 594 station–lead-time pairs could produce. Figure 8 shows many pairs where BOAs is worse than BOA; the Chamonix lead-48 case (Q95 3.06 vs 3.53) is favorable but selected from these 594 pairs, and no multiple-comparison control is reported. The paper provides no test that separates the contribution of the learned GBRT activation from a random awakening schedule with the same base rate. If the two are indistinguishable, the method's stated advantage is not the GBRT's signal but merely the SEF's willingness to occasionally up-weight biased experts — a property that could be obtained without any learned trigger.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes combining the Sleeping Expert Framework (SEF) with the second-order expert aggregation BOA, using gradient boosted regression trees (GBRT) to decide when to wake specialized quantile experts (Q10, Q30, Q70, Q90) in order to reduce large temperature forecast errors while preserving average performance. The authors give a regret bound for BOA in the SEF with unknown activation sets, and present experiments over 33 stations and 18 lead times, reporting that BOAs maintains the RMSE of BOA (1.24°C) while lowering the pooled 95th percentile of absolute error from 2.53°C to 2.52°C, and that a meta-aggregation FTL-BOA further improves Q95|e| to 2.51°C. A case study in Chamonix (lead time 48h) shows a larger improvement during a cold-spell event (RMSE 1.54 vs 1.68°C, Q95 3.06 vs 3.53°C).","tokens_in":25939,"tokens_out":6113,"duration_ms":62243,"significance":"If the central empirical claims were statistically robust, the paper would offer a practical mechanism for making expert aggregation more reactive to extreme events without sacrificing average skill, which is a genuinely useful goal in temperature forecasting. The paper is careful to describe an entirely online procedure for the GBRT once hyperparameters are fixed, and it makes code and data available on GitHub, which are strengths. The theoretical part adapts existing second-order bounds to the SEF, but the stated bound is conditional on the activation set containing the best expert and is expressed in terms of the GBRT's confusion counts, so it is not a standard sublinear regret guarantee. The empirical evidence for the flagship improvement is thin: the pooled Q95 gain is 0.01°C, the GBRT's skill is weak (ESS 0.11, hit rates 0.15 and 0.12 for large positive and negative errors), and no confidence intervals or permutation controls are given. The significance is therefore currently more methodological and potential than demonstrated.","major_comments":[{"comment":"The central empirical claim, that BOAs reduces the pooled Q95|e| from 2.53°C to 2.52°C while keeping RMSE at 1.24°C, rests on a 0.01°C difference over roughly 740,000 pooled predictions, with no confidence intervals, significance test, or adjustment for the 594 station–lead-time pairs shown in Figures 8 and 10. The paper's own diagnostics—ESS 0.11, hit rates of 0.15 for ê ≥ 2.5°C and 0.12 for ê ≤ −2.5°C—indicate that the GBRT misses most of the large-error events it is designed to detect, so the observed global Q95 gain is well within what a random or fixed-rate awakening schedule with the same base rate could produce. The authors should report bootstrap confidence intervals for the Q95 differences and run a permutation test comparing the GBRT-based awakening rule against a randomized awakening schedule to show that the learned trigger contributes beyond chance.","section":"§3 (Explaining the predictions of the GBRTs); Figures 8 and 10"},{"comment":"The GBRT hyperparameters are tuned on training data from 2023-09-04 to 2025-02-03, which is strictly after the test period (2020-03-30 to 2023-09-03) on which all reported scores are computed. This lookahead contradicts the paper's claim that the procedure is 'fully online' and could bias the comparison in favor of BOAs, because the hyperparameter choice is optimized using future information relative to the test period. The authors should either re-run the experiments with a chronological split (tuning on a period before the test set) or demonstrate insensitivity of the qualitative results to the hyperparameter grid, for example by reporting the Q95 and RMSE for several near-optimal hyperparameter configurations.","section":"§3 (GBRT hyperparameters); Figure 2"},{"comment":"The regret bound is stated conditionally on the assumption that the activation set E_t contains the best expert i_t for every t, which is exactly the property that the GBRT is supposed to learn and is not verified in the experiments. The bound in Eq. (16) is expressed in terms of the confusion counts n^T_{b_k,k} of the GBRT, so it is an oracle bound that scales with how often the wake-up rule makes mistakes; when the GBRT is unskilled (as the reported ESS suggests), the bound is not a useful sublinear regret guarantee. The derivation should be made fully rigorous with proofs or precise references for each inequality, especially the step from Eq. (13) to Eq. (14), and the price paid for the unknown activation times should be stated explicitly rather than hidden by the assumption that the activation set contains the best expert.","section":"§2, Eqs. (8), (11)–(18)"},{"comment":"The Chamonix improvement (RMSE 1.54 vs 1.68°C, Q95 3.06 vs 3.53°C for lead time 48h) is presented as a highlight, but this station–lead-time pair is selected from the 594 pairs displayed in Figure 8, and no multiple-comparison correction or out-of-sample replication is provided. Without such a control, the case study is anecdotal; the authors should report the full distribution of per-pair Q95 improvements, state how many pairs show a statistically significant improvement after correction, and ideally validate the Chamonix finding on a hold-out period not used in any part of the analysis.","section":"§3 (A special case study); Figure 8"},{"comment":"The regularized FTL-BOA uses a regularization constant of 0.0025 in the comparison of cumulative losses, but the paper does not state how this value was chosen, whether it was tuned on the training period, or how sensitive the conclusions are to it. Additionally, the reported overall gains of FTL-BOA (Q95 2.51°C, RMSE 1.23°C) again lack uncertainty quantification; given the weak GBRT signal, the difference from BOA's 2.53/1.24 could easily be sampling noise. A sensitivity analysis for the regularization parameter and a significance test for the Q95 difference would be needed to support the claim that the meta-aggregation 'almost completely avoids noise.'","section":"§3 (FTL-BOA); Algorithm 2"}],"minor_comments":[{"comment":"There are several cross-reference mismatches in the text and captions (e.g., references to 'Fig. 9' and 'Fig. 10' that do not match the actual figure numbering); please proofread all figure references.","section":"Figure 8 and surrounding text"},{"comment":"The timing in Algorithm 2 is inconsistent: step 5 says the true outcome y_t is observed before step 6 chooses between BOA and BOAs, while step 7 then says the environment reveals y_t again; clarify the ordering of steps.","section":"Algorithm 2"},{"comment":"The regret definitions use both L^s_t and L^s_T interchangeably in places; please make the time index consistent throughout Section 2.","section":"Eq. (6) and notation"},{"comment":"The paper refers to 'Part I' without giving a citation or arXiv identifier for it; add a reference so the reader can locate the companion paper.","section":"Introduction and references"},{"comment":"The variable list contains duplicate entries ('Standard deviation of all experts' appears twice); remove the redundancy.","section":"Appendix: List of variables"},{"comment":"There are occasional typos and punctuation errors (e.g., inconsistent spacing around mathematical expressions and an apparent missing word in the discussion of the abstention trick); a careful copyedit is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The most serious issue is the lookahead in hyperparameter tuning: tuning GBRT hyperparameters on data from 2023–2025 and then evaluating on 2020–2023 is a methodological flaw that, if not corrected, undermines the 'fully online' claim and the credibility of the empirical comparison. The empirical effect is also very small (0.01°C Q95 gain) and the GBRT's low skill (ESS 0.11) means the mechanism story is not yet supported. The theoretical section needs to be tightened: the current bound is conditional on the very property being learned and is expressed in terms of the learner's confusion counts, so it does not provide a standard sublinear regret guarantee. If the authors can fix the temporal split, add significance/permutation tests, and clarify the theory, the paper could be suitable for publication; otherwise the contribution is too fragile. The self-citation pattern (Part I and Wintenberger) is not problematic in itself, but the novelty of the SEF second-order bound should be positioned more carefully against Adamskiy et al. (2012) and Gaillard et al. (2014)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on arXiv:2506.15216. The genuinely new ingredient is using sleeping experts with BOA where the awakening is learned online by a GBRT—no pre-specified activation calendar, unlike Devaine et al. The pipeline is carefully built: fully online, no lookahead, GBRT retrained each iteration, and a regularized FTL meta-aggregation to switch between BOA and BOAs. The Chamonix December 2021 case shows the mechanism can matter; the oracle experiments separating perfect awakening from perfect specialization are the right diagnostic. Credit where due: the paper is honest about the GBRT's weakness (ESS 0.11, hit rates 0.15/0.12) and about the difficulty of improving already good postprocessed forecasts.\n\nThe soft spots are in the empirical claim. The global Q95 improvement is 0.01°C (2.53 to 2.52) over roughly 740,000 predictions, with no confidence intervals and no test that separates the GBRT's awakening from a random awakening with the same base rate. Given the GBRT misses about 85% of large-error events and Figure 8 shows many station–lead-time pairs where BOAs is worse, the 0.01°C gain could easily be noise or selection across 594 pairs. The Chamonix improvement is real but is a post-hoc pick; there is no multiple-comparison control. The regret bound also conditions on the best expert being awake, an assumption the learned trigger does not guarantee, so the theory's reach is narrower than the abstract implies. It is a fair adaptation of Adamskiy, Gaillard, and Wintenberger; no circularity, but it is a conditional bound and should be labeled as such.\n\nStill, this is a serious applied paper. It tackles a real operational problem, is analytically honest, and publishes the negative diagnostics (SHAP, ESS, oracle gaps) that a good applied paper should show. The weak trigger is not authorial sloppiness; it is an empirical finding. What is missing is the statistical verification that the learned trigger beats random activation. Without that, the central claim 'fewer large errors' is not established.\n\nWho it is for: statisticians and post-processing researchers working on online aggregation or extreme-error reduction. It deserves a serious referee, and I would send it to review, but request a random-awakening control, error bars or a permutation test, and a clearer statement of the activation-set assumption. Expect major revision.","headline":"Online sleeping-expert BOA with a weak GBRT trigger: the 0.01°C Q95 gain is plausible but statistically under-supported.","tokens_in":26452,"tokens_out":2711,"would_cite":false,"duration_ms":26805,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T05","90C25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Sleeping experts can make temperature forecasts more reactive without hurting their average accuracy.","keywords":["expert aggregation","sleeping experts","temperature forecasting","online learning","regret bounds","gradient boosting","ensemble post-processing","extreme events"],"falsifier":"Run the identical online protocol with the wake-up rule inverted (wake Q10/Q30 when the tree predicts too-cold, Q70/Q90 when too-hot). If the reported Q95 improvement or the Chamonix cold-spell gain persists, the improvement is not caused by waking the sign-appropriate specialized experts; if the gain turns into a loss, the wake-up signal is what carries the result.","tokens_in":25393,"feed_emoji":"🌡️","tokens_out":10893,"duration_ms":99173,"temperature":0.7,"pith_summary":"The paper asks whether an online temperature forecaster that is already good on average can be made reactive enough to catch short-lived extreme events. Its answer is yes, if the ordinary experts are supplemented by specialized low- and high-quantile experts that sleep most of the time and are woken only when a trained tree model predicts the base aggregation will err by more than 2.5°C. Across 33 stations and 18 lead times, the sleeping-expert version of the BOA aggregation matches its RMSE of 1.24°C and lowers the 95th percentile of absolute error from 2.53°C to 2.52°C. For the difficult Chamonix station at 48 hours, a December 2021 cold spell is markedly better captured: RMSE improves from 1.68°C to 1.54°C and Q95 from 3.53°C to 3.06°C.","feed_headline":"Sleeping experts cut worst temperature errors while keeping RMSE","feed_subtitle":"Waking quantile forecasters only on predicted 2.5°C misses trims tail errors and catches a Chamonix cold spell.","key_machinery":"The carrying object is the Sleeping Expert Framework together with the abstention trick: a sleeping expert is forced to predict exactly what the awake aggregation predicts, so it receives the aggregation's loss and never influences today's forecast, after which the usual BOA update runs on these modified losses. On top sits a three-class wake-up rule in which a gradient-boosted regression tree predicts whether the base aggregation's error $\\hat e_t$ is at most $-2.5°C$, between $-2.5°C$ and $2.5°C$, or at least $2.5°C$, waking Q70/Q90, no specialized expert, or Q10/Q30 respectively. The regret identity $R^s_T(i^T)=\\sum_{t=1}^T(\\ell_t(w^s_t)-\\ell_t(\\delta_{i_t}))\\mathbf{1}_{i_t\\in E_t}$ then rewrites the bound as a sum over wake-up categories weighted by the counts of wake-up mistakes, making the quality of the wake-up predictor part of the theoretical guarantee.","core_discovery":"In the paper's own terms, the earlier BOA aggregation competes with the best fixed convex combination of experts under a second-order regret bound, but in practice it quickly drives the weights of biased quantile experts (Q10, Q30, Q70, Q90) to zero, so it cannot react to short cold spells or heat waves where only those quantiles are right. The discovery is that running BOA inside the Sleeping Expert Framework, with sleeping experts forced to copy the aggregation's prediction and a gradient-boosted regression tree deciding each day whether to wake the low quantiles, the high quantiles, or none, makes the aggregation sharply more reactive. The theoretical result is a second-order regret bound in the sleeping-expert setting whose slack is a sum over three error categories of how often the wake-up rule is wrong, so a perfect wake-up rule recovers the BOA bound and an imperfect one pays a controlled price. Empirically the improvement is concentrated in the tail: the global 95th percentile of absolute error falls slightly, the Chamonix December 2021 cold spell is predicted much better, and root mean squared error is unchanged.","pith_inferences":["Nothing in the abstention trick, the wake-up rule, or the regret bound is temperature-specific, so the same construction is a candidate for wind, precipitation, or any post-processed ensemble output once a threshold and a wake-up signal for extremes are defined.","The two oracle scores suggest the bottleneck is knowing when, not which expert to wake; a confidence-weighted awakening in which the specialized expert's weight is scaled by the predicted error magnitude is a natural next experiment.","Because the global gains are small and concentrated in a few difficult stations and lead times, an operational deployment question is whether the added online-training complexity pays for itself outside those cases; the regularized FTL-BOA already implements a cautious version of that choice.","A stronger external test would be an independent cold-spell or heat-wave season, or the larger operational station grid, since the average wake-up skill score of 0.11 indicates the method's value may be concentrated in exactly the places where the ensemble spread and first-lead-time observations carry a clear signal."],"forward_implications":["An expert aggregation that is already optimal on average can be made reactive to short extreme events without sacrificing RMSE: global Q95 of absolute error falls from 2.53°C to 2.52°C, and the Chamonix cold-spell Q95 falls from 3.53°C to 3.06°C.","A regularized follow-the-leader meta-aggregation over BOA and BOAs removes most of the noise the sleeping-expert framework adds when its wake-up signal is weak, giving RMSE 1.23°C and Q95 2.51°C overall.","Oracle experiments bound the remaining headroom: perfect wake-up predictions would give RMSE 1.08°C and Q95 2.12°C, while perfectly specialized experts would give 1.16°C and 2.38°C, so imperfect timing of wake-ups is the dominant source of residual tail error.","The regret bound shows that second-order aggregation with sleeping experts does not require knowing the awake set in advance; the theory pays for each wake-up mistake explicitly through the counts $n_{\\hat b_k,k}$.","The procedure is fully online: the tree model is retrained at every iteration, and at the tested scale it still completes a station–lead-time pair in less than half an hour."],"supporting_citations":[{"why":"Supplies the expert-aggregation framework and the Follow-The-Leader guarantee used for the FTL-BOA meta-aggregation.","marker":"Cesa-Bianchi and Lugosi [2006]"},{"why":"Introduced the sleeping/specialized experts framework on which the reactive construction is based.","marker":"Freund et al. [1997]"},{"why":"Gives the sleeping-expert regret definitions and the electricity-load application template, adapted here to unknown wake-up times.","marker":"Devaine et al. [2013]"},{"why":"Provides BOA, the second-order regret-bound aggregation that is extended to the sleeping-expert setting.","marker":"Wintenberger [2024]"},{"why":"Supplies the abstention trick and compound-expert regret formulation that let the bound be written in terms of wake-up mistakes.","marker":"Mourtada and Maillard [2017]"},{"why":"The generic reduction that yields second-order regret bounds in the sleeping-expert and confidence framework.","marker":"Gaillard et al. [2014]"},{"why":"The gradient-boosted regression tree implementation used to predict the sign and magnitude of the aggregation's error.","marker":"Chen and Guestrin [2016]"},{"why":"Describes the post-processed PEARP quantile experts that constitute the specialized low- and high-quantile forecasters.","marker":"Taillardat and Mestre [2020]"}],"fun_headline_variants":["Wake-up rule for sleeping experts sharpens cold-spell forecasts","Sleeping experts react faster to weather extremes without RMSE loss","Second-order bounds with sleeping experts cut forecast tail errors","Biased quantile experts wake on cue to trim large temperature errors","Gradient-boosted wake-up makes expert aggregation more reactive"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole gain rests on the gradient-boosted trees being able to detect, from the ensemble's current spread and the difference between first-lead-time observations and predictions, the rare moments when the base aggregation is about to miss by more than 2.5°C and with which sign; the paper's own skill scores (0.11) and hit rates (0.15 and 0.12) show that signal is weak on average.","fun_headline_variants_meta":{"raw":{"variants":["Wake-up rule for sleeping experts sharpens cold-spell forecasts","Sleeping experts react faster to weather extremes without RMSE loss","Second-order bounds with sleeping experts cut forecast tail errors","Biased quantile experts wake on cue to trim large temperature errors","Gradient-boosted wake-up makes expert aggregation more reactive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001048,"raw_usage":{"total_tokens":4447,"prompt_tokens":1031,"completion_tokens":3416,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":3331}},"tokens_in":647,"tokens_out":3416,"duration_ms":24416,"temperature":1.0,"reasoning_tokens":3331,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:40:12.841010+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical online protocol with the wake-up rule inverted (wake Q10/Q30 when the tree predicts too-cold, Q70/Q90 when too-hot). If the reported Q95 improvement or the Chamonix cold-spell gain persists, the improvement is not caused by waking the sign-appropriate specialized experts; if the gain turns into a loss, the wake-up signal is what carries the result.","supporting_citations":[{"cited_title":"Prediction, Learning , and Games","cited_arxiv_id":null,"evidence_quote":"Supplies the expert-aggregation framework and the Follow-The-Leader guarantee used for the FTL-BOA meta-aggregation."},{"cited_title":"Stochastic online convex optimization","cited_arxiv_id":null,"evidence_quote":"Provides BOA, the second-order regret-bound aggregation that is extended to the sleeping-expert setting."},{"cited_title":"Efficient tracking of a growing number of experts","cited_arxiv_id":"1708.09811","evidence_quote":"Supplies the abstention trick and compound-expert regret formulation that let the bound be written in terms of wake-up mistakes."},{"cited_title":"A Second -order Bound with Excess Losses","cited_arxiv_id":null,"evidence_quote":"The generic reduction that yields second-order regret bounds in the sleeping-expert and confidence framework."}],"review_version":2}