{"id":"c528b3a9-6917-4e1e-b1d2-f86193c9a8c4","arxiv_id":"2607.22310","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A temporal coherent forecast combination pools multiple experts' forecasts across aggregation levels, enforces coherence, and outperforms base and reconciled forecasts on German and Spanish day-ahead prices.","lead":"This paper combines forecasts from several models across hourly, block, and daily electricity price levels into a single forecast that automatically respects aggregation rules, and shows it beats every individual model and the best reconciled model. The method's practical value comes with clear guidance on which error-covariance structure to use.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness assumption in Proposition 1 is untested and likely violated by day-ahead price forecasts; without it the minimum-variance-unbiased optimality does not apply to the empirical setting.","rationale":"The reader's weakest assumption is the unbiasedness of base forecasts. This is indeed the least secure condition for the central theoretical claim: Proposition 1's 'minimum variance among linear unbiased combinations' is only meaningful if the base forecasts are unbiased. The paper explicitly assumes this (Section 2.2) but provides no diagnostic evidence, and electricity price forecasting is a setting where bias is plausible. If the assumption fails, the combined forecast is not guaranteed to be unbiased, so the advertised optimality property does not transfer to the empirical application. This does not necessarily undercut the empirical finding that the coherent combination improves MAE/RMSE, but it does weaken the interpretation of the method as the optimal unbiased combination. The paper's robustness analysis is otherwise solid: it compares many covariance estimators, leave-one-out expert pools, and sequential baselines, and reports significance tests. However, the unbiasedness gap is a genuine soft spot. A simple holdout bias test would settle whether it lands; if biases are negligible, the concern is moot. Since the reader already conditions the verdict on this and related limitations, the appropriate final verdict remains CONDITIONAL, i.e., unchanged.","tokens_in":28603,"tokens_out":10531,"duration_ms":95882,"concrete_test":"Using the replication data, compute the mean out-of-sample error of each expert at each temporal level (and of the shrbe combined forecast) and test H0: mean = 0 with a HAC t-test. If biases are small and insignificant, the concern is resolved. If significant, re-estimate the combination after adding a bias-correction term; if the bias-corrected combination does not materially change MAE/RMSE, the empirical claim remains, but the theoretical optimality must be restated as conditional.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The theoretical guarantee of Proposition 1 (Eqs. 12–13) is optimality among linear unbiased combinations, which requires E[bz_q] = z for every expert (Section 2.2). The paper never checks this. Day-ahead prices in 2021–2024 include the energy crisis and price spikes; forecasting models are typically biased in such regimes. If E[bz_q] ≠ z, then E[ez_c] = z + S·G·bias, so the combined forecast is not unbiased and the trace-minimization result does not certify minimum variance or minimum MSE. The empirical MAE gains may survive, but the central 'minimum error variance among linear unbiased combinations' claim is then not established in the actual setting. The paper also does not report bias of base or combined forecasts, so the reader cannot tell whether the assumption is approximately met.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a temporal coherent forecast combination procedure for day-ahead electricity prices. Given Q experts who each produce forecasts for all levels of a temporal hierarchy (hourly through daily aggregates), the method combines the forecasts into one vector that satisfies the temporal aggregation constraints and, under a stated unbiasedness assumption, has minimum trace error variance among linear unbiased combinations. Proposition 1 gives closed-form structural and zero-constrained expressions, with the single-expert temporal reconciliation of Athanasopoulos et al. (2017) as the Q=1 special case. The empirical study uses the four model classes from Lipiecki et al. (2026) for Germany and Spain over 2021–2024, comparing the combined forecast with base and reconciled forecasts in terms of MAE/RMSE, Diebold–Mariano tests, and MCB Nemenyi ranks. The authors report that the coherent combination outperforms the best reconciled expert at nearly every temporal level, and that the by-expert correlation structure is the key covariance-modelling choice.","tokens_in":28795,"tokens_out":7642,"duration_ms":69771,"significance":"If the results hold, the paper makes a useful practical contribution: it solves forecast combination and temporal reconciliation in one step, provides a closed-form solution that nests existing methodology, and gives a thorough empirical comparison on public base forecasts from four very different model classes. The empirical design is credible: the base forecasts are taken from an existing replication package, the evaluation is out-of-sample, and significance is assessed both pairwise and through multiple-comparison plots. The robustness analysis across covariance estimators, leave-one-out expert pools, and sequential alternatives is a real strength. However, the central theoretical guarantee — minimum error variance among linear unbiased combinations — relies on an unbiasedness assumption that is not tested in the empirical setting and is arguably questionable for day-ahead electricity price forecasts during 2021–2024. The empirical gains may survive the failure of this assumption, but the paper's main claim as stated is conditional on an unverified condition.","major_comments":[{"comment":"The derivation assumes E[bz_q] = z for every expert q. This assumption is load-bearing: it justifies the unbiasedness constraint GS_Q = I_m, and the minimum-variance claim in Proposition 1 is only among linear unbiased combinations. If the base forecasts are biased, then E[ez_c] = z + S G bias, so the combined forecast is not unbiased and Eqs. (12)–(13) do not certify minimum variance or minimum MSE. The test period 2021–2024 includes the European energy crisis and price spikes, where day-ahead price forecasts are plausibly biased. The paper does not report the bias of the base forecasts, of the reconciled forecasts, or of the combined forecast, and does not test or correct for bias. Please add an explicit bias diagnostic (e.g., mean in-sample and out-of-sample errors for each expert and for the combined forecast) and either show the assumption is approximately satisfied, or qualify the","section":"Section 2.2, before Eq. (10); Proposition 1, Eqs. (12)–(13)"}],"minor_comments":[{"comment":"The superscripts 1 and 2 are explained in the note, but in the table body they appear as digits attached to the numerical values (e.g., '19.31' vs '19.312'). This makes the table difficult to read; typeset them as proper superscripts.","section":"Table 2"},{"comment":"The references for the nonlinear shrinkage estimators are inconsistent: the text cites Ledoit and Wolf (2022a) for lis/gis and (2022b) for qis, but Table 4 labels all three with 'Ledoit and Wolf, 2022b'. Please correct the attribution.","section":"Section 2.4 / Table 4"},{"comment":"The paper motivates shrinkage by high dimensionality, but in the application nQ = 240 and the calibration window has roughly N ≈ 1095 days, so the sample covariance is invertible. State the concentration ratio explicitly and clarify whether the shrinkage is needed for numerical conditioning and estimation error rather than singularity.","section":"Section 2.4, Eq. (15)"},{"comment":"The claim of 'significant' improvements is based on pairwise DM tests at the 1% level without adjustment for the number of comparisons, although the MCB Nemenyi procedure partially addresses this. It would be helpful to report whether the DM conclusions survive any standard multiple-testing correction, or to rely primarily on the MCB results.","section":"Section 3.2 / Figure 3"},{"comment":"In Germany, leaving out XGB leads to a small but systematic improvement over the full pool. This is an interesting qualification to the general message that combining more experts is beneficial; the text mentions it but could connect it more explicitly to the trade-off between signal and estimation noise.","section":"Section 4.3, Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a straightforward temporal specialization of the authors' earlier cross-sectional coherent forecast combination framework, so the theoretical novelty is modest. Its main value is the clean empirical demonstration on electricity prices, with public base forecasts and careful robustness analysis. The single largest risk is the untested unbiasedness assumption; if the authors can show that bias is negligible, or add an intercept/bias correction, the paper could be published. I did not find evidence of citation manipulation; the heavy self-citation is to directly relevant prior work and is acknowledged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper does something real: it gives a closed-form solution for combining Q forecasters across a temporal hierarchy in one step, and Proposition 1 nests the standard single-expert temporal reconciliation as the Q=1 special case. The derivation in Appendix A is self-contained and correct. The empirical work is also solid — the combination uses out-of-sample base forecasts from a public replication package, and the evaluation is genuinely ex ante. Hourly MAE gains of 4-24% over base forecasts, with the largest gains at coarser levels, and the leave-one-out analysis shows the result isn't carried by one expert.\n\nIt does several things well. The covariance estimator comparison is thorough and gives practical guidance: preserving each expert's error dependence across temporal levels matters more than the choice of shrinkage. The simultaneous combination-plus-reconciliation beats the sequential average-then-reconcile, which is a useful finding.\n\nThe soft spots are real but not fatal. The main one, raised by the stress-test, lands: Proposition 1 assumes the base forecasts are unbiased, and the paper never checks that. Over 2021-2024, with the energy crisis and price spikes, day-ahead forecasts are often biased. If so, the combined forecast isn't unbiased, and the minimum-error-variance-among-unbiased-combinations guarantee doesn't hold in the empirical setting. The MAE gains may persist, and the method may still be a reasonable pragmatic choice, but the theoretical claim is overstated. A bias-correction extension or even a reported bias table would help.\n\nA smaller issue: the reference shr_be estimator is chosen after seeing the test-period results. The paper is transparent and reports all estimators in Table 4, so it's not hidden, but the headline numbers use the best performer. Mild selection, not a fatal flaw.\n\nOverall: the central idea is new, the math is correct, and the empirical work is reproducible. I'd send it to peer review, and I'd cite it, but I'd want the unbiasedness issue addressed or at least clearly acknowledged before publication.","headline":"A real extension of coherent forecast combination to temporal hierarchies with solid empirical support, but the unbiasedness assumption behind the optimality claim is untested and likely violated in electricity prices.","tokens_in":29249,"tokens_out":3614,"would_cite":true,"duration_ms":31015,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a one-step temporal coherent combination that pools experts across all levels of an electricity price hierarchy, enforces aggregation constraints, and minimizes error variance; empirically it beats the best reconciled exp","keywords":["forecast combination","temporal hierarchies","forecast reconciliation","coherent forecasts","electricity price forecasting","day-ahead market","covariance shrinkage","temporal aggregation"],"falsifier":"In a period with known biased forecasts, such as price-spike episodes, compute the coherent combination with and without an explicit bias correction and compare hourly MAE; if the bias-corrected version is more accurate, the unbiasedness assumption, not the combination algorithm, is carrying the guarantee. A complementary simulation: generate Q experts with exactly known covariance but nonzero means and check that the combined forecast inherits the bias.","tokens_in":28489,"feed_emoji":"⚡","tokens_out":8747,"duration_ms":66781,"temperature":0.7,"pith_summary":"Day-ahead electricity prices are forecast by many competing models at many granularities, from single hours to daily baseloads, and these forecasts are typically incoherent—the average of the predicted hours does not match the predicted baseload—while no single model is the most accurate everywhere. The paper tries to solve both problems in one step: it combines the forecasts of Q experts across the whole temporal hierarchy into a single forecast that satisfies the aggregation constraints and has minimum error variance among all unbiased linear combinations. The solution is a closed-form formula driven by the covariance of the experts' stacked forecast errors. Tested on four very different model classes and two day-ahead markets over four years, the combined forecast significantly outperforms both the individual base forecasts and the best single-expert reconciled forecast at almost every temporal level, with the largest gains at the coarser aggregations used for block and baseload trading. If the result holds, forecasters no longer need to pick a single model or reconcile and combine in separate steps.","feed_headline":"Pool four experts and beat the best single model at nearly every level","feed_subtitle":"It reconciles and pools four model forecasts in one step, cutting hourly error by up to a quarter.","key_machinery":"The temporal hierarchy and the coherent-combination projector built on it. The hierarchy organizes the 24 hourly prices into non-overlapping averages at coarser levels—2-, 3-, 4-, 6-, 8-, and 12-hour blocks and the daily baseload—through a fixed aggregation matrix A, summarized by the structural matrix S=[A; I]. Proposition 1's projector, e z_c = S(S_QᵀW⁻¹S_Q)⁻¹S_QᵀW⁻¹ bz (equation 12), maps the stacked base forecasts onto the coherent subspace, weighting by the inverse covariance W⁻¹ of the stacked forecast errors. This single formula carries the argument: it chooses combination weights across experts and reconciles across temporal levels at the same time, and it collapses to single-expert","core_discovery":"The paper's central claim is that when Q experts each forecast every level of a temporal hierarchy, one linear combination of their stacked forecasts exists that is coherent (each aggregated level equals the corresponding average of the hourly forecasts) and has minimum error variance among all unbiased linear combinations. Proposition 1 gives this combination in closed form, in structural and zero-constrained representations (equations 12 and 13), and the single-expert case Q=1 reduces to standard temporal reconciliation. The formula depends on the covariance matrix of the stacked base forecast errors, so the paper evaluates estimators that cross four correlation structures with linear and","pith_inferences":["Inference: the method should transfer to other temporal hierarchies whose aggregates are averages or sums—electricity load, renewable generation, retail sales, or economic flows—and the testable prediction is that the one-step combination keeps beating each reconciled expert there.","Inference: since the optimality proof assumes unbiased base forecasts, markets with frequent price spikes or regime shifts may need a bias-correction step; a natural experiment would compare the coherent combination with and without bias correction in such periods.","Inference: the by-expert covariance finding suggests an operational rule—estimate each model's error dependence across aggregation levels and ignore cross-model error correlations—which, if confirmed on other datasets, greatly reduces the covariance estimation burden.","Inference: the closed-form projector makes a probabilistic extension natural: rather than combining point forecasts, one could combine predictive distributions and project them onto the coherent subspace, yielding coherent scenarios useful for risk management."],"forward_implications":["Forecasters can pool any number of models into one coherent forecast without a separate model-selection step; in the two markets studied, the pooled forecast beats the best reconciled single expert at nearly every temporal level.","The largest gains appear at coarser aggregations, directly improving the block and baseload products that are traded alongside hourly prices.","Combining and reconciling in one simultaneous step beats the sequential alternative of averaging the experts first and then reconciling the average.","Estimating the covariance structure matters more than the shrinkage technique: preserving each expert's within-expert temporal error dependence is the main driver, and sophisticated nonlinear shrinkage adds little over linear shrinkage.","Removing any one expert, even the strongest, does not destroy the gains; the reduced pool remains comparable to or better than the best reconciled expert in most cases, though a redundant expert can add estimation noise."],"fun_headline_variants":["Reconcile and pool forecasts to beat every expert","Beat every expert with one coherent forecast combination","Pool all experts' forecasts coherently in one formula","Stack forecasts, reconcile, and win every level"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The optimality guarantee assumes every expert's base forecasts are unbiased—that the expected forecast equals the true price—so if the individual forecasts are systematically biased, the minimum-variance-unbiased property of the combined forecast is no longer guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Reconcile and pool forecasts to beat every expert","Beat every expert with one coherent forecast combination","Pool all experts' forecasts coherently in one formula","Stack forecasts, reconcile, and win every level"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00133,"raw_usage":{"total_tokens":5215,"prompt_tokens":678,"completion_tokens":4537,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":4477}},"tokens_in":422,"tokens_out":4537,"duration_ms":25019,"temperature":1.0,"reasoning_tokens":4477,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:07:58.713799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a period with known biased forecasts, such as price-spike episodes, compute the coherent combination with and without an explicit bias correction and compare hourly MAE; if the bias-corrected version is more accurate, the unbiasedness assumption, not the combination algorithm, is carrying the guarantee. A complementary simulation: generate Q experts with exactly known covariance but nonzero means and check that the combined forecast inherits the bias.","supporting_citations":[],"review_version":1}