{"id":"b4054698-3a3e-4c3c-af4b-45fe5046a1a2","arxiv_id":"2411.17743","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For a battery storage trading strategy in the German day-ahead market, ranking probabilistic forecasting models by empirical coverage at the trading hours yields higher average per-trade profits than ranking by pinball loss.","lead":"This paper compares six ways of scoring probabilistic electricity price forecasts and asks which scoring rule picks the forecasting model that earns the most in a battery-trading simulation. The winner is a coverage score measured only at the two hours the trading strategy actually uses: charge when the price is near the forecast low, sell when it is near the forecast high.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The rolling-window length for model selection is given as 30 days in §4.1 but 182 days in §4.2; this unresolved contradiction directly undermines the reported 'clear winner' and must be settled before the central claim can be evaluated.","rationale":"The reader identified the 30-vs-182-day inconsistency in the rationale, but chose the 'single out-of-sample path without significance tests' as the weakest assumption. I regard the undisclosed and contradictory window length as the single most load-bearing concern because it directly specifies how the model-selection metric is computed. The headline result is a comparison of metrics; if the rolling window length is misreported, the comparison itself rests on an unspecified procedure. This is a concrete, checkable flaw that can be resolved by running the experiment with both window lengths. The lack of confidence intervals is also a serious problem, but it is secondary: even with significance tests, the result would still be uninterpretable if the core estimation procedure is not clearly defined. Since the reader's CONDITIONAL verdict already requests clarification and robustness checks, my concern reinforces that verdict rather than changing it. Therefore, I recommend no change to the reader's decision, while emphasizing that the window-length inconsistency must be one of the conditions for acceptance.","tokens_in":8390,"tokens_out":6556,"duration_ms":56918,"concrete_test":"Obtain the code or run the procedure twice, once with a 30-day rolling window and once with a 182-day rolling window, holding all other settings fixed. Reproduce Figure 5.1 for both configurations and check whether 'SP Coverage hours' remains the top metric across the majority of α values in both cases. If the ranking changes, the central claim is not robust to the window length and the manuscript must be revised to state and justify the correct choice. Also report bootstrap confidence intervals for the profit differences to assess whether the observed gaps are statistically significant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ranking by 'SP Coverage hours' yields superior average profit depends entirely on the rolling-window model selection described in §4.1. That section states 'a 30-day rolling window (see Figure 4.1) is applied' and the displayed sum runs from d=i to i+30. However, §4.2 states 'the third one for the calculation of statistical performance of probabilistic models for the considered metrics (182 days),' while Figure 4.1's caption (reproduced in the text) again says '30 days.' The manuscript therefore fails to specify whether the rolling average is computed over 30 or 182 days. This is not a cosmetic typo: the window length determines which model is selected each day, and different window lengths can rank the nine forecasting models differently. The headline 'clear winner' in §5 is the result of one specific but undisclosed choice. If the actual implementation used 182 days, the formula in §4.1 is wrong; if it used 30 days, the dataset description and the sentence in §4.2 are wrong. No sensitivity analysis over the window length is reported, so the reader cannot tell whether the superiority of SP Coverage hours is an artifact of this parameter. A second, related issue is the absence of any confidence interval or significance test in Figure 5.1: with a single out-of-sample path and no multiple-comparison control across the many α values, the observed gaps between metrics could be noise. The internal inconsistency compounds this because the ranking is not even based on a fully specified procedure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares statistical performance metrics for ranking probabilistic electricity price forecasting models, with the goal of identifying which ranking metric leads to the highest economic profit in a battery-storage trading strategy. Nine probabilistic forecasting models (built on a common point forecast) are ranked by six metrics derived from pinball loss and empirical coverage, including a metric computed only at the trading hours. A rolling out-of-sample selection exercise on German day-ahead prices (2015–2023) is used: each metric selects one model per day, and the selected model's forecasts drive the battery bid/offer decisions. Average per-trade profit is then compared across metrics for prediction-interval levels α = 50%–98%. The central claim is that ranking models by the coverage of quantile forecasts at the trading hours ('SP Coverage hours') yields the highest average profit.","tokens_in":8675,"tokens_out":3847,"duration_ms":36879,"significance":"If the central claim is robust, the paper makes a useful empirical contribution to the literature linking statistical forecast evaluation and economic value in electricity trading. The experimental design is a genuine rolling out-of-sample model selection exercise, rather than an in-sample fit, and the economic evaluation is transparently described. The comparison of several candidate metrics on a common set of forecasting models is a sensible way to study metric choice. However, the strength of the headline finding is currently difficult to assess because of an unresolved internal inconsistency in the rolling-window length, the absence of any uncertainty quantification for the profit differences in Figure 5.1, and the fact that the 'SP Coverage hours' metric is constructed almost exactly from the acceptance conditions that generate profit. The paper therefore needs additional work before its main conclusion can be accepted as stated.","major_comments":[{"comment":"The rolling-window length for computing the statistical performance metrics is stated inconsistently: §4.1 writes 'a 30-day rolling window' and the displayed sum runs from d=i to i+30, and the caption of Figure 4.1 says the third window is 30 days, while §4.2 states that the third rolling window for statistical performance is 182 days. This is not a purely cosmetic discrepancy, because the window length determines which forecasting model is selected each day, and the resulting ranking of metrics in Figure 5.1 depends on this choice. The authors must (i) state which length was actually used, (ii) correct the conflicting statements, and (iii) provide at least a sensitivity analysis over plausible window lengths, since the 'clear winner' status of SP Coverage hours may be specific to the reported implementation.","section":"§4.1 vs §4.2 and Figure 4.1"},{"comment":"The central comparison is presented without confidence intervals, significance tests, or any form of multiple-comparison control across the many values of α examined. The data are a single out-of-sample path (one market, one trading strategy, one calendar period), so the profit differences between metrics could be due to sampling variation. The authors should report uncertainty (e.g., block bootstrap or subperiod analysis) and, at minimum, discuss the magnitude of the profit gaps relative to their variability. Without this, the claim that SP Coverage hours 'exhibits a superior economic performance' is not statistically supported.","section":"§5, Figure 5.1"},{"comment":"The SP Coverage hours metric is defined as the joint event that the bid at hour h1 is accepted (actual price below the upper PI bound) and the offer at hour h2 is accepted (actual price above the lower PI bound). These are exactly the acceptance conditions in step (3) of the trading strategy that generate the profitable round-trip trade. Consequently, ranking models by historical SP Coverage hours is directly aligned with the frequency of profitable order pairs, and the 'superior economic performance' of this metric is at least partly mechanical rather than a finding about the generic value of coverage-based evaluation. The paper should explicitly discuss this endogeneity between the metric definition and the trading strategy, and should show whether the conclusion survives when the metric is not engineered to the strategy's acceptance conditions.","section":"§4.1 and §3.2.1"}],"minor_comments":[{"comment":"The second condition in the SP Coverage hours definition appears to have the wrong hour subscript: the text reads 'P_{d,h2} > \\hat P^{(1−α)/2}_{d,h1}', which should probably be '\\hat P^{(1−α)/2}_{d,h2}'.","section":"§4.1, displayed equation for SP Coverage hours"},{"comment":"The sentence 'The prices \\hat U^α_{d,h1} and \\hat L^α_{d,h1} are marked with black dots' should refer to \\hat L^α_{d,h2} rather than \\hat L^α_{d,h1}.","section":"§3.2.1, step (2)"},{"comment":"The conformal prediction formula has a stray bracket after 'for q > 50%]'.","section":"§2.2.2"},{"comment":"Two citations appear as unresolved '?' placeholders: the calibration-window reference in Conformal Prediction and the literature reference for Johnson's distribution. These should be completed.","section":"§2.2.2 and §2.2.3"},{"comment":"The word 'avarage' should be 'average'.","section":"§5"},{"comment":"The 'unlimited-bids benchmark' is described in Section 3.2.1 but never reported in the Results. Even if the paper's ranking claim is about relative performance across metrics, the benchmark is needed to position the economic magnitudes and to assess whether any statistical metric adds value over a naive price-taking strategy.","section":"§3.2.1 and §5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of econ.EM and addresses a timely question. The main concern for the editor is that the headline finding is built on an unresolved internal inconsistency (30 vs 182 days) and lacks inferential support. A further editorial consideration is that the 'SP Coverage hours' metric is so closely tailored to the trading strategy that the result may be interpreted as a validation of the metric construction rather than a general empirical law. I think the paper can be made acceptable after the window issue is resolved, the statistical significance is addressed, and the conceptual point about metric construction is discussed, but these are substantive revisions rather than copyediting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it does something genuinely useful: it compares statistical ranking metrics for probabilistic electricity price forecasts against the economic outcome (battery-storage trading profit) in a genuine out-of-sample rolling design. The SP Coverage hours metric is new and sensible—it checks whether the quantiles used for the actual bid and offer cover the realized prices. Second, the headline claim that this metric 'exhibits superior economic performance' is plausible but not yet proven, because the paper contains a load-bearing internal inconsistency and no statistical uncertainty quantification.\n\nThe experimental setup is a real strength. Nine forecasting models, a full year-by-year rolling calibration, and a realistic BESS trading strategy with battery state constraints. The out-of-sample period runs 2015–2023 on German day-ahead prices. That is a fair test bed, and the economic evaluation is not circular: profit is computed from actual accepted bids, and the ranking metrics are computed independently.\n\nThe soft spots are concrete. The most serious is the window-length contradiction. Section 4.1 says a 30-day rolling window is used to compute the statistical-performance metrics, and the equation sums from d=i to i+30. Section 4.2 says the third rolling window is 182 days. Figure 4.1's caption says 30 days. The manuscript never resolves this. It is not cosmetic: the window length determines which model gets selected each day, and different window lengths could rank the nine models differently. The 'clear winner' in Figure 5.1 is the output of one undisclosed choice. No sensitivity analysis over the window length is reported. This must be settled before the central claim can be evaluated.\n\nSecond, Figure 5.1 shows average profit differences across metrics and α values with no confidence intervals, significance tests, or multiple-comparison control. With a single out-of-sample path, those gaps could be noise. The stress-test note is right on both counts.\n\nThird, the unlimited-bids benchmark is described in Section 3.2.1 but never shown in the results. It would anchor the economic value of the quantile-based strategy. Minor: several citations are missing placeholders (e.g., '?'), which is sloppy but fixable.\n\nWho is this for? Researchers working on probabilistic electricity price forecasting, model ranking, and trading applications. It is a modest extension of Uniejewski (2024), not a breakthrough. If the window-length issue is resolved and uncertainty is addressed, the paper would be a solid applied contribution. As is, I would not take the headline result at face value, but it deserves a serious referee rather than a desk reject. Send it to review, and ask the authors to fix the inconsistency and add significance checks.","headline":"A plausible but under-supported empirical claim: trading-hour coverage ranks probabilistic price forecasts better than global pinball, but a window-length contradiction and missing uncertainty metrics need fixing before the result is trustworthy.","tokens_in":9198,"tokens_out":1552,"would_cite":false,"duration_ms":15154,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M20","91B84"],"pacs":[],"model":"deepseek-v4-flash","headline":"Ranking forecast models by the coverage of their traded quantiles maximizes battery-storage profit.","keywords":["probabilistic forecasting","electricity price forecasting","pinball loss","empirical coverage","quantile forecasts","model ranking","battery storage trading","day-ahead market"],"falsifier":"Resample the out-of-sample period with bootstrap and recompute the average per-trade profit for each ranking metric many times; if SP Coverage hours is not the top earner in the large majority of resamples, or if its margin over the runner-up overlaps with zero, then the 'clear winner' claim fails. Alternatively, repeating the comparison on a different day-ahead market and finding that another metric wins would refute the claimed superiority.","tokens_in":8159,"feed_emoji":"🔋","tokens_out":9422,"duration_ms":71866,"temperature":0.7,"pith_summary":"The paper asks which way of scoring probabilistic electricity-price forecasts, when used to pick the best forecasting model, leads to the largest profits from a battery-storage trading strategy. It introduces six ranking metrics built from pinball loss and empirical coverage, then tests them on German day-ahead prices from 2015 to 2023. The central finding is that ranking models by the coverage of the quantile forecasts that are actually used in the trading hours—the bid at the low-price hour and the offer at the high-price hour—beats all other metrics across nearly every prediction-interval level. This matters because it connects statistical forecast evaluation directly to economic decisions: the metric that rewards the right behavior in the trading strategy is the one that earns the most money.","feed_headline":"Traded-quantile coverage yields top battery-trading profit","feed_subtitle":"In German day-ahead data, ranking on the quantiles used in trading hours earns the most per trade.","key_machinery":"The key machinery is a set of six statistical performance metrics derived from pinball loss and empirical coverage, applied in a 30-day rolling window to select the best of nine probabilistic forecasting models for a battery-storage trading strategy. The strategy buys 1 MWh at the hour with the lowest predicted median price and sells at the highest, with bid and offer prices set at the $\\alpha$-level prediction interval bounds, using a 2.5 MWh battery with 90% round-trip efficiency. The decisive metric, 'SP Coverage hours', checks exactly what the trader cares about: for each day, whether the realized price is below the upper quantile of the prediction interval at the selected buying hour and above the lower quantile at the selling hour. By ranking models according to this targeted coverage instead of coverage over all hours or average pinball across all quantiles, the selection process rewards models whose intervals are reliable precisely where the trading decision is made.","core_discovery":"The paper's main claim is that, for the battery-storage trading strategy considered, ranking probabilistic forecasting models by 'SP Coverage hours'—an empirical-coverage metric that scores a day as a hit only when the realized price falls below the bid quantile at the buying hour and above the offer quantile at the selling hour—delivers the highest average profit per trade out-of-sample. This result holds across almost all values of the prediction-interval parameter $\\alpha$ from 50% to 98%. The paper further reports that all pinball-loss-based metrics perform similarly to each other, with average pinball across all hours and quantiles doing best among them, and that coverage averaged over all 24 hours is more variable. The conclusion is that the statistical metric most aligned with the economic objective—covering the quantiles that are actually used for bids and offers—is the one that identifies the best forecasting model.","pith_inferences":["An extension the paper leaves implicit is to test the same ranking metrics on other day-ahead markets (for example, Nordic or North American) to see whether the trading-hour coverage advantage is a stable property or an artifact of German price dynamics.","A bootstrap or subsample analysis of Figure 5.1 would tell whether the gap between SP Coverage hours and the next-best metric is statistically meaningful; the paper reports no uncertainty quantification.","The principle behind the winning metric generalizes naturally to any quantile-based trading rule: evaluate forecasts on the quantiles actually used in the decision, not on the full predictive distribution.","A separate testable extension is to vary the battery size, efficiency, or the choice of trading hours; if the advantage of trading-hour coverage depends on those design choices, the paper's practical recommendation becomes conditional."],"forward_implications":["A market participant using probabilistic day-ahead forecasts could switch from pinball-based model selection to trading-hour coverage and expect higher per-trade profits with no change to the forecasting models or the trading strategy.","Among pinball-loss-based metrics, the choice is less consequential, since they deliver similar profits for most prediction-interval levels.","Coverage averaged over all hours is a workable but more volatile criterion, which suggests that the information contained in the specific trading hours drives the economic gain.","The ranking metric can be computed from the same quantile forecasts already used for bidding, so the approach is operationally cheap to implement."],"supporting_citations":[{"why":"Supplies the battery-storage trading strategy and the pool of nine forecasting models whose rankings are compared.","marker":"Uniejewski (2024)"},{"why":"Provides the parsimonious autoregressive point forecast model that every probabilistic method builds on.","marker":"Ziel and Weron (2018)"},{"why":"Defines the pinball loss and proper scoring rules used to construct the pinball-based ranking metrics.","marker":"Gneiting and Raftery (2007)"},{"why":"Defines empirical coverage and interval hits, which underlie the coverage-based ranking metrics.","marker":"Chatfield (1993)"},{"why":"Introduces quantile regression averaging, one of the nine probabilistic models in the pool.","marker":"Nowotarski and Weron (2015)"},{"why":"Introduces smoothed quantile regression, the basis of the SQR averaging models.","marker":"Fernandes et al. (2021)"},{"why":"Supplies the conformal prediction approach used to generate one of the probabilistic forecast models.","marker":"Kath and Ziel (2021)"},{"why":"Provides the profit-per-MWh economic evaluation metric and related battery-storage trading context.","marker":"Maciejowska et al. (2024)"}],"fun_headline_variants":["Coverage of traded quantiles tops pinball loss in battery trading","Rank forecast models by traded-quantile coverage for best profit","Traded-quantile coverage wins over pinball in forecast ranking","Best battery trading profit from ranking on traded-hour coverage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a single out-of-sample path—German day-ahead prices from 2015 to 2023, one battery-storage trading strategy, and one set of nine forecast models—is enough to establish which ranking metric is best, and reports the winning metric without confidence intervals or significance tests.","fun_headline_variants_meta":{"raw":{"variants":["Coverage of traded quantiles tops pinball loss in battery trading","Rank forecast models by traded-quantile coverage for best profit","Traded-quantile coverage wins over pinball in forecast ranking","Best battery trading profit from ranking on traded-hour coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00073,"raw_usage":{"total_tokens":3191,"prompt_tokens":792,"completion_tokens":2399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":2328}},"tokens_in":408,"tokens_out":2399,"duration_ms":16205,"temperature":1.0,"reasoning_tokens":2328,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:38:48.894901+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Resample the out-of-sample period with bootstrap and recompute the average per-trade profit for each ranking metric many times; if SP Coverage hours is not the top earner in the large majority of resamples, or if its margin over the runner-up overlaps with zero, then the 'clear winner' claim fails. Alternatively, repeating the comparison on a different day-ahead market and finding that another metric wins would refute the claimed superiority.","supporting_citations":[{"cited_title":", author Weron, R","cited_arxiv_id":null,"evidence_quote":"Provides the parsimonious autoregressive point forecast model that every probabilistic method builds on."},{"cited_title":", year 1993","cited_arxiv_id":null,"evidence_quote":"Defines empirical coverage and interval hits, which underlie the coverage-based ranking metrics."},{"cited_title":", author Weron, R","cited_arxiv_id":null,"evidence_quote":"Introduces quantile regression averaging, one of the nine probabilistic models in the pool."},{"cited_title":", author Ziel, F","cited_arxiv_id":null,"evidence_quote":"Supplies the conformal prediction approach used to generate one of the probabilistic forecast models."},{"cited_title":", author Serafin, T","cited_arxiv_id":null,"evidence_quote":"Provides the profit-per-MWh economic evaluation metric and related battery-storage trading context."}],"review_version":1}