{"id":"b12b1041-0236-4808-8e5a-14c497234a39","arxiv_id":"2506.03744","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PC, the mean CRPS of isotonic distributional regression fitted post hoc to deterministic model output, offers a loss-function-independent way to compare AI and physics-based weather forecasts.","lead":"The paper introduces a new score, the potential continuous ranked probability score (PC), which compares single-valued AI and physics-based weather forecasts after applying the same statistical postprocessing to both. It shows that under this score, the data-driven GraphCast model outperforms the ECMWF HRES model on WeatherBench 2 data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample optimism in PC is not shown to be equal across models, so the headline ranking may be an artifact; a split-sample check is needed.","rationale":"I read the paper in good faith as proposing a well-defined, loss-function-independent measure of potential predictive ability, and the mathematical properties in Section 2.2 are sound. The reader's identified weakest assumption, isotonicity, is a legitimate structural premise, but it is explicitly stated and arguably defensible: for the univariate weather variables considered, stochastic monotonicity of the outcome in the deterministic forecast is a natural and interpretable constraint. A more load-bearing issue is the in-sample character of PC. Since EasyUQ is fitted and scored on the same evaluation data, the measure is an optimistically biased estimate of the expected CRPS of an EasyUQ-based probabilistic forecast, and nothing in the paper establishes that this optimism is equal across models with very different output characteristics. The paper even concedes the optimism in Section 2.1, but treats it as 'slightly' without supporting evidence. For the central claim that PC affords fair comparisons, the relevant null hypothesis is not 'the optimization is in-sample' but 'the in-sample optimization favors all models equally.' That hypothesis is untested. The proposed split-sample check would directly settle whether the headline ranking survives out-of-sample evaluation, and it would quantify the optimism gap per model. This concern is concrete, does not rely on contesting the stated definitions, and is consistent with the reader's CONDITIONAL verdict, though it points to a different primary weakness than the reader's isotonicity concern.","tokens_in":15885,"tokens_out":6600,"duration_ms":73941,"concrete_test":"Use a temporal split of the WeatherBench 2 IFS-analysis evaluation period: fit EasyUQ on the first half of 2020 and evaluate mean CRPS on the second half (or use blocked cross-validation with block length 2k to preserve serial dependence), for HRES, GC-IFS, and PW-IFS at each grid point and lead time; then compute latitude-weighted averages and compare rankings and gaps with Figure 2. Also compute the optimism gap as in-sample PC minus out-of-sample mean CRPS for each model. If GraphCast still dominates by a similar margin and the optimism gaps are approximately equal across models, the concern is resolved. If rankings change or the gaps shrink materially, the headline conclusion requires an out-of-sample or optimism-corrected version of PC.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that PC affords fair comparisons rests on the assumption that the in-sample optimism of EasyUQ is comparable across models. PC is defined in Section 2.1, eq. (4), as the mean CRPS of the EasyUQ solution fitted on the same evaluation data used for scoring ('post hoc (in-sample)'). This is an optimistically biased estimate of the expected CRPS of a fitted EasyUQ model, and the size of the bias depends on the number of observations n, the distribution and tie structure of the model output x, and the signal-to-noise relationship between x and the outcome y. The compared models in WeatherBench 2 (GraphCast, Pangu-Weather, HRES) have very different output characteristics, so there is no reason to assume the optimism cancels in the pairwise comparison. The paper acknowledges in Section 2.1 that PC is 'a slightly optimistic proxy' but provides no quantification of the optimism and no demonstration that it is approximately equal across models. Moreover, the headline empirical result in Figure 2, that GraphCast dominates both Pangu-Weather and HRES for all variables and lead times, is entirely in-sample and is presented without uncertainty quantification. If, for example, GraphCast's smoother or more finely resolved output permits more in-sample adaptation by isotonic regression, a substantial part of its apparent advantage could be an artifact of overfitting. This directly threatens the 'fair comparisons' claim, because the measure is meant to compare potential predictive ability, not in-sample training loss.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the potential continuous ranked probability score (PC), defined as the mean CRPS of the EasyUQ/IDR post-hoc probabilistic forecast fitted on the evaluation data itself, together with a skill score version PCS. The authors establish basic properties (nonnegativity, zero characterization, invariance under strictly increasing transformations of the model output, upper bound given by the climatological CRPS), illustrate them in a simulation, and apply the measures to WeatherBench 1 and WeatherBench 2 data. They report that in WeatherBench 2, GraphCast dominates Pangu-Weather and ECMWF HRES for all considered variables and lead times when evaluated with the IFS analysis as ground truth, and that PC for HRES aligns closely with the mean CRPS of the operational ECMWF ensemble. The paper argues that PC provides a fair, loss-function-independent comparison of deterministic AIWP and NWP backbones.","tokens_in":16097,"tokens_out":6451,"duration_ms":67042,"significance":"The proposal is attractive and original in its simplicity: applying the same CRPS-optimal isotonic postprocessing to every deterministic backbone removes the need to pre-specify a scoring function and addresses the unequal footing that arises when AI models are trained on the evaluation metric. The mathematical properties follow cleanly from established isotonic distributional regression theory, the computational cost is modest, and the authors provide code and replication material. The external check against the operational ensemble CRPS is a useful validation. The main risk is that PC is an in-sample fit of the evaluation objective, so its fairness across models depends on the optimism being comparable across models, which is neither shown nor tested. If this issue is resolved, the measure could be a valuable benchmark tool for the community.","major_comments":[{"comment":"The fairness claim rests on the assumption that the in-sample optimism of PC is comparable across models, but this is not demonstrated. PC is defined as the minimized in-sample mean CRPS of the EasyUQ fit, which is an optimistically biased estimate of the expected out-of-sample CRPS. The bias depends on n, the distribution and tie structure of the model output, and the signal-to-noise relationship; GraphCast, Pangu-Weather, and HRES outputs differ in these respects. The paper acknowledges that PC is \"a slightly optimistic proxy\" but provides no quantification or cross-model evidence of equal optimism. Because Figure 2 is entirely in-sample and lacks uncertainty quantification, the headline ranking could be partly an artifact of differential in-sample adaptability. A split-sample or cross-validated variant of PC (fit on training data, evaluate on holdout), or a bound/analysis of the optimism differences, is needed to support the \"fair comparisons\" claim. Alternatively, the claims should be weakened to comparisons of in-sample potential under the fitted isotonic model, rather than fair assessment of expected predictive ability.","section":"Section 2.1, Eq. (4) and Section 3.2, Figure 2"},{"comment":"The isotonicity constraint is a load-bearing assumption for the interpretation of PC as a measure of potential predictive ability. Property (b) shows that PC = 0 only when the outcome is a nondecreasing function of the model output, and the EasyUQ fit itself is constrained to be stochastically increasing in the model output. If a model encodes information non-monotonically (e.g., a change in sign or a U-shaped relationship), PC will understate that model's potential, and the comparison is no longer fair in the sense claimed in the abstract and Section 1. Since the paper explicitly targets \"across application domains,\" the scope of the fairness claim should be stated as applying only when isotonicity holds, or the paper should provide evidence that the variables and lead times considered satisfy this condition. This is not an optional caveat; it defines what PC measures.","section":"Section 2.2, property (b) and the definition of EasyUQ in Section 2"},{"comment":"The central empirical result that GraphCast dominates both Pangu-Weather and HRES for all considered variables and lead times is presented without any uncertainty quantification. The permutation test introduced in Section 3.2 and displayed in Figure 4 is applied to the comparisons in Figure 3 (ERA5 vs IFS ground truth) but not to the ranking in Figure 2. Since the differences in PC between models may be small at many grid points, and since PC itself is in-sample, the reader cannot assess whether the all-dominance pattern is statistically meaningful. I suggest adding block-bootstrap confidence intervals for the latitude-weighted PC differences or a permutation test for the Figure 2 comparisons, following the paper's own method.","section":"Section 3.2, Figure 2 and Figure 4"},{"comment":"The proxy claim for the operational ensemble CRPS is not fully supported by the evidence presented. The paper does not state how the operational CRPS values in Figure 7 were obtained (which dataset, what ground truth or observations, what period), nor does it explain the grid-point matching. The heuristic that optimistic and pessimistic effects \"balance each other\" is plausible but is only tested for HRES, not for AIWP models. Please specify the data source and computation of the operational CRPS, and consider showing similar scatterplots for GraphCast and Pangu-Weather if such data are available; otherwise, soften the claim to a case study for HRES.","section":"Section 3.3, Figure 7"}],"minor_comments":[{"comment":"There are several typos: \"ground thruth\" should be \"ground truth,\" and \"PC-ERA5\" in the sentence comparing HRES, GC-ERA5, and PC-ERA5 should presumably be \"PW-ERA5.\" Also, Section 2.1 has \"addresss\" for \"addresses.\"","section":"Section 3.2"},{"comment":"The block permutation procedure is described very tersely; please clarify how p-values are computed at the grid-point level, especially with respect to the two daily runs and the handling of the block length of 2k.","section":"Footnote 2 of Section 3.2"},{"comment":"The paragraph that calls PC a \"slightly optimistic proxy\" and then a \"pessimistic proxy\" is confusing; the two directions of bias refer to different baselines (in-sample fitting versus operational product sophistication) and should be stated more explicitly and consistently.","section":"Section 2.1"},{"comment":"The latitude-weighted PC values are spatially and temporally correlated, yet the figure shows no measure of uncertainty; even a brief statement about the dependence structure and its implications would be helpful.","section":"Section 3.2, Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written and the core idea is appealing. The main concern is the unquantified in-sample optimism of PC; if the authors can provide a split-sample analysis or a theoretical argument that the optimism is comparable across models, this would substantially strengthen the fairness claim. I would be open to accepting after a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"PC is a sensible and clearly-argued measure: take the mean CRPS of the EasyUQ isotonic fit computed in-sample on the evaluation data, and use that as a benchmark for deterministic forecasts. The authors are explicit that this is a potential score, not an out-of-sample estimate, and they prove the natural invariance and zero-optimality properties. The WeatherBench 2 application is careful about matching ground-truth analyses, and the proxy validation against the operational ECMWF ensemble for HRES is a genuine empirical asset. The paper deserves a serious referee.\n\nWhere it's soft: first, the headline result in Figure 2, that GC-IFS beats PW-IFS and HRES everywhere, comes with no uncertainty quantification. The permutation tests used later for the mixed-ground-truth comparisons would have been appropriate here; without them, the 'dominates' claim is a point estimate. Second, the fairness claim implicitly assumes the in-sample optimism of EasyUQ is comparable across models. The paper acknowledges the optimism but doesn't check whether it's balanced. A split-sample or cross-validation analysis, where the EasyUQ fit is built on a training segment and scored on a holdout, would directly address whether the ranking survives. This is the most consequential gap. The isotonicity constraint is the structural limitation: if a model's output is informative but non-monotone, PC understates its potential. The authors note the constraint but don't discuss when it might be a problem; weather variables are mostly monotone, so I'd treat this as a caveat rather than a fatal flaw. Minor: the simulation example is clean but does no more than illustrate invariance; the 'slightly optimistic proxy' phrasing is accurate but the size of the optimism is never quantified.\n\nWho this is for: anyone building or using benchmarks for deterministic forecast models, especially in the AIWP/NWP debate. A statistically mature reader will see the measure as a functional of the data rather than a predictive score, and will want the robustness checks above. With those, it could become the default way to compare deterministic backbones. I would send this to peer review. The core idea is sound and the paper is transparent. The referee should ask for uncertainty bounds on Figure 2 and a split-sample check on the optimism balance.","headline":"A clear, honest proposal for a potential-score benchmark, but the headline GraphCast result needs uncertainty bounds and the optimism-parity assumption should be tested.","tokens_in":16697,"tokens_out":4684,"would_cite":true,"duration_ms":45034,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M20","62G05","62G08","62P12"],"pacs":["92.60.Ry"],"model":"deepseek-v4-flash","headline":"This paper introduces the potential continuous ranked probability score (PC), a loss-function-free measure built by post-processing deterministic forecasts with isotonic distributional regression, and reports that under PC the data-driven…","keywords":["potential CRPS","isotonic distributional regression","EasyUQ","forecast evaluation","AIWP","NWP","WeatherBench 2","continuous ranked probability score"],"falsifier":"Fit PC on data where the outcome is a strictly decreasing function of the model output. The score will be positive even though the forecast is perfect, because isotonicity forbids decreasing relationships; that outcome would refute any claim that PC fairly compares arbitrary model encodings.","tokens_in":15657,"feed_emoji":"🌦️","tokens_out":9172,"duration_ms":84766,"temperature":0.7,"pith_summary":"This paper introduces the potential continuous ranked probability score ($\\mathrm{PC}$), a measure meant to make single-valued forecasts from AI-based and physics-based weather models comparable without choosing a loss function. The trick is to post-process every deterministic forecast in the evaluation set with the same automated method, isotonic distributional regression, and take the mean continuous ranked probability score of the resulting probabilistic forecasts. The paper argues this yields a fair, loss-independent yardstick: $\\mathrm{PC}$ is invariant under strictly increasing transformations of model output, is measured in the units of the weather variable, and is zero exactly when the outcome is a fixed non-decreasing function of the forecast. On WeatherBench 2 data, the paper finds that GraphCast dominates Pangu-Weather and ECMWF's high-resolution model under $\\mathrm{PC}$ for all three studied variables and all lead times, and that HRES's $\\mathrm{PC}$ closely tracks the operational ECMWF ensemble's mean CRPS. If the measure is accepted, it changes how AIWP and NWP backbones are benchmarked, removing the objection that AI models are tuned to the metric being reported.","feed_headline":"GraphCast tops ECMWF under new fair score","feed_subtitle":"It post-processes every deterministic forecast identically, then scores skill without a pre-chosen loss function.","key_machinery":"The load-bearing object is the EasyUQ solution, a special case of isotonic distributional regression: given paired model output and outcomes, it finds the unique predictive distributions that minimize the average CRPS subject to the quantile functions being nondecreasing in model output. Pool-adjacent-violators algorithms compute this solution in about $O(n \\log n)$ operations, and the $\\mathrm{PC}$ score is then the mean CRPS of these fitted distributions. This machinery gives $\\mathrm{PC}$ its invariance under strictly increasing transformations, its zero-in-the-perfect-monotone-predictor property, and its interpretation as a slightly optimistic proxy for a real-time postprocessed probabilistic forecast.","core_discovery":"The paper's central claim is that the $\\mathrm{PC}$ measure, defined as the in-sample mean CRPS of the EasyUQ/IDR postprocessed predictive distribution, gives a principled and fair way to compare deterministic forecast output across models trained under different objectives. Under this measure, the paper reports, GraphCast outperforms both Pangu-Weather and ECMWF HRES for mean sea level pressure, 2 m temperature, and 10 m wind speed at lead times from one to ten days in WeatherBench 2 when all models are evaluated against the IFS analysis; Pangu-Weather generally beats HRES for pressure and wind, while HRES beats Pangu-Weather for temperature. The paper also claims that HRES's $\\mathrm{PC}$ is a close proxy for the mean CRPS of the operational ECMWF ensemble, with slight optimism at one day and pessimism at longer leads. The measure achieves these properties by construction: $\\mathrm{PC}$ is invariant under strictly increasing transformations of model output, equals zero if and only if the outcome is a non-decreasing function of the forecast, and is bounded above by the in-sample climatological CRPS.","pith_inferences":["Inference: the isotonicity constraint means PC actually measures monotone information content; a non-monotone model would be undervalued, so PC is best read as a lower bound on potential skill under monotone post-processing.","Inference: the same construction transfers to other single-valued forecast settings where competitors train under different loss functions, such as energy load, finance, or epidemiology, whenever an evaluation set of a few hundred to a few thousand paired forecasts and outcomes is available.","Inference: replacing CRPS in the defining criterion with threshold- or quantile-weighted CRPS would yield a tail-focused version of PC aimed at extreme weather, a direction the paper sketches but does not implement.","Inference: a directed stress test would invert a skilful model's output, making the forecast-outcome relation strictly decreasing; a fully fair skill measure should still credit it, while PC would not unless the ordering is restored."],"forward_implications":["Benchmarks of deterministic weather forecasts can be run without pre-specifying a loss function, because PC treats every model's output through the same post-processing lens.","AI models lose the advantage of having been trained on the exact RMSE or other metric used for evaluation, since PC is defined by the evaluation data alone.","PC values for an operational deterministic backbone can serve as a check on the expected skill of the operational ensemble, flagging lead times where the ensemble adds little.","Because PCS is a normalized skill score with a meaningful climatology baseline, forecasters can compare model families across variables and lead times on a common 0-to-1 scale.","The invariance property means ranking by PC is stable under monotone rescaling of forecasts, so unit changes or calibrations of model output do not distort comparisons."],"supporting_citations":[{"why":"Supplies isotonic distributional regression, whose optimal CRPS solution under the isotonicity constraint is the mathematical core of PC.","marker":"Henzi et al. (2021)"},{"why":"Introduces EasyUQ, the automated post-processing used to turn deterministic output into predictive distributions; PC is the in-sample mean CRPS of this solution.","marker":"Walz et al. (2024a)"},{"why":"Gives the CRPS and its representation used to derive PC's units, bounds, and scoring properties.","marker":"Gneiting and Raftery (2007)"},{"why":"Provides WeatherBench 2, the benchmark data, ground-truth analysis products, and model runs on which the empirical comparisons rest.","marker":"Rasp et al. (2024)"},{"why":"Defines GraphCast, the data-driven model that the paper reports beats Pangu-Weather and HRES under PC.","marker":"Lam et al. (2023)"},{"why":"Defines Pangu-Weather, the second data-driven model in the WeatherBench 2 comparison.","marker":"Bi et al. (2023)"},{"why":"Proposes the lagged-ensemble probabilistic benchmark that motivates and contrasts with PC as a proxy for ensemble CRPS.","marker":"Brenowitz et al. (2025)"},{"why":"Supplies the convention of restricting direct comparisons to models using the same analysis product, which the paper follows in its main WeatherBench 2 analysis.","marker":"Ben Bouallègue et al. (2024)"}],"fun_headline_variants":["PC score: GraphCast beats ECMWF on fair terms","New fair measure crowns GraphCast over ECMWF","GraphCast wins under transformation-free PC metric","No preset loss? PC score still picks GraphCast","AIWP edges NWP with equalized post-processing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison depends on the assumption that all useful predictive information in a deterministic forecast can be extracted by a monotone transformation; if a model encodes information non-monotonically, PC understates its potential and the comparison is no longer fair.","fun_headline_variants_meta":{"raw":{"variants":["PC score: GraphCast beats ECMWF on fair terms","New fair measure crowns GraphCast over ECMWF","GraphCast wins under transformation-free PC metric","No preset loss? PC score still picks GraphCast","AIWP edges NWP with equalized post-processing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000777,"raw_usage":{"total_tokens":3498,"prompt_tokens":1073,"completion_tokens":2425,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":689,"completion_tokens_details":{"reasoning_tokens":2351}},"tokens_in":689,"tokens_out":2425,"duration_ms":19574,"temperature":1.0,"reasoning_tokens":2351,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:56:03.891840+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit PC on data where the outcome is a strictly decreasing function of the model output. The score will be positive even though the forecast is perfect, because isotonicity forbids decreasing relationships; that outcome would refute any claim that PC fairly compares arbitrary model encodings.","supporting_citations":[{"cited_title":"F., and Gneiting, T","cited_arxiv_id":null,"evidence_quote":"Supplies isotonic distributional regression, whose optimal CRPS solution under the isotonicity constraint is the mathematical core of PC."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides WeatherBench 2, the benchmark data, ground-truth analysis products, and model runs on which the empirical comparisons rest."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines GraphCast, the data-driven model that the paper reports beats Pangu-Weather and HRES under PC."},{"cited_title":"D., Cohen, Y., Pathak, J., Mahesh, A., Bonev, B., Kurth, T., Durran, D","cited_arxiv_id":null,"evidence_quote":"Proposes the lagged-ensemble probabilistic benchmark that motivates and contrasts with PC as a proxy for ensemble CRPS."},{"cited_title":"C., Magnusson, L., Gascon, E., Maier-Gerber, M., Janou s ek, M., Rodwell, M., Pinault, F., Dramsch, J","cited_arxiv_id":null,"evidence_quote":"Supplies the convention of restricting direct comparisons to models using the same analysis product, which the paper follows in its main WeatherBench 2 analysis."}],"review_version":1}