REVIEW 3 major objections 5 minor 5 cited by
The paper claims that forecast realism splits into three distinct forms — per-instance accuracy, statistical consistency, and physical plausibility — and that data-driven forecasts need a falsification test, not just scores, to check the th
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 05:55 UTC pith:ERWHP6JX
load-bearing objection A useful conceptual taxonomy of forecast realism, but the new 'falsification' leg is a placeholder until F=f(X_i,K) is instantiated. the 3 major comments →
"What is a realistic forecast?" Assessing data-driven weather forecasts, a journey from verification to falsification
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that forecast realism, like forecast goodness before it, should be understood through three distinct lenses rather than one. Type 1, functional realism, is the per-instance closeness of a forecast to an observation, assessed by scoring functions; Type 2, structural realism, is the average statistical consistency between forecast and observation, assessed by diagnostics; Type 3, physical realism, is the compatibility of a forecast with scientific knowledge, to be assessed by a falsification test F=f(X_i,K). The paper argues that while Types 1 and 2 are routinely evaluated, Type 3 becomes decisive for data-driven forecasts because inductive training can produce outputs tha
What carries the argument
The load-bearing mechanism is the falsification equation F=f(X_i,K), where K is a formal representation of the knowledge base of physics. The paper defines three measures — V=v(x_i,y_i) for verification, D=d(X,Y) for diagnostics, F=f(X_i,K) for falsification — and assigns each to one type of realism. The new element is the explicit introduction of K as a reference for evaluation: rather than comparing forecasts only against observations, one compares them against a codified understanding of what is physically possible. The paper uses this device to draw a hierarchy among realism types and to frame hallucinations in generative models as errors that scores alone will not reveal.
Load-bearing premise
The framework stands or falls on whether a knowledge base K of consolidated physics can be formalized precisely enough to run a falsification test F=f(X_i,K) that separates physically possible from impossible forecasts; the paper itself concedes that quantitative physical realism is a source of debate at the time of writing.
What would settle it
Take a data-driven weather model and run it on a set of initial states in which a conserved quantity (for example total energy) is perturbed; if the forecasts evolve without respecting the conservation law while still scoring well under standard verification and diagnostics, then falsification is demonstrably a separate, necessary evaluation step. If, however, no such case can be found — if every physically impossible forecast is already flagged by scores or reliability metrics — then the paper's claim that falsification adds something beyond verification and diagnostics would be refuted.
If this is right
- Evaluating a data-driven weather model will require a three-part report: verification scores, reliability diagnostics, and a falsification check against physical constraints.
- A model that tops the leaderboard in RMSE can still be physically unrealistic; model rankings should not rely on accuracy alone.
- The 'accuracy versus activity' trade-off is a manifestation of the tension between Type 1 and Type 2 realism, and physical realism adds an independent third constraint that must be reported.
- For applications like climate projection, where generalisation beyond the training sample matters, physical realism becomes a necessary condition for trust, not an optional extra.
- A quantitative falsification test, once K is specified, would allow comparison of models and definition of acceptable realism thresholds per application.
Where Pith is reading between the lines
- If K is built from conservation laws and dynamical constraints, falsification could be automated and would likely catch artifacts such as non-physical energy drift or negative precipitation that current scoring functions miss.
- The tripartite realism scheme extends naturally beyond weather: any data-driven model in a physical science — ocean, climate, hydrology — could adopt the same split between per-instance accuracy, statistical consistency, and physical plausibility.
- The hierarchy described in the thought experiment suggests a development ladder for ML weather models: physical realism comes first, then structural, then functional; that ordering could guide where to invest evaluation effort.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position/commentary paper asks what it means for a (data-driven) weather forecast to be realistic, and proposes a three-part taxonomy inspired by Murphy's (1993) analysis of forecast goodness: functional realism (§2.1), the instance-wise closeness of forecast to observation, measured by scoring functions; structural realism (§2.2), the statistical consistency of forecast and observation distributions, assessed by diagnostics such as bias, activity bias, or spectral measures; and physical realism (§2.4), compatibility with a knowledge base K of consolidated physics, checked by a falsification test F=f(X_i,K). The paper argues that evaluation of data-driven weather forecasts should combine verification, diagnostics, and falsification, discusses relationships among the three types, a perfect-forecast hierarchy (§3.1), fit-for-purpose considerations (§3.2), validation and case-study analysis (§3.3), trust and interpretability (§3.4), information content (§3.5), and closes with a 'typical journey' of three evaluation stations (§3.7).
Significance. If the proposed taxonomy were accepted, it would give the AI/NWP verification community a much-needed shared vocabulary, extending Murphy's goodness framework to the currently central but imprecise notion of 'realism.' The paper is a conceptual contribution rather than an empirical or formal one: there are no data, no fitted parameters, and the equations are simple and internally consistent. Its main strengths are the clear borrowing from Murphy, the explicit separation of per-instance accuracy from distributional consistency, and the honest admission — repeated in the text — that quantitative physical-realism assessment is not yet attained. The paper also usefully brings hallucination detection and subjective verification into the same discussion as standard verification metrics. However, the central new formal object, F=f(X_i,K), is only sketched, and the paper itself labels its operationalization 'a source of debate at the time of writing.' This gap is the main barrier to accepting the paper as a complete framework.
major comments (3)
- [§2.4, Eq. (3)] The only novel formal element of the paper, the falsification test F=f(X_i,K), is not operationalized. No null model, no tolerance, no error rate, no representation of K, and no aggregation procedure are specified. The text explicitly concedes that 'how to attain this objective (and wether it is attainable) is a source of debate at the time of writing.' Since the title and abstract promise falsification as a complementary evaluation process, this concession leaves the central claim as a programmatic metaphor rather than an assessment method. Please either provide a concrete recipe for at least one nontrivial F (e.g., conservation-law residual with a threshold and a statistical test), or clearly restrict the claim to a qualitative, science-driven review step.
- [§2.4, hallucination paragraph] The paper's only concrete detection mechanism for hallucinations — the phenomenon Type-3 realism is meant to catch — is 'a thorough review of individual forecasts' by knowledgeable humans, described as subjective verification. This is a legitimate and time-honored practice, but it is not a quantitative hypothesis test of the form Eq. (3). The cited exemplars (Hakim and Masanam 2024; Bonavita 2024) are bespoke one-off dynamical checks, not a general f parameterized by a knowledge base. The claimed complementarity of verification, diagnostics, and falsification therefore rests on an unestablished link between subjective review and the formal notation of Eq. (3). The author should either make this link explicit or present subjective review as one legitimate implementation of Type-3 assessment alongside, not conflated with, a formal falsification procedure.
- [§3.1 and §3.7] The perfect-forecast hierarchy in §3.1 is a useful thought experiment, but it sidesteps a structural issue: Eq. (3) is defined per instance x_i, while §3.7 applies falsification to a model as a whole ('a comparison of a forecast with our understanding of reality... to identify artifacts'). It is unclear whether falsification is a property of a single forecast, a sample, or a model, and how the results are aggregated (fraction of cases flagged? maximum violation? areal threshold?). The treatment of observations that themselves fall outside K — record extremes, or states that are physically possible but outside a finite knowledge base — is also absent. Adding an explicit discussion of the unit of analysis and of observational error in Eq. (3) would materially strengthen the framework.
minor comments (5)
- [§2.4] Typo: 'wether' should be 'whether'; also 'a so called falsification test' should be 'a so-called falsification test.'
- [Fig. 2] Panel C is labeled 'Not valid given K,' which is the output of a falsification test, not the process itself. The caption could clarify what V, D, and F actually return (a score, a summary measure, a binary decision).
- [§2.1] The discussion of elicitable functionals and proper scoring rules is brief to the point of being cryptic for a reader outside verification. A single sentence defining elicitability, or a pointer to Gneiting and Raftery, would improve readability.
- [References] The DOI for Ben Bouallègue and the AIFS team (2024) reads '0.21957/8b50609a0f'; this looks like a missing '10.' before '0.21957'.
- [§3.3] The sentence 'Type 3 of realism is closely related to Type 1 of goodness' is slightly confusing because Murphy's Type 1 goodness is about forecaster's true belief, not scientific knowledge; the connection would benefit from a one-sentence unpacking.
Circularity Check
No significant circularity: the paper is a conceptual framework, not a derivation from fitted inputs.
full rationale
The paper does not derive predictions from fitted parameters or from inputs that already contain the conclusions. Its three equations are definitional placeholders: V=v(x_i,y_i), D=d(X,Y), and F=f(X_i,K) formalize existing or proposed evaluation activities rather than being used to compute novel results. The taxonomy of functional, structural, and physical realism is explicitly inspired by Murphy (1993), an external source, and the paper acknowledges the limits of the third type: Section 2.4 states that a quantitative assessment of physical realism and whether it is attainable "is a source of debate at the time of writing." That is an admitted open problem, not a circular move. The paper cites several works by the same author and ECMWF colleagues, but these citations are contextual (e.g., demonstrating ML skill, operationalization, the accuracy-versus-activity trade-off) and do not carry the central claim. The trade-off reference is also anchored in Murphy (1993), so it is not a self-citation chain standing in for independent support. There is no fitted input renamed as prediction, no uniqueness theorem imported by fiat, and no ansatz smuggled in via citation. The skeptical concern that the falsification test is underspecified is a completeness/operationalization issue, not circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Observations can stand as 'truth' for forecast verification.
- ad hoc to paper A knowledge base K of consolidated physics can demarcate physically possible from impossible forecasts.
- domain assumption Murphy's three types of forecast goodness map one-to-one onto three types of forecast realism.
- standard math Proper scoring rules guarantee that improving reliability improves functional scores.
read the original abstract
The artificial intelligence revolution is fuelling a paradigm shift in weather forecasting: forecasts are generated with machine learning models trained on large datasets rather than with physics-based numerical models that solve partial differential equations. This new approach proved successful in improving forecast performance as measured with standard verification metrics such as the root mean squared error. At the same time, the realism of data-driven weather forecasts is often questioned and considered an Achilles' heel of machine learning models. How forecast realism can be defined and how this forecast attribute can be assessed are the two questions simultaneously addressed here. Inspired by the seminal work of Murphy (1993) on the definition of forecast goodness, we identify 3 types of realism: a functional realism measured by scoring functions, a structural realism related to the statistical characteristics of the forecasts, and a physical realism that is apprehended through the lenses of our scientific knowledge. This conceptual setting serves as a basis for the design of a new framework for the evaluation of data-driven weather models where falsification arises as a complementary process to the well-established diagnostic and verification tasks.
Figures
Forward citations
Cited by 5 Pith papers
-
WP-MIP: An Artificial Intelligence, Hybrid, and Physically Based Model Intercomparison Project for Weather Prediction
WP-MIP creates a multi-center archive of physical, AI, and hybrid model forecasts and shows AI models match large-scale skill while suppressing small-scale variability.
-
Geometric coherence of single-cell CRISPR perturbations reveals regulatory architecture and predicts cellular stress
Shesha quantifies directional coherence of single-cell CRISPR responses as mean cosine similarity of shift vectors, correlating with magnitude while identifying pleiotropic regulators and stress associations across fi...
-
Geometric coherence of single-cell CRISPR perturbations reveals regulatory architecture and predicts cellular stress
Shesha quantifies directional coherence of single-cell CRISPR responses, correlates strongly with effect magnitude, distinguishes pleiotropic from lineage-specific regulators, and predicts chaperone activation after m...
-
Geometric coherence of single-cell CRISPR perturbations reveals regulatory architecture and predicts cellular stress
Shesha perturbation stability Sp measures directional coherence of single-cell CRISPR shifts and predicts UPR stress and directional reproducibility beyond magnitude and Song et al.'s PS.
-
WP-MIP: An Artificial Intelligence, Hybrid, and Physically Based Model Intercomparison Project for Weather Prediction
WP-MIP creates a centralized forecast database and evaluation framework to compare physically based, machine-learning, and hybrid weather prediction models across global centers.
Reference graph
Works this paper leans on
-
[1]
https://philsci-archive.pitt.edu/26075/
Andrews, M., 2023: The devil in the data: Machine learning & the theory-free ideal. https://philsci-archive.pitt.edu/26075/
2023
-
[2]
Bauer, P., A. Thorpe, and G. Brunet, 2015: The quiet revolution of numerical weather prediction.Nature,525, 47–55. Ben Bouallègue, Z., M. Clare, and M. Chevallier, 2024: Prévisions météorologiques reposant sur l’intelligence artificielle : Une révolution peut en cacher une autre.La Météorologie,126, 48–52, https://doi.org/10.37053/lameteorologie-2024-0058...
-
[3]
Bjerknes, V ., 1904: Das Problem der Wettervorhersage, betrachtet vom Standpunkte der Mechanik und der Physik (The problem of weather prediction, considered from the viewpoints of mechanics and physics).Meteorologische Zeitschrift,21, 1–7 (translated and edited by VOLKEN E. and S. BR¨ONNIMANN. – Meteorol. Z. 18 (2009), 663–667), https://doi.org/https://do...
-
[4]
Bonavita, M., 2024: On some limitations of current machine learning weather prediction models.Geophysical Research Letters,51 (12), e2023GL107 377, https://doi.org/https://doi.org/10.1029/2023GL107377
-
[5]
Boucher, E., G. Moldovan, Z. Ben Bouallègue, N. Raoult, M. Maier-Gerber, M. Clare, and J. D. L. Brazidec, 2025: La prévision du temps par IA devient opérationnelle au CEPMMT.La Météorologie,130, 4–7, https://doi.org/10.37053/lameteorologie-2025-0051. Bröcker, J., 2009: Reliability, sufficiency, and the decomposition of proper scores.Quarterly Journal of t...
-
[6]
Charlton-Perez, A., H. Dacre, and S. e. a. Driscoll, 2024: Do AI models produce better weather forecasts than physics-based models? A quantitative evaluation case study of Storm Ciarán.npj Clim Atmos Sci 7, 93, https://doi.org/10.1038/s41612-024-00638-w
-
[7]
Ebert-Uphoff, I., and Coauthors, 2025: Measuring Sharpness of AI-Generated Meteorological Imagery.Artificial Intelligence for the Earth Systems,4 (3), e240 083, https://doi.org/10.1175/AIES-D-24-0083.1
-
[8]
Gneiting, T., T. Biegert, K. Kraus, E.-M. Walz, A. I. Jordan, and S. Lerch, 2025: Probabilistic measures afford fair comparisons of AIWP and NWP model output. https://arxiv.org/abs/2506.03744
Pith/arXiv arXiv 2025
-
[9]
Hakim, G. J., and S. Masanam, 2024: Dynamical tests of a deep learning weather prediction model.Artificial Intelligence for the Earth Systems,3 (3), e230 090, https://doi.org/10.1175/AIES-D-23-0090.1
-
[10]
Harrison, D. R., A. McGovern, C. D. Karstens, A. Bostrom, J. L. Demuth, I. L. Jirak, and P. T. Marsh, 2025: An assessment of how domain experts evaluate machine learning in operational meteorology.Weather and Forecasting,40 (3), 393 – 410, https://doi.org/10.1175/W AF- D-24-0144.1
doi:10.1175/w 2025
-
[11]
Hersbach, H., and Coauthors, 2020: The ERA5 global reanalysis.Quarterly Journal of the Royal Meteorological Society,146 (730), 1999– 2049, https://doi.org/https://doi.org/10.1002/qj.3803
doi:10.1002/qj.3803 2020
- [12]
-
[13]
Lynch, P., 2008: The origins of computer weather prediction and climate modeling.Journal of Computational Physics,227 (7), 3431–3444, https://doi.org/https://doi.org/10.1016/j.jcp.2007.02.034
-
[14]
Magnusson, L., 2017: Diagnostic methods for understanding the origin of forecast errors.Quarterly Journal of the Royal Meteorological Society,143 (706), 2129–2142, https://doi.org/https://doi.org/10.1002/qj.3072
-
[15]
McGovern, A., R. Lagerquist, D. J. Gagne, G. E. Jergensen, K. L. Elmore, C. R. Homeyer, and T. Smith, 2019: Making the black box more transparent: Understanding the physical implications of machine learning.Bulletin of the American Meteorological Society,100 (11), 2175 – 2199, https://doi.org/10.1175/BAMS-D-18-0195.1
-
[16]
Moldovan, G., and Coauthors, 2025: AIFS 1.1.0: An update to ECMWF’s machine-learned weather forecast model AIFS.EGUsphere,2025, 1–23, https://doi.org/10.5194/egusphere-2025-4716
-
[17]
3rd ed., https://christophm.github.io/interpretable-ml-book
Molnar, C., 2025:Interpretable Machine Learning. 3rd ed., https://christophm.github.io/interpretable-ml-book
2025
-
[18]
Murphy, A. H., 1993: What Is a Good Forecast? An Essay on the Nature of Goodness in Weather Forecasting.Weather and Forecasting, 8 (2), 281 – 293, https://doi.org/10.1175/1520-0434(1993)008<0281:WIAGFA>2.0.CO;2
-
[19]
Olivetti, L., and G. Messori, 2025: Whose weather is it? a fairness framework for data-driven weather forecasting.Environmental Research Letters,20 (12), https://doi.org/10.1088/1748-9326/ae21f5
-
[20]
Popper, K., 1959:The Logic of Scientific Discovery
1959
-
[21]
https://arxiv.org/abs/2504.08526, 2504.08526
Rathkopf, C., 2025: Hallucination, reliability, and the role of generative ai in science. https://arxiv.org/abs/2504.08526, 2504.08526
arXiv 2025
-
[22]
Richardson, D. S., 2000: Skill and relative economic value of the ecmwf ensemble prediction system.Quarterly Journal of the Royal Meteorological Society,126 (563), 649–667, https://doi.org/https://doi.org/10.1002/qj.49712656313
-
[23]
Rodwell, M. J., M. C. A. Clare, S.-J. Lock, K. Lonitz, and M. Chevallier, 2025: Power spectra of physics-based and data-driven ensembles. Meteorological Applications,32 (5), e70 071, https://doi.org/https://doi.org/10.1002/met.70071
-
[24]
Selz, T., W. Bruinsma, G. C. Craig, S. Markou, R. Turner, and A. Vaughan, 2025: On the effective resolution of AI weather prediction models. https://doi.org/10.22541/essoar.174139239.94807670/v1
arXiv 2025
-
[25]
Sha, Y ., J. S. Schreck, W. Chapman, and D. J. G. II, 2025: Improving AI weather prediction models using global mass and energy conservation schemes. 2501.05648
Pith/arXiv arXiv 2025
-
[26]
Wilson, and W
Stanski, H., L. Wilson, and W. Burrows, 1989: Survey of Common Verification Methods in Meteorology (WMO Research Report No. 89-5)
1989
-
[27]
University of Chicago Press, Chicago, IL
Winsberg, E., 2010:Science in the Age of Computer Simulation. University of Chicago Press, Chicago, IL. 11
2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.