{"id":"8a8c7539-3246-4bfd-bfd8-e9767562d3c3","arxiv_id":"2412.11973","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A hybrid differentiable GCM trained online on satellite precipitation data simulates and forecasts precipitation more accurately than CMIP6 models, ERA5, and ECMWF's ensemble.","lead":"Researchers built a hybrid weather model by wiring a neural network into a physics-based atmospheric model and training it to match satellite rain observations. The model produces more realistic rain patterns, extremes, and daily timing than standard climate models, and beats a top forecasting system for rain forecasts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Extremes and diurnal-cycle claims are evaluated only against the training reference (IMERG); without an independent reference for these metrics, 'substantially surpasses' is not yet established.","rationale":"Choosing between the candidate weaknesses: seed selection, evaporation diagnosis, sub-6-hour artifacts, and training-target evaluation. The evaporation issue (Fig. S21) is acknowledged by the authors and is not directly falsifying: the precipitation output can be realistic while the diagnosed residual evaporation is noisy, and the central claim concerns precipitation, not evaporation. Seed selection is a limitation but the selected model's stability across 37 initial conditions and 731/732 22-year runs is independently demonstrated. The sub-6-hour artifacts are explicitly flagged in the supplementary and the paper conditions its diurnal claim to the harmonic representation. The training-target evaluation is the most load-bearing because it bears directly on the central claim's 'substantially surpasses ... extremes, and diurnal cycle.' The model's IMERG CRPS optimization makes agreement with IMERG partially a matter of curve fitting. The paper provides GPCP checks for mean and forecasts, which is good, but notably omits GPCP or another independent product for the two headline categories (extremes and diurnal cycle) that are most product-sensitive. A concrete, feasible check—recomputing Rx1day/percentile and diurnal phase against GPCP and CMORPH/GSMaP—would settle whether the headline claim generalizes. This is an addressable gap rather than an internal inconsistency; the paper is otherwise well-executed, with public code/checkpoints and substantial evidence. I therefore maintain the reader's CONDITIONAL verdict (no change), with the condition made explicit: the extreme/diurnal-cycle superiority claims require an independent reference.","tokens_in":25377,"tokens_out":8205,"duration_ms":72073,"concrete_test":"Recompute the two headline diagnostics for the same 20-year NeuralGCM runs (and the same ERA5/CMIP6 baselines) against references not used in training: (i) for extremes, use GPCP One-Degree Daily (already downloaded by the authors, coarsened to 2.8°) to compute Rx1day, 99.9th percentile, and the 24-hour rate frequency distribution; (ii) for the diurnal cycle, use an independent sub-daily product such as CMORPH or GSMaP (e.g., 3-hourly, coarsened to 2.8°) to compute the phase of maximum precipitation and diurnal/semi-diurnal harmonic amplitudes for the same summertime definition as Fig. 6. Then compare the NeuralGCM MAE values against the baselines. If NeuralGCM's advantage persists (e.g., MAE still below ERA5 and below the best CMIP6 baseline) the central claim is robust; if the advantage disappears or reverses, the headline claim is an artifact of evaluating on the training target.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim enumerates mean state, extremes, and diurnal cycle as precipitation aspects in which NeuralGCM 'substantially surpasses' GCMs and ERA5. Mean-state support is partly independent: Fig. S15 repeats the Fig. 4 MAE comparison against GPCP, a dataset not used in training, and the forecast section includes a GPCP-based CRPS/RMSB check (Fig. S5). But the extreme-precipitation evidence (Fig. 5 frequency distributions, Rx1day, 99.9th percentile; Figs. S22, S23) and the diurnal-cycle evidence (Fig. 6; Figs. S16–S18) are computed exclusively against IMERG, which is the very precipitation product used in the CRPS training loss. The model was optimized to reduce CRPS against IMERG precipitation; therefore its IMERG scores are not an independent test of physical accuracy. This is not a circularity in the loss itself—the climate diagnostic (e.g., Rx1day, diurnal phase) is a different functional than the CRPS training loss—but it is a form of training-target evaluation: a model optimized to match IMERG will naturally resemble IMERG more than models that were not. The paper's own Fig. S1 shows non-negligible IMERG–ERA5 differences at 2.8°, so IMERG-specific biases (e.g., overestimation of heavy precipitation, as cited in Ref. 45) could be learned and then reported as 'improvement.' The independent GPCP checks are reassuring for mean and forecast skill, but they do not cover extremes or diurnal phase/amplitude, which are the headline aspects most sensitive to observational product choice. Without an independent reference for these metrics, the central claim overstates what is demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a hybrid neural general circulation model (NeuralGCM) trained end-to-end on satellite-based precipitation observations from IMERG, together with ERA5 atmospheric fields. Precipitation is predicted by a small neural network, and evaporation is diagnosed from the column water budget, so that precipitation is consistent with the model's moisture tendencies. The authors report substantial improvements over CMIP6 models, ERA5, and the X-SHiELD global cloud-resolving model in simulated mean precipitation, extreme precipitation, and diurnal cycle, and also report that a 50-member NeuralGCM ensemble outperforms the ECMWF ensemble in precipitation forecast skill out to 15 days.","tokens_in":25704,"tokens_out":4593,"duration_ms":43021,"significance":"If the claims hold, the paper is a significant advance: it demonstrates that a differentiable GCM can be trained directly on an observational precipitation product, that the resulting model remains stable in multi-decadal simulations, and that it produces competitive or superior precipitation forecasts at 2.8 degrees resolution. The work is supported by extensive evaluation on two datasets (IMERG and GPCP), open code and data availability, and a proof-of-concept large ensemble for extreme-precipitation sensitivity. The central claims, however, are weakened by the fact that the headline metrics for extremes and the diurnal cycle are evaluated only against IMERG, the very dataset used in the training loss, and by the paper's own documentation of unrealistic sub-6-hourly precipitation variability and unrealistic instantaneous evaporation. These issues are acknowledged in the supplementary material but need to be addressed in the main text and through additional independent evaluation.","major_comments":[{"comment":"The extreme-precipitation and diurnal-cycle metrics are computed exclusively against IMERG, which is the precipitation product used in the CRPS training loss. The model was optimized to match IMERG precipitation, so its IMERG-based scores are not an independent test of physical accuracy. The paper does provide independent GPCP-based evaluation for mean precipitation (Fig. S15) and forecast skill (Fig. S5), but not for extremes or diurnal phase/amplitude. Since the central claim in the Discussion is that the model 'substantially surpasses' GCMs and ERA5 in 'mean state, extremes, and the diurnal cycle,' the authors should either add GPCP-based evaluation of extremes (e.g., Rx1day or 99.9th percentile, which are computable from GPCP daily data) and diurnal cycle if possible, or explicitly limit the claim to IMERG-defined skill and discuss the risk that IMERG-specific biases (e.g., overestimation of heavy precipitation) are being learned and reported as improvement.","section":"Results, 'Precipitation extremes and precipitation rate distribution' and 'Diurnal cycle of precipitation'; Figs."},{"comment":"The paper's own text states that 'the diurnal cycle in NeuralGCM exhibits unrealistic features, with certain times of day experiencing significantly more precipitation than others (Figs. 6e-g, S7), likely due to the model being optimized for 6-hourly precipitation accumulation,' and recommends against using this configuration at frequencies higher than 6-hourly. Yet the abstract and Discussion claim 'a more accurate diurnal cycle' and 'substantially surpasses ... the diurnal cycle.' The MAE-based harmonic metrics (Fig. 6a-d) may still favor NeuralGCM, but the model's diurnal precipitation shape is visibly unrealistic at sub-6-hour scales (Fig. 6e-g). The claim should be qualified to distinguish harmonic phase/amplitude skill from the model's unphysical high-frequency variability, and the implications for the 'accurate diurnal cycle' statement should be discussed in the main text.","section":"Results, 'Diurnal cycle of precipitation'; Figs."},{"comment":"The diagnosed evaporation field E = NN_precip(X) - (1/g)∫∑(dq/dt)_NNtend ps dσ has the potential to be unphysical if the learned moisture tendencies from the physics network carry ERA5 biases. The supplementary material itself shows unrealistic instantaneous evaporation fields for this configuration (Fig. S21). If the precipitation improvement is achieved by compensating, unphysical evaporation errors, the attribution of the improvement to learning from observations is weakened. The authors should validate the diagnosed evaporation more directly, or at least analyze whether the water-budget closure relies on cancellation between the precipitation network and the moisture tendencies. As written, the paper's central mechanism—training precipitation while enforcing column conservation—is not fully supported for the diagnosed evaporation component.","section":"Methods, Eq. (1); Supplementary Eq. (S4); Supplementary Fig. S21"},{"comment":"The statement that 'the stable model was still obtained by training several models with varying random seeds and choosing the most stable one' is a selection procedure that could bias the reported stability. The 37-initial-condition and 732-run stability experiments are impressive and partially mitigate this concern, but the selection procedure and its potential effect on the stability and climate statistics should be described in the Methods and discussed. Without this, the 'remains stable for decadal simulations' claim is presented more strongly than the training procedure warrants.","section":"Discussion, 'Our work retains some noteworthy limitations'"}],"minor_comments":[{"comment":"The caption contains a typo: 'NeruralGCM-evap' should be 'NeuralGCM-evap'.","section":"Supplementary Figure S24"},{"comment":"The phrase 'Mean absoulute error' should be 'Mean absolute error'.","section":"Supplementary Figure S22"},{"comment":"The description says the precipitation network predicts precipitation 'at 1-hour intervals,' but the training loss is on 6-hour accumulated precipitation. Clarify whether the network is trained on instantaneous 1-hour rates or on 6-hourly accumulations, since this distinction is relevant to the sub-6-hour precipitation artifacts.","section":"Methods, 'Neural network for predicting precipitation'"},{"comment":"The sentence 'We find similar conclusions when studying the 99.9th percentile (Fig. S22)' would benefit from the numerical MAE reduction for the 99.9th percentile, to match the quantitative context given for Rx1day.","section":"Results, 'Precipitation extremes and precipitation rate distribution'"},{"comment":"The phrase 'stochastic training approach of [29]' is slightly ambiguous; it could be read as stochastic gradient descent rather than the rollout-length randomization described later. Consider rewording to 'the training approach of [29], which progressively increases rollout length.'","section":"Introduction, third paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate for publication if the evaluation gaps are addressed. The central results are likely to be of broad interest, and the authors are commendably transparent about limitations in the supplementary material. However, the main-text claims currently outrun the independent evidence for extremes and diurnal cycle, and the diagnosed-evaporation artifacts raise a physical-consistency question that the authors should confront explicitly. A revision that adds GPCP-based extreme/diurnal evaluation (or clearly re-scopes the claims) and discusses the evaporation and seed-selection issues would make the paper substantially more convincing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real: a hybrid GCM trained online against IMERG precipitation, with evaporation diagnosed from the column water budget, that runs 20-year simulations stably and beats CMIP6, ERA5, X-SHiELD, and ECMWF ENS on the precipitation metrics tested. That's a genuine advance in the NeuralGCM program, not a brand-new framework—the differentiable core and online training recipe are inherited from Kochkov et al. 2024. What's new is the precipitation network, the water-budget closure that makes P a state variable with E diagnosed, and the demonstration that an observational target is usable for end-to-end training of a hybrid GCM.\n\nCredit where it's due: the paper ships code and model outputs, compares against a real operational ensemble on proper probabilistic scores, and includes GPCP as an out-of-training reference for mean precipitation (Fig S15) and forecast skill (Fig S5). The mean-state and forecast claims hold up under that independent check. The limitations section is unusually candid—seed-dependent stability, unrealistic instantaneous evaporation (Fig S21), sub-6-hour oscillations, and an explicit recommendation against using this configuration at frequencies higher than 6-hourly. That honesty buys goodwill.\n\nSoft spots, in proportion. The stress-test concern is correct: extremes (Figs 5, S22, S23) and diurnal cycle (Fig 6, S16–S18) are evaluated only against IMERG, the training target. The model was optimized to match IMERG's CRPS, so those scores are not independent tests of physical accuracy. This doesn't sink the paper—GPCP covers mean and forecast, and the authors flag the diurnal artifacts themselves—but 'substantially surpasses' is established for mean and forecast, and only suggested for extremes and diurnal cycle until a GPCP or gauge-based check appears. Second, the most stable seed was selected post hoc; all-seed statistics aren't reported, so the stability claim is conditional. Third, down-weighting ERA5 specific humidity by 100x means the model is pulled away from reanalysis moisture; the diagnosed E in Fig S21 shows the cost. One point in the paper's favor: the alternative NeuralGCM-evap configuration, which predicts E and diagnoses P, actually does better on most precipitation metrics (Figs S25–S27), suggesting the skill is not an artifact of the specific P/E split.\n\nWho this is for: anyone working on ML parameterization or hybrid GCMs, and anyone who cares about precipitation fidelity in climate models. It deserves a serious referee. The fixes are straightforward—report all-seed stats, add GPCP numbers for extremes and diurnal phase, state the valid time scales—and I'd expect a conditional accept after those additions.","headline":"A real advance in the NeuralGCM program: training a hybrid GCM on satellite precipitation works, the forecast and mean-state claims survive an independent GPCP check, but the extremes and diurnal-cycle wins are only shown against the training target.","tokens_in":26271,"tokens_out":2201,"would_cite":true,"duration_ms":20708,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a hybrid general circulation model directly on satellite-based precipitation observations can make its precipitation climatology, extremes, and diurnal cycle substantially more realistic than conventional GCMs, reanalysis, and a…","keywords":["NeuralGCM","precipitation parameterization","hybrid machine learning climate model","differentiable dynamical core","IMERG satellite precipitation","diurnal cycle of precipitation","extreme precipitation","ensemble weather forecasting"],"falsifier":"Examine the diagnosed evaporation fields in a long integration and compare them to direct flux observations or to a high-resolution water-budget product; if the fields show grid-scale artifacts that track precipitation errors, or if the model's precipitation skill collapses when the learned moisture tendencies are replaced by a trusted reanalysis budget, the water-budget closure is doing compensating work and the central claim is falsified.","tokens_in":25148,"feed_emoji":"🌧️","tokens_out":7940,"duration_ms":66947,"temperature":0.7,"pith_summary":"General circulation models have persistent, societally important precipitation errors: mean biases are comparable in size to projected changes, extremes are poorly distributed, and the daily cycle peaks too early. This paper tries to remove those errors by training a hybrid atmospheric model directly on satellite-based rain observations instead of relying only on reanalysis or high-resolution simulation targets. The model is a differentiable dynamical core with a learned physics parameterization, augmented by a learned precipitation network; evaporation is diagnosed from the column water budget so the hydrologic cycle stays closed. Across 20-year simulations, the authors report that the model cuts mean precipitation error by about 40% relative to a 37-model CMIP6 AMIP ensemble, improves extreme and diurnal precipitation, and outperforms a 50-member operational ensemble in medium-range precipitation forecasts. The upshot is that observation-based training can directly improve a quantity that conventional parameterization development has struggled to fix for decades.","feed_headline":"Satellite rain data trains a climate model that beats ECMWF","feed_subtitle":"NeuralGCM trained on IMERG rainfall cuts mean precipitation bias by 40% and captures the diurnal cycle.","key_machinery":"The load-bearing object is the column water budget closure, Eq. (1): $P - E = \\frac{1}{g}\\int\\sum_i \\left(\\frac{dq}{dt}\\right)^{\\mathrm{NNtend}}_i p_s\\, d\\sigma$, with $P$ the learned precipitation rate, $E$ the diagnosed evaporation, and the sum running over water-vapor, cloud-ice, and cloud-liquid tendencies from the learned physics network. A small Encode-Process-Decode network with a ReLU output predicts non-negative precipitation from the atmospheric column plus static embeddings, while the differentiable dynamical core and tendency network evolve the state. Enforcing this budget couples the precipitation loss to the physics network's moisture tendencies, which shifts the model's precipitable-water distribution toward observations and is what allows an observable (rain) to train the unresolved physics without breaking water conservation.","core_discovery":"NeuralGCM was originally a differentiable dynamical core coupled to a learned physical-tendency network trained on ERA5. This paper retrains it end-to-end while adding a compact precipitation network that outputs the hourly precipitation rate, with evaporation then diagnosed as the residual of the column water budget, $P - E = \\frac{1}{g}\\int\\sum_i \\left(\\frac{dq}{dt}\\right)^{\\mathrm{NNtend}}_i p_s\\, d\\sigma$. The precipitation network is trained against 6-hourly IMERG accumulations, while atmospheric state, evaporation, and other fields are trained against ERA5; the specific-humidity loss is down-weighted because ERA5 humidity itself is known to deviate from observations. In 20-year, SST-forced simulations started from 37 initial conditions, every run remains stable, and the model reproduces the IMERG precipitation-rate frequency distribution, annual-maximum daily precipitation (Rx1day), and diurnal phase more closely than ERA5, CMIP6 AMIP and historical runs, and GFDL's X-SHiELD cloud-resolving model. In forecasting, a 50-member NeuralGCM ensemble improves on the 50-member ECMWF ensemble for all 15 forecast days in CRPS, root-mean-square bias, spread-skill ratio, and Brier score at the 0.95 quantile, with the result holding against GPCP, a dataset not used in training.","pith_inferences":["The two configurations explored here have complementary failure modes—predicted precipitation with diagnosed evaporation creates unrealistic sub-6-hour oscillations, while predicted evaporation with diagnosed precipitation produces negative rain. Enforcing both fields as learned outputs with a hard water-budget constraint and non-negativity is the natural next step and might remove both artifacts.","The diagnosed-evaporation model shows unrealistic instantaneous evaporation fields (Fig. S21), so a key test is whether the precipitation skill survives when the learned moisture tendencies are replaced by a physically trusted budget; if it does not, the improvement may be partly owed to compensating errors rather than to a better representation of convection.","Because the model was trained jointly on ERA5 and IMERG with the humidity loss deliberately weakened, the diurnal-cycle improvement is likely attributable mainly to the IMERG precipitation loss; training identical models with and without the precipitation term would make that causal attribution quantitative.","The reported global sensitivity of annual-maximum precipitation, $4.2\\%\\,\\mathrm{K}^{-1}$, is lower than the $5$\\textendash$10\\%\\,\\mathrm{K}^{-1}$ range often quoted, and the authors note resolution suppresses the most extreme tail; running the same training recipe at higher resolution is a direct way to test whether the sensitivity rises toward that range."],"forward_implications":["Training on satellite precipitation removes the need to inherit precipitation biases from reanalysis or high-resolution simulation targets, which is the main route by which current hybrid models pick up errors.","A coarse 2.8° model can beat the operational ensemble in precipitation forecast skill, implying that resolution is not the only binding constraint for rain prediction and that learned physics can compensate.","Observation-trained hybrid models can be stable over 20-year climate integrations when the loss is tuned and stability is screened across random seeds.","The computational speed (about 1200 simulated years per TPU-day) makes large ensembles practical, as demonstrated by 732 twenty-two-year runs used to estimate the sensitivity of annual-maximum precipitation to temperature.","The same end-to-end recipe can be applied to any observable whose loss can be computed from model output, such as clouds or radiative fluxes."],"supporting_citations":[{"why":"Supplies the differentiable NeuralGCM dynamical core, the learned physics parameterization, and the online stochastic training procedure this paper modifies.","marker":"[29]"},{"why":"IMERG V07 final satellite-gauge precipitation data are the training target for the precipitation network and the primary evaluation benchmark.","marker":"[34]"},{"why":"ERA5 provides the atmospheric-state, evaporation, and surface training targets, and serves as the main reanalysis baseline.","marker":"[30]"},{"why":"GPCP one-degree daily precipitation is the independent observation-based benchmark not used in training.","marker":"[37]"},{"why":"WeatherBench2 supplies the evaluation code and procedure for scoring the 50-member forecast ensemble.","marker":"[50]"},{"why":"X-SHiELD is the global cloud-resolving model whose precipitation statistics NeuralGCM is compared against.","marker":"[51]"},{"why":"The continuous ranked probability score (CRPS) is both the stochastic training loss and a main forecast skill metric.","marker":"[33]"}],"fun_headline_variants":["Satellite-trained GCM beats ECMWF, cuts rain bias 40%","NeuralGCM tuned on IMERG rain beats ECMWF forecasts","Satellite-trained model slashes rain bias 40%, beats ECMWF","Climate model trained on satellite rain outperforms ECMWF ensemble","Rain-trained AI climate model beats ECMWF 15-day forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the learned physics network's moisture tendencies are accurate enough that the evaporation diagnosed from the column water budget is physically meaningful; if those tendencies merely compensate for the mismatch between ERA5 and IMERG, the precipitation gains could come with unphysical evaporation errors.","fun_headline_variants_meta":{"raw":{"variants":["Satellite-trained GCM beats ECMWF, cuts rain bias 40%","NeuralGCM tuned on IMERG rain beats ECMWF forecasts","Satellite-trained model slashes rain bias 40%, beats ECMWF","Climate model trained on satellite rain outperforms ECMWF ensemble","Rain-trained AI climate model beats ECMWF 15-day forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000902,"raw_usage":{"total_tokens":3889,"prompt_tokens":962,"completion_tokens":2927,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":2836}},"tokens_in":578,"tokens_out":2927,"duration_ms":20817,"temperature":1.0,"reasoning_tokens":2836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:24:45.476430+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Examine the diagnosed evaporation fields in a long integration and compare them to direct flux observations or to a high-resolution water-budget product; if the fields show grid-scale artifacts that track precipitation errors, or if the model's precipitation skill collapses when the learned moisture tendencies are replaced by a trusted reanalysis budget, the water-budget closure is doing compensating work and the central claim is falsified.","supporting_citations":[],"review_version":1}