{"id":"fadbfe23-60ed-4d2e-80b6-c7843949f549","arxiv_id":"2510.20769","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A transformer-based ensemble weather model trained on ERA5 with IMERG precipitation labels beats GEFS precipitation skill at short-to-medium lead times in January and July 2023.","lead":"The paper introduces CSU-PCAST, a deep-learning ensemble system that predicts global 6-hour precipitation out to 15 days from GFS initial conditions, and reports that it beats the operational GEFS ensemble in most tested metrics for January and July 2023. A generalist might read it because it is a concrete attempt to replace physics-based medium-range precipitation ensembles with a fast learned model, though the evaluation is limited and uses the same satellite product for tr","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"IMERG serves as both training label and verification reference; the Appendix A.2 control shows that when the label/reference is switched to ERA5 with GFS initialization, the CSU-PCAST advantage over GEFS disappears, suggesting the reported skill may be target-matching rather than genuine precipitati","rationale":"The reader's weakest_assumption identifies the same issue, and I agree. The reason I regard this as the most load-bearing concern is the existence of the ERA5-label control in A.2. It is not hypothetical: under the operational GFS initialization, changing only the label/verification product from IMERG to ERA5 flips the result from 'better than GEFS' to 'marginally worse than GEFS.' That is exactly the signature of target-product matching. A model that genuinely improves precipitation representation ought to retain at least some advantage over a dynamical baseline when verified against a different but reasonable precipitation reference; A.2 shows it does not. This does not make the paper worthless: the model may indeed be the best available IMERG-reproduction system, and IMERG is arguably a better global precipitation reference than ERA5. But the central claim as worded—'more skillful for precipitation than GEFS'—is broader than what is demonstrated. I therefore keep the reader's CONDITIONAL verdict rather than moving to REJECT; the issue can be settled by an independent verification. Secondary issues noted by the reader (two-month evaluation vs full-year abstract, no confidence intervals, GFS fine-tune selection on the test period, unsupported ECMWF claim) reinforce this but are not the single most load-bearing point.","tokens_in":17813,"tokens_out":8781,"duration_ms":79282,"concrete_test":"Recompute the Jan/Jul 2023 comparison of CSU-PCAST versus GEFS, verifying both against an independent gauge-based precipitation analysis (e.g., CPC Unified 0.25° or MSWEP) using the same CSI thresholds (0.1, 1, 10, 20 mm) and CRPS over lead times 0–15 days. If the advantage persists, the target-matching objection is weakened; if it shrinks or reverses, the central claim should be restated as skill at reproducing IMERG, and the verdict should be CONDITIONAL pending independent verification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 fixes IMERG Final as the training label, and Sections 2 and 4.3 verify all precipitation metrics against the same IMERG product. The model's objective explicitly minimizes CRPS and log1pMSE against IMERG at every grid point, while GEFS is an independent NWP system. The Appendix A.2 control sharpens this: when the precipitation label/reference is switched to ERA5 and the model is initialized from GFS, CSU-PCAST performs marginally worse than GEFS across all metrics (Figs. 9–11). Thus the main-operational advantage (IMERG labels, GFS initialization, IMERG verification) is not a robust property of the forecasting system but is tied to the identity of the training/verification product. The reported CSI/CRPS gains may therefore reflect the model's ability to reproduce IMERG's systematic retrieval biases (e.g., orographic overestimation, light-rain detection, IR-based estimates) rather than superior skill for true precipitation. The absence of an independent reference means the central claim is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CSU-PCAST, a dual-branch Swin Transformer framework for 15-day, 30-member ensemble precipitation forecasting. The model is trained on ERA5 atmospheric fields with IMERG precipitation as labels, initializes from operational GFS analyses at inference, and uses an autoregressive rollout with stochastic noise for ensemble diversity. The authors evaluate deterministic and probabilistic precipitation skill against GEFS for January and July 2023, reporting higher CSI, lower RMSE and CRPS, and improved Brier Scores. An appendix control experiment switches training labels and verification reference to ERA5 precipitation.","tokens_in":18022,"tokens_out":2961,"duration_ms":31421,"significance":"If the headline result held, CSU-PCAST would be a notable contribution: an ML ensemble at 0.25° that beats an operational NWP ensemble in precipitation skill. The architecture is well described, the evaluation is temporally out-of-sample, and the authors honestly report a control experiment in Appendix A.2. However, that same control experiment sharply limits the claim: when the training/verification reference is changed from IMERG to ERA5 and the model is initialized from GFS, CSU-PCAST becomes marginally worse than GEFS. Since the main evaluation uses IMERG as both training target and verification reference, the reported skill advantage may reflect target-product matching rather than a robust forecasting improvement. In addition, the abstract claims a full-year 2023 evaluation while the body evaluates only January and July, and no uncertainty quantification is provided for score differences.","major_comments":[{"comment":"The abstract states evaluation 'over the full year of 2023,' but Sections 2 and 4.3 describe evaluation only for January and July 2023 (two representative months). The conclusion also refers to 'both winter and summer cases.' The abstract's claim is not supported by the body. Either perform a full-year evaluation or revise the abstract to match the two-month evidence.","section":"Abstract; Section 2; Section 4.3"},{"comment":"The training objective (Eqs. 3–5) minimizes CRPS and log1pMSE against IMERG, and all main precipitation metrics are verified against the same IMERG product. GEFS, by contrast, is an independent NWP system. The control experiment in Appendix A.2 is the key test: with ERA5 precipitation as label and reference, and with GFS initialization, CSU-PCAST performs marginally worse than GEFS (Figs. 9–11). This indicates that the main reported advantage is tied to the identity of the training/verification product, not an intrinsic forecast-skill gain. The authors need to either evaluate against an independent reference (e.g., gauge-adjusted products or a second satellite product) or substantially weaken the central claim and discuss this circularity explicitly.","section":"Section 3.1; Section 2; Eqs. (3)–(5); Appendix A.2"},{"comment":"GEFS forecasts at native 0.5° resolution are bilinearly interpolated to 0.25° prior to verification, while CSU-PCAST is trained and evaluated at 0.25°. Bilinear interpolation smooths the GEFS precipitation field, which can inflate apparent CSU-PCAST advantages in categorical metrics such as CSI at high thresholds and in CRPS. The authors should quantify this effect, e.g., by also evaluating CSU-PCAST on the 0.5° grid after conservative remapping, or by reporting both native- and common-grid comparisons.","section":"Section 3.2"},{"comment":"All claims of 'consistently higher' or 'consistently lower' skill rely on point estimates over two months. No confidence intervals, significance tests, or block-bootstrap intervals are provided, despite strong spatial and temporal correlations in precipitation fields. The CSI differences in Figure 1 are often small (e.g., <0.05 at several lead times/thresholds), and the BS differences in Figure 4 cross zero at long lead times. The authors should add uncertainty quantification to support the strength of the claims.","section":"Section 2; Figures 1–4"},{"comment":"The introduction states that CSU-PCAST 'outperforms ECMWF and GEFS ensembles,' but the paper only evaluates against GEFS. No ECMWF ENS comparison is presented anywhere in the manuscript. This unsupported claim should be removed or qualified to 'GEFS' only.","section":"Section 1"}],"minor_comments":[{"comment":"Typos: 'NVIDA' in Section 2 and 'Cooperative Intitute' in Acknowledgments. The abstract says 'Precipitation Forecasting' but the title uses 'precipitation Forecasting'; use consistent capitalization.","section":"Abstract; Section 1"},{"comment":"'2-meter dewpoint temperature' is abbreviated 2D, but the standard abbreviation is 2Td or 2D; clarify. Also 'Total column water vapor' is TCWV; fine. Consider defining abbreviations in the table caption.","section":"Section 3.1, Table 1"},{"comment":"Figure panels are small and the text refers to 'POD' and 'FAR' in Section 2.1, but these metrics are not shown in the figures or explicitly defined in Section 4.3. Either add the figures or remove the reference.","section":"Figures 1, 2, 4"},{"comment":"The statement that 'precipitation inherently follows an autoregressive dependency' and therefore one-step training is sufficient is plausible but not demonstrated. The model is rolled out 60 steps without multi-step precipitation training; the paper should cite evidence or provide an ablation (e.g., fine-tuning with 2–3 steps) to support stability.","section":"Section 4.2.3"},{"comment":"The manuscript does not mention code or data availability. Given the reproducibility expectations for ML-based forecasting papers, please add a statement or repository link.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The Appendix A.2 control is the crux. The authors deserve credit for reporting it, but as written it directly undermines the headline claim: the advantage over GEFS disappears when the training/verification product is changed to ERA5 with GFS initialization. This is not an internal inconsistency, but it is a correctness-risk concern that cannot be dismissed without an independent reference or a much more cautious framing. The abstract's 'full year 2023' mismatch and the absence of ECMWF comparison further overstate the evidence. I recommend major revision, not rejection, because the architecture and evaluation framework are sound and the issues can be addressed with additional analysis or revised claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a real attempt at a hard problem. It presents a FuXi-style Swin transformer with a dual-branch decoder that separates precipitation from other variables, stochastic noise via conditional layer norm/FiLM, trained on ERA5 with IMERG labels, initialized from GFS, and produces 30-member 15-day forecasts. Evaluation is a genuine temporally out-of-sample comparison against GEFS for Jan and Jul 2023, and the model wins on CSI, CRPS, RMSE, and BS when verified against IMERG.\n\nWhat the paper does well: the architecture is clearly described, the dual-branch idea is sensible, and the noise injection is a plausible way to generate ensemble members. The authors also include a control experiment (Appendix A.2) and state plainly that both ensembles are underdispersive. That honesty is worth something.\n\nThe soft spot that matters: IMERG serves as both training target and verification reference. The loss explicitly minimizes CRPS and log1pMSE against IMERG; GEFS does not learn from IMERG. The abstract claims a full-year 2023 evaluation, but the body shows only January and July. More importantly, Appendix A.2 undercuts the headline: when the training label and verification reference are switched to ERA5 and the model is initialized with GFS, CSU-PCAST becomes marginally worse than GEFS across all metrics. That pattern is consistent with target-matching rather than generalizable skill. The authors present the ERA5 experiment as showing ERA5 is a worse training target, but the implication for their main result is the opposite: the reported advantage is tied to the identity of the reference product.\n\nOther issues: no confidence intervals or significance tests on the score differences; GEFS is bilinearly interpolated from 0.5 to 0.25 degrees; the GFS fine-tuned variant was rejected based on observed CSI, possibly on the same test period; and the introduction claims outperformance against ECMWF without evaluating ECMWF.\n\nWho this is for: the ML-for-weather community, and anyone doing precipitation verification. The paper deserves a serious referee; it's not a desk reject. But the referee should require a full-year evaluation, an independent reference (gauges, radar, or a different satellite product), uncertainty quantification, and code or weights. Without those, the central claim should be treated as a hypothesis.\n\nMy recommendation: engage with it. Send it to review, but push the authors to confront the A.2 control directly. If the IMERG advantage survives against an independent reference, the paper becomes much more convincing. As it stands, the headline result is underdetermined.","headline":"A real architecture and a genuine two-month evaluation, but the A.2 control shows the IMERG advantage is likely target-matching, not generalizable skill.","tokens_in":18600,"tokens_out":3594,"would_cite":false,"duration_ms":34986,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learning ensemble trained on reanalysis and satellite precipitation can beat the operational GEFS at forecasting rain up to 15 days out.","keywords":["ensemble precipitation forecasting","deep learning weather prediction","Swin Transformer","IMERG","GEFS","medium-range forecast","CRPS","autoregressive model"],"falsifier":"Run CSU-PCAST and GEFS forecasts for a full year and verify against an independent precipitation analysis not used in training, such as a dense gauge network over the United States or a reanalysis-independent satellite product. If the CSI and CRPS advantages shrink or reverse, the reported skill is likely an artifact of matching IMERG's error structure rather than a genuine improvement in precipitation forecasting.","tokens_in":17602,"feed_emoji":"🌧️","tokens_out":2709,"duration_ms":27524,"temperature":0.7,"pith_summary":"The paper presents CSU-PCAST, a 30-member ensemble precipitation forecasting system built on a Swin Transformer with stochastic noise conditioning, and claims it outperforms the operational GEFS ensemble in medium-range precipitation forecasts. Using 21 years of ERA5 atmospheric data with IMERG satellite precipitation as labels, the model is initialized from GFS analyses and predicts 6-hourly rainfall out to 15 days. Evaluated on January and July 2023, it reports higher Critical Success Index at thresholds from 0.1 to 20 mm, lower RMSE, lower CRPS, and better Brier Scores across lead times. If correct, this shows that a data-driven ensemble can match or exceed a traditional numerical ensemble for a high-impact, hard-to-predict variable like precipitation.","feed_headline":"ML ensemble beats GEFS on rain forecasts to 15 days","feed_subtitle":"A Swin Transformer with noise-conditioned training improves CSI, CRPS, and Brier scores across all thresholds.","key_machinery":"The central architecture is a patch-based Swin Transformer V2 backbone whose feature representations are modulated by conditional layer normalization (FiLM) driven by both time embeddings and injected Gaussian noise. This noise conditioning generates ensemble spread without perturbing initial conditions. A dual-branch decoder separates the total-precipitation channel from the 57 other atmospheric and surface variables, allowing the precipitation branch to be trained with a precipitation-specific loss that emphasizes rain above 5 and 10 mm via intensity weighting. The model is trained autoregressively, first on non-precipitation variables, then fine-tuned for precipitation, and during inferen","core_discovery":"CSU-PCAST consistently outperforms GEFS when both are initialized from GFS analyses and verified against IMERG. In January and July 2023, the model achieves higher CSI at 0.1, 1, 5, 10, and 20 mm thresholds, lower precipitation RMSE, and lower CRPS across the full 15-day window. The largest gains appear at moderate-to-heavy rainfall and at longer lead times, where GEFS skill degrades quickly. Brier Score differences are also favorable through about day 10, and a Typhoon Sanba case shows improved spatial structure and exceedance probabilities. The paper argues that a purpose-built precipitation branch, trained with a combined CRPS and weighted log1p MSE loss, is key to this advantage.","pith_inferences":["The evaluation covers only two months of one year, so the claim of general 15-day superiority is an extrapolation; testing across all seasons would clarify whether the advantage holds year-round.","Because the model is trained to match IMERG, the comparison against GEFS may partly measure how well each system reproduces IMERG's retrieval biases; verification against independent gauge or radar data would test this.","The intensity-weighted log1p loss could be transferred to other extreme-event forecasting problems, such as heat waves or ocean wave heights, where rare large values need emphasis.","If the model's spread is produced purely by noise conditioning, it may be possible to tune spread on the fly without retraining, enabling cheap calibration for downstream users."],"forward_implications":["If the claimed skill generalizes, data-driven ensembles could provide operationally useful precipitation guidance at a fraction of the computational cost of traditional ensemble NWP.","The noise-conditioning approach offers a lightweight alternative to diffusion or perturbed-initial-condition ensembles for generating spread.","The dual-branch decoder suggests a general recipe for handling heavy-tailed variables: train a shared backbone for smooth variables, then attach a specialized branch with a tailored loss.","The reported improvements at heavier thresholds and longer lead times are exactly where operational GEFS is weakest, so CSU-PCAST could complement existing guidance.","Consistent with the paper's own statement, both systems remain underdispersive, so the reliability gain is relative, not absolute."],"fun_headline_variants":["AI rain model beats GEFS on CSI, CRPS in 2023","ML ensemble improves rain forecasts vs GEFS to day 15","Transformer rain model outperforms GEFS in 2023 test","Rain forecasting AI tops GEFS on skill scores"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"IMERG Final (version 07) precipitation is treated as the ground truth for both training labels and verification, so reported skill is skill at reproducing IMERG, which the model was trained to match, not necessarily skill at reproducing true precipitation.","fun_headline_variants_meta":{"raw":{"variants":["AI rain model beats GEFS on CSI, CRPS in 2023","ML ensemble improves rain forecasts vs GEFS to day 15","Transformer rain model outperforms GEFS in 2023 test","Rain forecasting AI tops GEFS on skill scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001051,"raw_usage":{"total_tokens":4300,"prompt_tokens":845,"completion_tokens":3455,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":3380}},"tokens_in":589,"tokens_out":3455,"duration_ms":20555,"temperature":1.0,"reasoning_tokens":3380,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:22:31.639649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CSU-PCAST and GEFS forecasts for a full year and verify against an independent precipitation analysis not used in training, such as a dense gauge network over the United States or a reanalysis-independent satellite product. If the CSI and CRPS advantages shrink or reverse, the reported skill is likely an artifact of matching IMERG's error structure rather than a genuine improvement in precipitation forecasting.","supporting_citations":[],"review_version":1}