{"id":"8739187e-25cc-46d3-9e4e-5a2ddef83564","arxiv_id":"2506.11698","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A coordinate-based diffusion model fuses multi-source precipitation records and corrects biases in unseen operational forecasts.","lead":"PRIMER is a generative AI system that merges satellite, reanalysis, and gauge-style precipitation data into a single high-resolution estimate. It corrects biases in existing weather products, downscales coarse forecasts, and adapts to new forecast models without retraining.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation 'ground truth' is the same gridded gauge-satellite product used for fine-tuning, so PRIMER's claimed bias corrections may amount to fitting that product's own biases.","rationale":"The reader's weakest assumption—that the 'gauge observations' are not independent rain gauges but a gridded gauge-satellite merged product—is exactly the most load-bearing concern for this paper. The paper's method itself (coordinate-based diffusion, two-stage training, Bayesian posterior sampling) is internally plausible and the zero-shot operational forecast demonstration is interesting, but the headline claim of statistically significant error reductions depends entirely on the validity of the evaluation truth. Section 4.6 explicitly states that both fine-tuning targets and evaluation ground truth are grid cells from Shen et al. (2014), so the concern is not speculation but documented in the manuscript. The consequence is concrete: if the merged product's bias field correlates with the bias fields of ERA5 or IMERG, PRIMER can reduce apparent error by learning that product's systematic mapping, especially since its Stage 2 training uses the same product. The proposed test with raw station observations is the definitive check because it removes the circularity. If the independent evaluation confirms the reported improvements, the method's central claim is supported; if not, the paper needs major re-evaluation. I therefore keep the reader's CONDITIONAL verdict unchanged, with the raw-station validation as the explicit condition for acceptance. Other concerns (extreme-event selection, lack of code/data, absence of comparison to existing merged products) are secondary and do not change this verdict.","tokens_in":62,"tokens_out":2246,"duration_ms":42512,"concrete_test":"Obtain raw hourly automatic weather station (AWS) observations from CMA for a subset of the 150 test events (at minimum, the three case-study events and 20 randomly sampled test timestamps), without gridding or merging. For each event, evaluate PRIMER's posterior ensemble mean and spread at the exact station coordinates against these raw station values, using the same MAE and CRPS metrics as in the paper. If the reported improvements over ERA5, IMERG, and baseline posterior distributions largely persist, the concern is resolved. If the improvements shrink substantially or reverse, the error reductions reported in Figs. 4–6 are artifacts of evaluating against the same gauge-satellite merged product used for fine-tuning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that PRIMER corrects biases using gauge observations as ground truth—rests on the assumption that the evaluation targets are independent, trustworthy point measurements. Section 4.6 shows this is not the case. Fine-tuning uses grid cells from the Shen et al. (2014) gauge-satellite merged product with at least one assimilated AWS; evaluation uses cells from the same product with at least four AWS. This product is not raw gauge data but a gridded analysis that already blends satellite retrievals with gauge information. Consequently, Stage 2 training and the reported test metrics share the same interpolation/merging methodology and likely share systematic biases. If the merged product's bias field is correlated with IMERG or ERA5 (e.g., through its satellite component), PRIMER's 'improvement' may reflect learning to reproduce that product's gridded values rather than correcting real precipitation error. The phrase 'approximately 1,000 independent rain gauges' in the Introduction is therefore misleading: these are amalgamated grid-cell values, not independent stations. The test-set selection (Appendix C.2) further compounds this by choosing timestamps based on the same product's station intensities, potentially favoring events where the product's bias is systematic. Unless PRIMER is validated against raw station observations not used in any way during training or fine-tuning, the headline quantitative claims are not substantiated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PRIMER, a coordinate-based diffusion model for fusing precipitation records from gridded reanalysis (ERA5), satellite retrievals (IMERG), and sparse gauge-type observations. The method trains in two stages: first, it learns separate priors for ERA5 and IMERG; second, it fine-tunes on a merged gauge-satellite product to obtain an updated prior P*(x). Posterior sampling with this prior is then used for bias correction, downscaling, and zero-shot correction of unseen operational forecasts (HRES). The authors report improved MAE, CRPS, intensity distributions, and spatial coherence against what they call gauge observations, based on 150 selected precipitation events from 2016.","tokens_in":27292,"tokens_out":5805,"duration_ms":58910,"significance":"If the central claims are substantiated, PRIMER would be a valuable contribution: the coordinate-based formulation is a principled way to handle irregular gauge locations without destructive interpolation, the two-stage training idea is plausible, and the zero-shot generalization to HRES is an interesting demonstration. The paper also provides useful methodological details, including pseudo-code, a spectral diagnostic, and a theoretical justification in the supplement. The main weakness is that the evaluation target is not independent of the fine-tuning data: Section 4.6 states that both fine-tuning and evaluation use grid cells from the same Shen et al. (2014) gauge-satellite merged product, so the quantitative improvements in Figs. 4 and 5 may partly reflect fitting the product's own biases. The stress-test concern therefore lands, and the current evidence does not support the abstract's claim of statistically significant corrections against independent gauge observations.","major_comments":[{"comment":"The evaluation target is not independent of the fine-tuning data. The text states that fine-tuning uses grid cells from the Shen et al. (2014) gauge-satellite merged product with at least one assimilated AWS, and that evaluation uses cells from the same product with at least four AWS. Because this product is a gridded analysis that already blends satellite retrievals with gauge information, the updated prior P*(x) is calibrated to the same analysis used for scoring. The temporal split (2015/2017 training versus 2016 testing) does not remove shared product-level biases, so the improved MAE and CRPS in Figs. 4 and 5 may reflect reproducing the merged product's gridded values rather than correcting real precipitation error. The authors should either validate against raw station observations that are not used in training or in the construction of the merged product, or explicitly reframe all headline results as bias correction relative to the Shen et al. analysis rather than to gauges.","section":"Section 4.6 and Fig. C5"},{"comment":"The statement in the Introduction that evaluations use 'approximately 1,000 independent rain gauges' is misleading given Section 4.6: the evaluation is performed at grid cells of a merged product, not at independent station records. In addition, the 150 test events in Appendix C.2 are selected from the same product's station intensities, specifically the 100 timestamps with the highest individual station intensities and the 50 with the highest average intensity. This concentrates the test set on extreme events and may overstate improvements that are particular to heavy-rain conditions. Please report performance on a random or complete temporal sample from 2016, and quantify how the event-selection criteria affect the headline error reductions.","section":"Abstract, Section 1, and Appendix C.2"},{"comment":"The abstract claims 'statistically significant error reductions at most stations', but no significance test is described anywhere in the Methods or figure captions. The maps in Figs. 4 and 5 show mean differences, but there is no account of spatial or temporal dependence among stations and events, no confidence intervals, and no multiplicity control. Please add a paired significance test across stations or events (e.g., a bootstrap or permutation test), state the null hypothesis, and report the fraction of stations for which the improvement is significant after accounting for spatial correlation.","section":"Section 4.5 and Figs. 4-5"},{"comment":"The paper's baselines are its own posterior samples from PERA5(x) and PIMERG(x), not established bias-correction or fusion methods such as quantile mapping, CDF matching, or simple gauge interpolation. As a result, the current evaluation does not establish that PRIMER's advantage comes from the generative/Bayesian machinery rather than from the gauge information injected in Stage 2. A comparison against at least one standard non-generative fusion baseline would make the central claim more convincing. In the same vein, the theoretical justification in SI B.1 applies an ambient-diffusion bound that assumes noisy observations of a common Ptrue with known isotropic noise; ERA5 and IMERG are not independent noisy observations of the same target, and the merged product is not clean point truth, so the bound's relevance to the two-stage procedure should be stated more cautiously.","section":"Section 2.3 and SI B.1"}],"minor_comments":[{"comment":"The entity embedding e3 is described as 'gauge observations', but the data used for fine-tuning are grid cells from a merged gauge-satellite product; please use consistent terminology throughout.","section":"Section 4.2 and Section 4.6"},{"comment":"The sentence 'the ensemble-mean ΔMAE decreases from 0.46 mm/hr for P*(x|OERA5) to 0.14 mm/hr for PERA5(x|OERA5)' appears to have the sign or ordering reversed relative to the definition of ΔMAE in Section 2.3, where positive values indicate improvement of P* over the baseline; please clarify.","section":"Section 2.2"},{"comment":"Reporting that batch size and learning rate 'varied between... due to intermittent training interruptions' is not reproducible; please provide the final settings or a precise schedule for all reported experiments.","section":"Table B1 and SI B.7"},{"comment":"The SDEdit noise level τ is tuned on a single IMERG event (13 June 2016 at 23:00 UTC) and then applied to all experiments; a cross-event sensitivity check, or a rationale for why this choice transfers, would strengthen the results.","section":"Section 4.4"},{"comment":"The code availability statement says code will be released upon acceptance; an anonymized repository or model checkpoint available during review would help verify the implementation and reproducibility.","section":"Code availability"},{"comment":"Please fix the typo 'Acknowledgemenrs' in the Declarations section.","section":"Declarations"}],"recommendation":"major_revision","confidential_remarks":"The central evaluation-design issue is the main obstacle to publication. If the authors can validate against raw station observations that are not used in training or in the merged product, the paper could become acceptable. Otherwise, the quantitative claims in the abstract and Figs. 4-5 should be substantially softened and reframed as demonstrating consistency with the Shen et al. merged analysis. The paper fits the journal's scope; I do not see a novelty disclosure concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The method is worth taking seriously. Coordinate-based diffusion with source embeddings is a clean way to fuse gridded and irregular data, and the two-stage fine-tuning (pretrain on ERA5/IMERG, then tune on gauge-like data) is well motivated. The zero-shot demonstration on HRES forecasts is the most interesting part: it suggests the learned prior has some genuine transferability. The authors also put real effort into the theory (mollification, Wasserstein bound) and the architecture details are given in sufficient depth to reproduce, once code ships.\n\nThe soft spot is the evaluation, and it is not minor. The 'independent rain gauges' are grid cells from the Shen et al. (2014) gauge-satellite merged product—the same product used for Stage 2 fine-tuning. The temporal holdout (2015/2017 vs 2016) helps, and requiring at least four AWS per evaluation cell means those cells are better constrained by gauges. But the product's own systematic biases are likely present in both training and test periods, so PRIMER may simply be learning to reproduce that product's values. The claim in the abstract that PRIMER corrects biases 'using gauge observations as ground truth' is therefore overstated. The test set selection also targets extremes, which can inflate apparent improvement. And the paper says 'statistically significant' but gives no significance testing details.\n\nWhat's missing is the obvious control: comparison against the raw Shen product itself, or any standard merged product like MSWEP, at the same evaluation locations. If PRIMER beats those, the claim gets real support. Without that, the improvement over ERA5/IMERG is partly a comparison against uncalibrated baselines. The zero-shot HRES case is nice but it inherits the same evaluation flaw.\n\nThat said, this is not a fatal flaw. The method is novel, the architecture is sensible, and the authors are transparent about many limitations (oceanic coverage, China-only, computational constraints). They just don't address the circularity of their own evaluation, which is the one thing a referee has to force them to fix. If they release code and validate against raw station observations held out from all stages of training, this could be a genuinely useful contribution.\n\nI'd send it to peer review. The methodological novelty warrants it, and the evaluation problem is fixable. But I would make the revision requirements explicit: raw-gauge validation, comparison to existing merged products, and significance testing. This is a good paper for a hydrology or ML-for-science reading group, precisely because the evaluation pitfalls are instructive.","headline":"Solid method, shaky evaluation: the ground truth is the same merged gauge-satellite product used to fine-tune the model.","tokens_in":27805,"tokens_out":1922,"would_cite":false,"duration_ms":21773,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a coordinate-based diffusion prior trained on reanalysis, satellite retrievals, and sparse rain gauges produces precipitation estimates that outperform any single source, and that the same prior corrects biases in…","keywords":["precipitation data fusion","diffusion model","coordinate-based representation","Bayesian inference","bias correction","downscaling","zero-shot generalization","gauge–satellite merging"],"falsifier":"Re-evaluate PRIMER's posterior outputs at raw automatic weather stations that were excluded from the merged product and from all training data, in a region with dense independent gauges, and compare mean absolute error and CRPS against raw ERA5, raw IMERG, and the original merged product; if the advantage over the raw products shrinks or vanishes under this independent comparison, the claimed bias correction is largely an artifact of the merged ground truth.","tokens_in":26823,"feed_emoji":"🌧️","tokens_out":6765,"duration_ms":61548,"temperature":0.7,"pith_summary":"PRIMER is a generative model that treats precipitation as a continuous spatial field rather than a fixed grid, allowing dense satellite and reanalysis grids and sparse rain-gauge point observations to fuse at their native sampling structures without destructive interpolation. The paper's central claim is that a diffusion model trained in two stages—first on gridded ERA5 and IMERG data, then fine-tuned on gauge observations—yields a prior distribution that outperforms priors built from any single source. Conditioning this prior on an existing precipitation product through posterior sampling corrects biases, downscales coarse fields, and restores realistic high-frequency structure, with statistically significant error reductions at most evaluated stations and improved spatial coherence. The same prior also corrects biases in operational HRES forecasts it never saw during training, which the authors present as evidence of zero-shot generalization. A sympathetic reader should care because precipitation datasets currently disagree with each other by as much as the signal itself, and PRIMER offers a principled Bayesian route to combining imperfect sources into a single full-coverage estimate with quantified uncertainty.","feed_headline":"Fusing rain data beats any single source, and fixes unseen forecasts","feed_subtitle":"One diffusion prior from gauges, satellites, and reanalysis cuts rain errors and corrects unseen forecasts.","key_machinery":"The load-bearing object is the coordinate-based diffusion prior, a score-based generative model defined over the Hilbert space $L^2([0,1]^n \\to \\mathbb{R}^d)$, so precipitation is represented as a continuous function of spatial coordinates rather than a fixed-resolution tensor. The forward process mollifies white noise with a Gaussian kernel, implemented via Fourier transforms and a Wiener filter, to keep trajectories in $L^2$; the reverse process uses a hybrid network of SparseConvResBlocks and a U-Net, modulated by a source embedding that tells the model which dataset a sample came from. Two-stage training first learns large-scale climatology from ERA5 and IMERG, then fine-tunes with gauge observations through a weighted loss that preserves the gridded priors while grounding the manifold in local measurements. Posterior inference is carried out with inpainting and SDEdit, which balance the learned prior against the conditioning observations, and a noise-level parameter controls how much the posterior may deviate from the input field.","core_discovery":"The central discovery claimed is that an informative prior over precipitation fields can be constructed from imperfect, heterogeneous records by exploiting the spectral progression of diffusion models: as Gaussian noise gradually corrupts the target, the model learns low-frequency structure first and high-frequency details later. PRIMER therefore learns conditional source-specific priors for ERA5 and IMERG, then refines them with gauge observations to obtain a gauge-calibrated prior, using source embeddings and shared weights so each dataset contributes at its natural scale. Under posterior sampling conditioned on an existing product, the gauge-calibrated prior consistently lowers mean absolute error and continuous ranked probability score relative to the original ERA5 or IMERG fields, improves the tail of the intensity distribution, and makes the spatial lag-correlation structure closer to that of gauge observations. It also corrects biases in HRES forecasts without retraining. In the authors' framing, this turns the heterogeneity of imperfect data from a limitation into a strength.","pith_inferences":["Editorial inference: If the gauge-calibrated prior really captures shared climatology rather than the idiosyncrasies of the merged product, the same architecture could serve as a universal prior for many satellite, reanalysis, and forecast products beyond ERA5, IMERG, and HRES; this is a testable extension, not something the paper demonstrates.","Editorial inference: A fair stress test would be to fine-tune on one region's gauges and evaluate on another region with an independent national network, to see whether improvements come from generalizable physics or from memorizing the merged product's biases.","Editorial inference: The spectral view the authors adopt suggests a quantitative prediction: the largest gains from gauge fine-tuning should appear at wavelengths where gridded products are structurally deficient, so applying PRIMER to a dataset with different spectral error characteristics should produce a different improvement profile.","Editorial inference: If zero-shot correction of HRES is real, the framework could be deployed as an operational post-processing layer for numerical weather prediction ensembles, but the practical value would depend on the gauge network density available in the region of interest."],"forward_implications":["Posterior sampling from the calibrated prior yields ensemble precipitation fields with a spread, so users can obtain both a bias-corrected mean and an uncertainty estimate for risk assessment.","Because the prior is source-agnostic after fine-tuning, the same trained model can downscale coarse reanalysis, correct satellite retrieval biases, and correct operational deterministic forecasts, as demonstrated on HRES.","Additional gauge observations can be injected during sampling, so the framework can act as a lightweight data-assimilation tool without retraining.","The two-stage training recipe—large noisy gridded data for structure, sparse accurate data for local refinement—is transferable to other Earth-system variables with similar observation trade-offs.","If the learned prior reproduces reference climatology and spectra, it can be used to generate physically plausible precipitation realizations for studies of extremes and for evaluating risk under rare event scenarios."],"supporting_citations":[{"why":"Supplies the gauge-satellite merged precipitation analysis used for fine-tuning and as ground truth in evaluation.","marker":"[29]"},{"why":"Establishes the infinite-resolution diffusion framework with mollified states that PRIMER's coordinate-based architecture builds on.","marker":"[59]"},{"why":"Provides the score-based SDE formulation and reverse-time sampling used for generation and posterior inference.","marker":"[56]"},{"why":"Supplies the denoising diffusion training objective that PRIMER minimizes with a Hilbert-space loss.","marker":"[36]"},{"why":"Provides the IMERG satellite retrieval data used in Stage 1 pretraining.","marker":"[77]"},{"why":"Provides the ERA5 reanalysis data used in Stage 1 pretraining.","marker":"[78]"},{"why":"Supplies the SDEdit posterior-sampling procedure used for bias correction by adding controlled noise before denoising.","marker":"[73]"},{"why":"Provides the inpainting posterior-sampling method used to condition on partial gauge observations.","marker":"[70]"},{"why":"Supplies the operational HRES forecasts that PRIMER corrects zero-shot in the generalization experiment.","marker":"[50]"},{"why":"Motivates the personalization-style fine-tuning strategy that preserves gridded priors while adapting to gauge observations.","marker":"[68]"}],"fun_headline_variants":["AI merges rain data from gauges, satellites, models to fix errors","Diffusion model fuses precipitation sources, corrects unseen forecasts","Merging rain records with generative AI cuts errors and bias","One generative model blends rain data, fixes biases even on new forecasts","Rain data fusion: diffusion prior fixes forecasts it never saw"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gridded gauge–satellite merged product used as ground truth is itself trustworthy; because that product provides both the fine-tuning target and the evaluation target, any systematic bias in it would be treated as truth, and the paper's reported error reductions could reflect alignment to that bias rather than to real precipitation.","fun_headline_variants_meta":{"raw":{"variants":["AI merges rain data from gauges, satellites, models to fix errors","Diffusion model fuses precipitation sources, corrects unseen forecasts","Merging rain records with generative AI cuts errors and bias","One generative model blends rain data, fixes biases even on new forecasts","Rain data fusion: diffusion prior fixes forecasts it never saw"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000539,"raw_usage":{"total_tokens":2593,"prompt_tokens":961,"completion_tokens":1632,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":1543}},"tokens_in":577,"tokens_out":1632,"duration_ms":11601,"temperature":1.0,"reasoning_tokens":1543,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:35.269590+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-evaluate PRIMER's posterior outputs at raw automatic weather stations that were excluded from the merged product and from all training data, in a region with dense independent gauges, and compare mean absolute error and CRPS against raw ERA5, raw IMERG, and the original merged product; if the advantage over the raw products shrinks or vanishes under this independent comparison, the claimed bias correction is largely an artifact of the merged ground truth.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the gauge-satellite merged precipitation analysis used for fine-tuning and as ground truth in evaluation."},{"cited_title":"J.et al.Nasa global precipitation measurement (GPM) integrated multi-satellite retrievals for GPM (IMERG).Algorithm theoretical basis document (ATBD) version4, 30 (2015)","cited_arxiv_id":null,"evidence_quote":"Provides the IMERG satellite retrieval data used in Stage 1 pretraining."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ERA5 reanalysis data used in Stage 1 pretraining."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the inpainting posterior-sampling method used to condition on partial gauge observations."},{"cited_title":"URL https://www.ecmwf","cited_arxiv_id":null,"evidence_quote":"Supplies the operational HRES forecasts that PRIMER corrects zero-shot in the generalization experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the personalization-style fine-tuning strategy that preserves gridded priors while adapting to gauge observations."}],"review_version":1}