{"id":"ee8dc826-62e6-4a10-acf6-790952c95fb1","arxiv_id":"2507.04930","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"RainShift is a global benchmark showing that precipitation downscaling models lose up to 30% accuracy when applied to unseen Global South regions, and input quantile mapping recovers some of that loss.","lead":"RainShift is a new benchmark dataset that tests whether precipitation downscaling models trained in data-rich regions can generalize to data-poor regions in the Global South. The paper's measurements show substantial accuracy drops out-of-distribution, and that a simple quantile-mapping alignment recovers part of the gap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The controlled-setup premise is untested: regional ERA5/IMERG biases, plus IMERG-derived input clipping, may drive the reported OOD drops rather than geographic generalization.","rationale":"The reader identified the same weakest assumption, and I agree. This is the most load-bearing point because the entire framing of the benchmark as a 'geographic distribution shift' depends on the paired data sources being neutral across regions. If the drop is due to product artifacts, the benchmark still contains useful data and infrastructure, but its central scientific conclusion is not established. I considered other concerns: single-run/no-error-bar CRPS, the 30%/17% vs 37% inconsistency, the overstated QM claim, and the unavailability of code. These are real but secondary; they affect confidence in specific numbers and rankings, not the benchmark's foundational interpretation. The independent-target rerun is decisive and feasible for at least two regions; the regression diagnostic is a cheap complement. The paper deserves credit for a large, presumably useful dataset and for acknowledging some limitations, but a benchmark whose main message is 'OOD drops are large' should not be accepted without testing the product-bias confound. The verdict remains CONDITIONAL; I would not reject, because the concern is addressable and potentially refutable.","tokens_in":15622,"tokens_out":12053,"duration_ms":147549,"concrete_test":"Rerun a compact version of the benchmark with an independent high-resolution precipitation target for a subset of evaluation regions (e.g., Tibetan Plateau and Melanesia), using a gauge- or radar-based reference product (or MSWEP V2 as a merged alternative) at the coarsest common temporal resolution, with ERA5 inputs and the A1/A4 training configurations otherwise unchanged. Compare the OOD CRPS drop relative to in-distribution training under this independent target with the drop measured on IMERG. If the drops remain in the same range, the geographic-shift interpretation survives; if they largely disappear, the RainShift headline is confounded by regional IMERG bias. As a cheaper secondary check, regress the six per-region OOD drops on ERA5-minus-IMERG bias metrics (mean error, quantile slopes) computed over 2001-2020; high correlation would corroborate the confound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"RainShift's headline claim is that the benchmark measures geographic generalization of downscaling models. This requires the paired ERA5/IMERG products to be globally consistent in their errors, so that a model trained in one region and tested in another experiences only a change in climate, not in data quality. The Methods assert that 'the use of globally consistent satellite and reanalysis data enables a controlled benchmark setup,' but no evidence is given. Both products have known, strongly region-dependent biases: ERA5 is constrained by a heterogeneous observation network and degrades in data-sparse regions, while IMERG Final Run includes gauge calibration whose station density varies geographically; both products are particularly uncertain over complex terrain and convective/high-precipitation regimes, which are exactly the six evaluation regions (Amazon, Tibetan Plateau, Melanesia, etc.). Therefore the 30%/17% (and Discussion's 37%) OOD CRPS drops may be partly a regional product-bias artifact rather than a pure geographical/climatic gap. The preprocessing step compounds this: ERA5 precipitation is clipped using thresholds derived from IMERG min/max, injecting target-product statistics into the input stream, and the quantile-mapping experiment then aligns target CDFs to training CDFs. What looks like 'data alignment helps' may be correcting product bias rather than climatic shift. This concern does not invalidate the dataset, but it does undermine the clean interpretation of the headline numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RainShift, a large-scale benchmark dataset and evaluation framework for studying geographic generalization of deep-learning-based precipitation downscaling. The dataset pairs ERA5 reanalysis fields (coarse input) with IMERG satellite precipitation (high-resolution target) over 12 training regions and 6 evaluation regions, and defines four hierarchical training scenarios (A1–A4) to simulate increasing availability of high-resolution observations. The authors evaluate a deterministic ResNet, a WGAN-GP, a diffusion model, and bilinear interpolation, using CRPS as the primary metric, and supplement the main results with in-region training baselines and a quantile-mapping input-alignment experiment. The headline findings are that all learned models outperform bilinear interpolation, generative models outperform the deterministic ResNet, out-of-distribution performance drops by up to 30% (A1) and 17% (A4) relative to in-distribution training, expanding the training domain helps only partially, and quantile mapping improves performance in most target regions.","tokens_in":15825,"tokens_out":3728,"duration_ms":43673,"significance":"If the benchmark's premise holds, RainShift fills a real gap: existing downscaling benchmarks largely evaluate within a single region or on a limited set of regions, whereas cross-geography generalization is a central operational problem for global applicability. The paper's strengths include a substantial public dataset (300 GB of Zarr data), a clear temporal train/validation/test split, multiple model classes, in-region upper-bound comparisons, and a concrete domain-alignment proposal with public code. The quantitative results, while preliminary in places, provide a useful reference point for future method development. However, the central interpretative claim — that the benchmark measures geographic/climatic distribution shift rather than regionally varying data-product biases — is not yet supported by direct evidence, and several reporting issues (an unexplained extreme outlier, lack of repeated-seed statistics, and an overstatement of quantile-mapping gains) limit the confidence one can place in the specific numerical conclusions.","major_comments":[{"comment":"The claim that 'globally consistent satellite and reanalysis data enables a controlled benchmark setup' is not substantiated. Both ERA5 and IMERG have known, region-dependent error characteristics: ERA5 is constrained by a heterogeneous assimilation network and is less reliable in data-sparse and complex-terrain regions, while IMERG Final Run incorporates gauge calibration whose station density varies geographically. The six evaluation regions (Amazon, Tibetan Plateau, Melanesia, etc.) are precisely regions where such product errors are expected to be largest. Furthermore, the preprocessing step clips ERA5 precipitation using IMERG-derived min/max thresholds, injecting target-product statistics into the input stream. Consequently, the reported out-of-distribution CRPS drops (Figures 5–6, Table 4) may partly reflect regional product-bias artifacts rather than pure geographic/climatic shift. To support the benchmark's central premise, the authors should provide diagnostic evidence of product consistency, for example by comparing ERA5 and IMERG against available local gauge or radar measurements in the evaluation regions, and by reporting the sensitivity of the main results to the clipping step (e.g., omitting clipping or varying thresholds).","section":"Methods (Input data, Target data, Data processing); Table 4"},{"comment":"The ResNet A1 entry for the Tibetan Plateau (E5) is reported as CRPS = 18113.558, which is several orders of magnitude larger than all other values. The text attributes this to numerical instabilities during inference, but the numerical value is still included in the table, and the caption describes it as showing mean pixel-wise CRPS. This outlier should either be excluded and replaced with a placeholder (e.g., 'N/A' or 'unstable') or analyzed quantitatively so that the reader can understand whether it is a single divergent sample, an overflow in the CRPS calculation, or a genuine model failure. As reported, the number obscures the comparison among models for that cell and could distort any aggregate analysis computed from Table 4.","section":"Table 4, Results ('Probabilistic models outperform deterministic ones')"},{"comment":"All results in Table 4 and Figures 5–6 appear to be based on a single training run per model and training scenario; no repeated-seed statistics or confidence intervals are reported. Given that GANs and diffusion models are stochastic both in training and in sampling, and that many performance differences in Table 4 are small (e.g., GAN vs. diffusion differences of 0.01–0.02 mm/h), it is not clear that the qualitative conclusions (generative models outperform ResNet; expansion from A3 to A4 helps in some regions) are robust to training seed variability. The authors should provide mean and standard deviation (or at least seed-level results) across several training runs, or otherwise justify why the 8-sample CRPS ensemble is sufficient to establish the reported patterns.","section":"Evaluation ('Quantitative evaluation'), Training details"},{"comment":"The quantile mapping results are overstated relative to the data. The text says quantile mapping 'greatly improves performance' and 'across nearly all regions,' while Table 1 shows that the diffusion model improves only in E3 and E5, degrades in E1 and E6, and is unchanged in E2 and E4; the GAN improves in E3, E5, and E6, but degrades in E1 and E3 (E3: 0.093 to 0.080, actually improvement; E1: 0.075 to 0.093 is a degradation; E3 improves; E5 improves; E6 improves; E2 unchanged; E4 unchanged). The caption itself notes 'except Cape Horn,' but the main text's phrasing is stronger than the evidence. The paper should either soften the claim to 'improves performance for some models and regions' or provide a more nuanced per-model characterization.","section":"Methods ('Quantile mapping for geographical generalization'), Table 1, Results ('Geographical factors dominate…"}],"minor_comments":[{"comment":"The manuscript contains several typographical errors, including 'atomospheric' (Methods, RainShift dataset), 'particulary' (Methods, Data processing), 'high-reslution' (Background), and 'Y et' (Abstract). These should be corrected.","section":"General"},{"comment":"The text refers to 'absolute (see Figure 4)' when reporting A4 GAN and diffusion performance, but Figure 4 is a map of training and evaluation regions, not a performance plot. The absolute CRPS values are in Table 4; the citation should be corrected.","section":"Results ('Probabilistic models outperform deterministic ones')"},{"comment":"The heatmaps in Figures 5 and 6 are described in the text as showing percentage improvements/drops, but the color scale and exact numerical values are not defined in the captions or in the text. Adding a colorbar with units and describing how the percentages are computed (e.g., relative to which baseline) would improve interpretability.","section":"Figure 5 and Figure 6"},{"comment":"The caption states that 'Applying quantile mapping equals or improves performance across most models and regions (except Cape Horn),' but the table shows that for the diffusion model, Melanesia (E6) also degrades (0.295 to 0.310), and for the GAN, E1 degrades (0.075 to 0.093). The caption should be updated to reflect the full pattern of results.","section":"Table 1 caption"},{"comment":"The repository is said to be 'made available upon acceptance,' which is acceptable for a submission, but the authors should clarify the planned license for the benchmark code and dataset, as well as whether the dataset can be downloaded directly from Hugging Face without additional steps.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful benchmark and the public dataset is a clear contribution. The main risk is not the benchmark itself, but the interpretation of the results as measuring geographic distribution shift without addressing regional product biases in ERA5 and IMERG. This is fixable with additional analyses, so I recommend major revision rather than rejection. I would also encourage the editor to consider whether the paper's level of statistical rigor (single runs, no error bars) meets the standards of the target venue; the authors should be asked to add repeated-seed experiments or explicitly discuss the computational cost and why it is prohibitive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Three things you should know up front. First, RainShift is the first global benchmark built explicitly around geographic distribution shift for precipitation downscaling, with 18 regions, four hierarchical training scenarios, three model families, and sensible baselines. That is a real gap, and the paper fills it. Second, the headline empirical claims (generative models beat ResNet; OOD performance drops; expanding training areas helps unevenly) are consistent with the reported tables and are the kind of results the community needs. Third, the paper is not ready as-is: no code release, single runs without error bars, and a main table containing a numerically unstable CRPS of 18113.558 that is dismissed as N/A in the figures. Those are fixable, but they are exactly the things a referee should push on.\n\nThe strongest part of the paper is the dataset and the framing of the problem: zero-shot downscaling to data-sparse regions, with the Global South as the motivation, is important and the hierarchy of training scenarios is well designed. The comparison of GAN versus diffusion versus ResNet under the same protocol is useful even if the absolute numbers shift with more careful runs.\n\nNow the soft spots, in proportion. The missing error bars are the most immediate technical problem: every model ranking in Table 4 rests on a single run, so I would not let the specific ordering of GAN and diffusion be cited as a firm result without repeated seeds. The absence of the repository means I cannot verify the preprocessing steps, the exact patch sampling, or the quantile-mapping implementation from the paper alone. The quantile-mapping claim is slightly overstated relative to Table 1: it helps in most regions but hurts in Cape Horn, and the text should say so.\n\nThe stress-test concern about the controlled setup is legitimate and I think it lands. The paper asserts that globally consistent ERA5 and IMERG data enable a controlled benchmark, but both products have well-documented region-dependent biases, especially over complex terrain and in gauge-sparse areas like the Tibetan Plateau and Melanesia. The preprocessing also clips ERA5 precipitation to IMERG's min/max, which injects target-product statistics into the input stream and can renormalize the exact distribution shift the benchmark claims to measure. So the reported 30%/17% (or 37% in the Discussion) drops are likely a compound of climatic shift and product-bias shift. That does not destroy the benchmark's comparative value—all models face the same confound—but it does undermine the clean interpretation of the magnitude of the drops, and it raises the question of whether quantile mapping is correcting climate mismatch or simply compensating for product bias.\n\nMy verdict: worth a serious referee, conditional accept at best. I would ask for code release, multi-seed runs, and an explicit sensitivity analysis or at least a discussion of regional data-product biases before trusting the headline claims. The dataset and benchmarking effort deserve to be in the literature; the current draft overtitles its clean setup.","headline":"RainShift is a genuinely useful benchmark for cross-geography downscaling and deserves reviewer time, but the 'controlled setup' framing is untested and the headline OOD drops likely mix geographic shift with regional product biases.","tokens_in":16436,"tokens_out":2274,"would_cite":true,"duration_ms":27581,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RainShift, a global benchmark built from reanalysis and satellite rainfall, shows that state-of-the-art precipitation downscaling models degrade by up to 30% when applied to unseen regions, and that quantile-mapping input alignment…","keywords":["precipitation downscaling","geographic generalization","distribution shift","deep learning","benchmark dataset","ERA5","IMERG","quantile mapping"],"falsifier":"Measure ERA5-minus-IMERG residuals against independent gauge or radar data separately in each training and evaluation region; if those residual differences across regions are comparable in size to the reported out-of-distribution CRPS drops, the benchmark's geographic shift is partly an artifact of data-product bias rather than pure generalization.","tokens_in":1477,"feed_emoji":"🌧️","tokens_out":2057,"duration_ms":84614,"temperature":0.7,"pith_summary":"RainShift is a dataset and benchmark for testing whether deep-learning precipitation downscaling models trained in data-rich regions, mostly in the Global North, can be applied to data-scarce regions, mostly in the Global South. The paper pairs coarse ERA5 reanalysis fields with high-resolution IMERG satellite rainfall across 12 training and 6 evaluation regions, and evaluates a deterministic ResNet, a Wasserstein GAN, and a diffusion model. Its central finding is that every learned model beats bilinear interpolation in unseen regions, but out-of-distribution performance drops by up to 30% relative to training on the target region, and the drop remains up to 17% even with the largest training domain. The paper also shows that aligning the input rainfall distributions of target regions to the training region with quantile mapping improves or matches performance in most regions, suggesting that distribution shift, not model architecture, is the main barrier to geographic generalization. If right, this gives the community a standard way to measure and improve transfer of downscaling models to regions with scarce observations.","feed_headline":"Rain downscaling models lose up to 30% on unseen regions","feed_subtitle":"A global benchmark finds AI downscaling degrades on unseen regions, and input alignment recovers much of the loss.","key_machinery":"The load-bearing object is the benchmark itself: a set of 18 fixed 20-degree-by-20-degree patches, 12 training and 6 evaluation, built from paired hourly ERA5 reanalysis inputs (nine atmospheric variables plus land-sea mask and orography) and IMERG satellite precipitation targets at 2.5x upsampling. Training configurations A1 through A4 are hierarchical subsets of the training patches, simulating increasing observational coverage. Evaluation uses pixel-wise CRPS with eight samples, against bilinear interpolation as a lower bound and in-region training as an upper bound. The paper's corrective mechanism is multiplicative quantile mapping, which maps the target region's historical input CDF onto the training region's CDF and applies that transfer function to future inputs, aligning distributions before normalization.","core_discovery":"The paper's central claim is that geographic distribution shift is the dominant obstacle to using learned downscaling models in new regions, and that this shift is measurable and partly correctable. Across six target regions and four hierarchical training configurations, all three learned models improve on bilinear interpolation, yet CRPS relative to in-distribution training falls by up to 30% for the single-region setup and up to 17% for the full Global North setup. Adding more training regions helps in some areas but not uniformly, and on-target training is not always best. The paper further shows that applying multiplicative quantile mapping to the low-resolution precipitation inputs before inference reduces the distributional mismatch and improves or equals the unaligned baseline in all target regions except Cape Horn, even stabilizing a ResNet that otherwise produces numerically unreliable CRPS on the Tibetan Plateau. The benchmark is offered as a standardized zero-shot task: train on the training regions, evaluate on the evaluation regions, with CRPS as the headline metric.","pith_inferences":["If the same splits were evaluated with a local radar or gauge target instead of IMERG, the out-of-distribution drops could be larger, because regional product biases would add to the geographic shift.","Extending quantile mapping from the single precipitation input to all nine ERA5 variables might close more of the remaining gap; the paper reports only the precipitation-input variant.","The strong correlation between mean precipitation and error suggests region difficulty could be predicted from climatology, allowing future benchmarks to stratify results by expected difficulty rather than treating regions as exchangeable."],"forward_implications":["All learned downscaling models beat bilinear interpolation in every evaluation region, so the learned coarse-to-fine mapping transfers at least partially even when the target region is unseen.","Out-of-distribution CRPS drops of up to 30% in the single-region setup and up to 17% in the largest setup show that simply adding training regions does not eliminate the geographic gap, meaning zero-shot use in data-sparse regions carries a measurable accuracy penalty.","Probabilistic models, both GAN and diffusion, reliably outperform the deterministic ResNet, with the diffusion model more robust to a small training domain and the GAN benefiting more from added regions.","Quantile-based input alignment improves or matches unaligned performance in all target regions except Cape Horn, including stabilizing a ResNet that otherwise produces numerically unreliable scores on the Tibetan Plateau.","Performance differences between GAN and diffusion are small relative to differences between geographic regions, so regional climate variation is the dominant factor limiting generalization."],"supporting_citations":[{"why":"Supplies the coarse-resolution ERA5 reanalysis fields used as model input.","marker":"[28]"},{"why":"Defines the IMERG satellite precipitation product used as the high-resolution target.","marker":"[30]"},{"why":"Documents the IMERG V07 Final Run product used to build the target data.","marker":"[31]"},{"why":"Provides the generative deep-learning downscaling architecture and variable choice used for baselines.","marker":"[13]"},{"why":"Supplies the Wasserstein GAN with gradient penalty used as the GAN baseline.","marker":"[21]"},{"why":"Provides the diffusion-based downscaling framework used as the diffusion baseline.","marker":"[22]"},{"why":"Supplies the quantile-mapping method adapted for aligning input distributions across regions.","marker":"[54]"},{"why":"Motivates the A2 training configuration that combines North American and European regions.","marker":"[14]"},{"why":"Motivates the choice of training regions and the transferability evaluation.","marker":"[23]"}],"fun_headline_variants":["Geographic shift cuts AI rain downscaling accuracy by 30%","New benchmark tests rain downscaling across global regions","AI rain models degrade on unseen geographies; alignment helps","Rain downscaling fails on new regions unless inputs aligned","RainShift: global test reveals 30% drop for AI downscaling"],"cache_read_input_tokens":18560,"weakest_assumption_plain":"The load-bearing premise is that ERA5 reanalysis and IMERG satellite rainfall are globally consistent enough that a model trained in one region and tested in another isolates geographic distribution shift rather than regional data-product bias.","fun_headline_variants_meta":{"raw":{"variants":["Geographic shift cuts AI rain downscaling accuracy by 30%","New benchmark tests rain downscaling across global regions","AI rain models degrade on unseen geographies; alignment helps","Rain downscaling fails on new regions unless inputs aligned","RainShift: global test reveals 30% drop for AI downscaling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2349,"prompt_tokens":968,"completion_tokens":1381,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":1294}},"tokens_in":584,"tokens_out":1381,"duration_ms":9127,"temperature":1.0,"reasoning_tokens":1294,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:35:50.520337+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure ERA5-minus-IMERG residuals against independent gauge or radar data separately in each training and evaluation region; if those residual differences across regions are comparable in size to the reported out-of-distribution CRPS drops, the benchmark's geographic shift is partly an artifact of data-product bias rather than pure generalization.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the coarse-resolution ERA5 reanalysis fields used as model input."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the IMERG satellite precipitation product used as the high-resolution target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the IMERG V07 Final Run product used to build the target data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the generative deep-learning downscaling architecture and variable choice used for baselines."},{"cited_title":"J., Sobie, S","cited_arxiv_id":null,"evidence_quote":"Supplies the quantile-mapping method adapted for aligning input distributions across regions."},{"cited_title":"Further analysis of cGAN: A system for Generative Deep Learning Post-processing of Precipitation","cited_arxiv_id":"2309.15689","evidence_quote":"Motivates the A2 training configuration that combines North American and European regions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the choice of training regions and the transferability evaluation."}],"review_version":1}