{"id":"f4a69474-9d7a-44ec-b10c-6eeb49002147","arxiv_id":"2412.16407","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A machine-learning downscaling of 13 CMIP6 models projects that Bangladesh's 100-year extreme rainfall increases by roughly 50 mm/day by mid-century and 100 mm/day by 2100 under SSP5-8.5.","lead":"This study downscales 13 global climate models to project extreme rainfall risk in Bangladesh under four warming scenarios. It finds that 100-year daily rainfall could rise by about 50 mm/day by mid-century and 100 mm/day by end-century under the highest-emissions scenario.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 50/100 mm/day return-level increases rest entirely on an untested stationarity assumption: the downscaling map trained on 1981–2019 ERA5/CHIRPS is applied unchanged to future CMIP6 climates, and present-climate validation cannot constrain this.","rationale":"The reader's weakest_assumption already identifies the stationarity/invariance assumption, and my stress-test confirms it is the most load-bearing condition for the central claim. The paper's present-climate validation, while useful, cannot validate the future application because present-day fields are individually bias-corrected and the training data come from reanalysis and observations. The authors themselves flag the invariance assumption as requiring further validation in the Discussion, and they suggest the natural paired-simulation test without performing it. Thus the headline increase estimates are best read as conditional on that assumption rather than as robust point predictions. The multi-model and multi-scenario spread does provide some sense of projection uncertainty, and the qualitative conclusion of increased extreme rainfall risk in northeastern and southeastern Bangladesh is consistent with the broader literature, so the reader's CONDITIONAL verdict remains appropriate. No verdict change is needed, though the concrete perfect-model test would materially strengthen the paper if it passes, or would reveal a serious systematic bias if it fails.","tokens_in":6489,"tokens_out":4480,"duration_ms":43219,"concrete_test":"Run a perfect-model stationarity test using a model with paired coarse- and high-resolution simulations for both historical and future periods, such as a CMIP6/HighResMIP pair. Fit the two-stage downscaling on the historical coarse-to-high-resolution pairs, apply it to that model's future coarse fields, and compare the downscaled future change in the 100-year rainfall return level against the model's own high-resolution future change. If the discrepancy exceeds the bootstrap uncertainty shown in Figures 5–6, the invariance assumption fails in at least one plausible climate state, and the 50/100 mm/day estimates would need substantially larger uncertainty bounds or revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—roughly 50 mm/day and 100 mm/day increases in the 100-year daily rainfall return level by mid- and end-century—depends on the assumption stated in Section II: 'The same downscaling function is used for present and future climate projections, assuming that the downscaling function remains unchanged in the warming scenario.' The downscaling map is a two-stage GAN trained on ERA5-to-ERA5-Land and upscaled-CHIRPS-to-CHIRPS pairs from 1981–2019, then applied to coarse CMIP6 fields after bicubic interpolation. Present-climate skill (Figures 2–4) does not test this assumption robustly because each model's present-day output is individually bias-corrected against observations, so agreement in the historical period can be achieved even if the learned coarse-to-fine relationship does not generalize. Under warming, changes in storm type, convective organization, or the partitioning of orographic versus non-orographic rainfall can alter the relationship between coarse predictors and local extremes; if GAN-1/GAN-2 distort the CMIP6 model's own future coarse-scale change signal, the projected return-level increments are not reliable. The authors explicitly acknowledge this in the Discussion: the invariance assumption 'requires further validation.' Because no such validation is provided, the headline numbers are conditional on an untested functional invariance, and the uncertainty communicated by the inter-model spread in Figures 5–6 does not include this source of error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the authors' previously developed statistical-physical adversarial downscaling method (Saha and Ravela, 2024a,b) to outputs from thirteen CMIP6 ScenarioMIP models over Bangladesh, producing 0.05° daily rainfall fields for the present climate and for four SSP scenarios through 2100. It validates the downscaling against CHIRPS and ERA5/ERA5-Land in the present climate (Figures 2–4), then uses the same downscaling function to project extreme rainfall risk. The central quantitative claim is that the 100-year return level of daily maximum rainfall increases by roughly 50 mm/day by mid-century and 100 mm/day by end-century under SSP5-8.5, with the largest increases in northeastern and southeastern Bangladesh. The authors explicitly acknowledge in Section IV that the stationarity of the downscaling function requires further validation.","tokens_in":6770,"tokens_out":4369,"duration_ms":37414,"significance":"If the projections are robust, the paper provides actionable, high-resolution risk information for a highly climate-vulnerable region and demonstrates a computationally efficient way to downscale a multi-model, multi-scenario ensemble. Strengths include the physics-plus-adversarial-learning formulation, the use of thirteen CMIP6 models and four SSPs, the present-day validation showing clear improvement over coarse ERA5 in capturing extreme rainfall risk, and the candid discussion of limitations. The paper's value, however, hinges on the untested assumption that the downscaling function trained on 1981–2019 data remains valid in future climates; the headline numbers are conditional on this invariance and on the adequacy of the return-level extrapolation from short records.","major_comments":[{"comment":"The load-bearing assumption is stated as: 'The same downscaling function is used for present and future climate projections, assuming that the downscaling function remains unchanged in the warming scenario.' The present-day validation (Figure 4) cannot constrain this assumption, because each CMIP6 model's present-day output is individually bias-corrected against observations; agreement in the historical period can be achieved even if the learned coarse-to-fine relationship does not generalize. The authors acknowledge in Section IV that this invariance 'requires further validation,' but no validation or sensitivity analysis is provided. I ask the authors to add a concrete test—for example, training on historical CMIP6 data and comparing against observations, using a pseudo-reality experiment with high-resolution future simulations, or at least quantifying how the projected return-level increments change when the training period or downscaling architecture is varied. Without this, the 50/100 mm/day increments are unsupported beyond the assumption.","section":"Section II, 'Data and Methods'"},{"comment":"The headline increases are 100-year return levels fitted from 20–30 year periods (1985–2014, 2031–2050, 2081–2100) using a two-parameter Generalized Pareto distribution. The paper does not report confidence intervals for the fitted return levels or the extrapolation uncertainty; the shaded regions in Figures 5–6 show only inter-model spread, not GP parameter uncertainty or the uncertainty from extrapolating to a 100-year return period. Please provide uncertainty bounds on the fitted return levels (e.g., via bootstrapping as in Figure 4) and state the threshold used for the GP fits. Without this, the quantitative 50/100 mm/day claims are not adequately supported.","section":"Figures 5–6 and abstract"},{"comment":"The statement 'Since present climate data is individually bias-corrected against observations for each model, the variation between models is minimal' indicates that the bias-correction step removes inter-model differences in the present climate. If the same optimal-estimation bias correction is applied to future projections using present-climate statistics, it may also suppress or distort part of the climate-change signal if model biases are non-stationary. Please clarify how the bias correction is applied in the future period—whether it uses only present-climate statistics—and discuss the potential distortion of the projected change signal. A sensitivity test that applies the bias correction only to the present period and leaves future fields uncorrected would help bound this effect.","section":"Section III, 'Results'"},{"comment":"The manuscript does not specify precisely how the return levels are computed from the downscaled daily fields: are they derived from annual maxima, peaks over threshold, or daily regional maxima pooled across the domain, and over what spatial aggregation? The abstract's phrase 'daily maximum rainfall for a 100-year return period' is ambiguous. Please define the variable, the spatial pooling, and the fitting procedure (including threshold selection and number of exceedances) in Section II or III. This information is necessary to interpret Figures 4–8 and to assess the statistical reliability of the extrapolation.","section":"Sections II and III"}],"minor_comments":[{"comment":"The scenario list contains the typo 'SSSP3-7.0'; it should read 'SSP3-7.0'.","section":"Section II"},{"comment":"The text says 'shown in Figure 2f, is upscaled in Figure 2d, and then downscaled to Figure 2e by GAN-2,' but the figure caption does not list a panel (d) and appears to skip it, and the panel order in the caption is (a), (b), (c), (e), (f), (g). Please correct the panel labels and the cross-references.","section":"Figure 2 and accompanying text"},{"comment":"The data availability statement mentions HighResMIP data, but the manuscript uses ScenarioMIP CMIP6 data; this appears to be a leftover from the authors' previous work and should be updated.","section":"Data Availability Statement"},{"comment":"Reference [2] is an arXiv preprint (arXiv:2408.11790) and reference [11] is a prior paper by one of the authors; if the journal requires, please provide a more complete description of the method in the main text so readers are not forced to consult both prior papers to understand the downscaling pipeline.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is largely an application of previously published methods; the novel contribution is the multi-scenario, multi-model CMIP6 ensemble for Bangladesh and the resulting projections. The editor should weigh whether the paper's applied contribution is sufficient for the journal's scope. The central issue is the stationarity assumption, which the authors themselves flag; I do not see this as fatal, but it requires substantial additional work—either a validation experiment or a clearly quantified sensitivity analysis—before the headline numbers can be regarded as reliable. The self-citation to the authors' own prior papers is heavy, but the method description is at least summarized. If the authors can address the major comments with concrete tests, the paper could become acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of arXiv:2412.16407. The useful new content is the projection: thirteen CMIP6 ScenarioMIP models, four SSPs, through 2100, downscaled to 0.05° for Bangladesh. Headline numbers are about 50 mm/day and 100 mm/day increases in the 100-year daily rainfall return level by mid- and end-century under SSP5-8.5, with the largest rises in the northeast and southeast. That is a legitimate extension of their earlier HighResMIP single-scenario work, and it gives planners a concrete ensemble spread.\n\nWhat the paper does well: the present-day validation (Figure 4) shows the downscaled product tracks CHIRPS return levels much better than coarse ERA5, and the spatial patterns in Figures 7-8 are plausible. The method is fast enough to downscale many models, which is a real advantage. The authors are also upfront in the Discussion that the stationarity assumption 'requires further validation.'\n\nThe soft spot is exactly that assumption, and it is load-bearing. The downscaling map is trained on 1981-2019 ERA5/CHIRPS pairs, then applied unchanged to future CMIP6 fields. Because each model's present-day output is individually bias-corrected against observations, historical agreement does not test whether the coarse-to-fine relationship holds under warming. If storm type, convective organization, or orographic partitioning changes, the projected return-level increments are not reliable. The inter-model spread in Figures 5-6 shows model and scenario uncertainty, but it does not include this structural error. Also, fitting GPDs to 20-30 years of daily maxima and extrapolating to 100-year return levels adds tail uncertainty that is only partially shown as shading.\n\nI do not think the stress-test note overstates this. The central qualitative conclusion is probably right—extreme rainfall risk in Bangladesh increases with warming, especially in the northeast and southeast—and that matches the broader literature. But the specific 50 and 100 mm/day figures are conditional on an untested functional invariance. The paper says so itself, which is honest, but it does not fix the problem.\n\nWho this is for: someone working on downscaling extremes or Bangladesh climate risk will want to know this exists; it is a useful ensemble projection. It deserves peer review, but a referee should push for either a transferability test (e.g., train on one period, validate on another) or at least a clear statement that the point estimates are conditional on an assumption that could change the numbers substantially. My vote is to send it to review, not desk reject.","headline":"A useful multi-model, multi-scenario ensemble projection of Bangladesh rainfall extremes, but the headline return-level increases rest on an untested stationarity assumption that the paper itself flags.","tokens_in":7338,"tokens_out":3193,"would_cite":true,"duration_ms":24091,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under the highest-emissions scenario, Bangladesh's 100-year daily rainfall is projected to be about 50 mm/day higher by mid-century and up to 100 mm/day higher by 2100.","keywords":["extreme rainfall","statistical downscaling","Bangladesh","return period","flood risk","climate projections","scenario uncertainty"],"falsifier":"Train the identical downscaling function on the 1981-1999 period and validate its extreme-rainfall return levels against observed data for 2000-2019; systematically growing error with warming would falsify the stationarity assumption and undermine the 50/100 mm/day projections. Alternatively, compare the paper's downscaled future return levels against those from a high-resolution regional climate model run for the same emissions scenarios; divergence would call the projections into question.","tokens_in":6238,"feed_emoji":"🌧️","tokens_out":11292,"duration_ms":86404,"temperature":0.7,"pith_summary":"Bangladesh's extreme-rainfall risk is projected to grow substantially under climate change, according to this study, which applies a fast downscaling pipeline to thirteen global climate models across four emissions scenarios. The central quantitative claim is that the daily rainfall of a 100-year event could rise by about 50 mm/day by mid-century and near 100 mm/day by end-century under the highest-emissions scenario, with the largest increases in the northeastern hilly region and the southeastern coast. This matters because Bangladesh is densely populated, flood-prone, and current coarse climate-model output underestimates extremes. The authors show their downscaled fields reproduce observed spatial patterns and present-day return levels, then apply the same function to future projections and report substantial model-to-model and scenario-to-scenario uncertainty.","feed_headline":"Bangladesh 100-year rainfall could jump 100 mm/day by 2100","feed_subtitle":"Downscaling 13 models puts Bangladesh's biggest rainfall-risk rises in the northeast hills and southeast coast.","key_machinery":"The central mechanism is a fixed downscaling function that combines three components: an ensemble-approximated conditional Gaussian process regressor that gives a first-guess rainfall field, a linear spectral orographic-precipitation model that injects terrain-driven structure, and two stacked generative adversarial networks (machine-learning pairs of a generator and a discriminator) that refine the field first to an intermediate resolution and then to the final high-resolution grid, followed by an optimal-estimation bias correction tuned to observed return levels. The function is trained on 1981-2019 data and then applied unchanged to each model and scenario, so the whole future-risk argument rides on this transfer.","core_discovery":"The paper establishes that a two-stage adversarial downscaling system—first from a coarse meteorological reanalysis to an intermediate land reanalysis, then from upscaled observations to the native high-resolution grid—can turn coarse climate-model rainfall into realistic high-resolution fields, and that this transfer holds when the same trained function is applied to thirteen global climate models under four emissions scenarios. In the present climate, the downscaled fields match observed spatial patterns and generalized-Pareto return-level curves far better than the coarse model alone. Under future scenarios, the downscaled ensemble projects a nationwide rise in 100-year daily rainfall maxima, strongest in the northeast hilly region and the southeast coast; for the highest-emissions scenario the 100-year event intensity increases by roughly 50 mm/day by mid-century and 100 mm/day by end-century. The paper reports that inter-model and inter-scenario spread is large and that the same downscaling function is assumed to remain valid in a warmer climate.","pith_inferences":["Because the paper applies one fixed downscaling function to all future climates, its numbers depend on the stationarity assumption; a natural stress test is to train on an early period and validate on a later observed period, something the paper leaves for future work.","The speed of the method suggests it could be transferred to other data-sparse, flood-prone regions, provided a high-resolution observed rainfall record exists for training; that extension is not in the paper.","The large inter-model spread means that for adaptation decisions, the scenario and model uncertainty may dominate the downscaling uncertainty, so robust-decision approaches may be more appropriate than point estimates."],"forward_implications":["Under the highest-emissions scenario, the 100-year daily rainfall maximum in Bangladesh increases by about 50 mm/day by mid-century and about 100 mm/day by end-century, relative to 1985-2014.","Under the two lower-emissions scenarios, the mid-century increase does not grow much further by end-century, suggesting that limiting emissions would cap the rise in extreme-rainfall risk.","The largest increases are concentrated in the northeastern hilly region and the southeastern coastal zone, regions already exposed to floods and cyclones.","Coarse model output underestimates present-day extreme risk, while the downscaled fields track observed return levels, so risk assessments based on raw coarse output would understate future hazard.","The multi-model, multi-scenario ensemble yields explicit uncertainty bounds on future return levels, allowing adaptation plans to be stress-tested against the spread rather than a single projection."],"supporting_citations":[{"why":"introduces the statistical-physical adversarial downscaling framework that this paper applies.","marker":"[1]"},{"why":"prior application of the method to Bangladesh and source of the schematic figures used here.","marker":"[2]"},{"why":"supplies the high-resolution observed rainfall reference used for training and validation.","marker":"[3]"},{"why":"provides the coarse reanalysis fields on which the first-stage downscaling is trained.","marker":"[4]"},{"why":"provides the intermediate high-resolution land reanalysis used as the first-stage reference.","marker":"[5]"},{"why":"gives the linear orographic-precipitation theory used to inject physics into the first-guess field.","marker":"[8]"},{"why":"supplies the optimal-estimation formalism used for bias correction against observed extremes.","marker":"[9]"},{"why":"defines the future scenario simulations from the global climate models used in the projections.","marker":"[10]"}],"fun_headline_variants":["Bangladesh extreme rain risk to rise 100 mm/day by 2100","Bangladesh 100-year rain could climb 100 mm/day by 2100","Bangladesh's 100-year storm rain up 100 mm/day by 2100","Bangladesh worst-day rain may jump 100 mm/day by 2100"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole future projection rests on the assumption that the statistical relationship between coarse-scale weather and local rainfall learned from 1981-2019 data stays the same in a warmer climate; if it changes, the projected 100-year return-level increases would be systematically biased.","fun_headline_variants_meta":{"raw":{"variants":["Bangladesh extreme rain risk to rise 100 mm/day by 2100","Bangladesh 100-year rain could climb 100 mm/day by 2100","Bangladesh's 100-year storm rain up 100 mm/day by 2100","Bangladesh worst-day rain may jump 100 mm/day by 2100"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001015,"raw_usage":{"total_tokens":4276,"prompt_tokens":926,"completion_tokens":3350,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":3260}},"tokens_in":542,"tokens_out":3350,"duration_ms":21506,"temperature":1.0,"reasoning_tokens":3260,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:36:20.104412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the identical downscaling function on the 1981-1999 period and validate its extreme-rainfall return levels against observed data for 2000-2019; systematically growing error with warming would falsify the stationarity assumption and undermine the 50/100 mm/day projections. Alternatively, compare the paper's downscaled future return levels against those from a high-resolution regional climate model run for the same emissions scenarios; divergence would call the projections into question.","supporting_citations":[{"cited_title":"first-guess","cited_arxiv_id":null,"evidence_quote":"introduces the statistical-physical adversarial downscaling framework that this paper applies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"prior application of the method to Bangladesh and source of the schematic figures used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the high-resolution observed rainfall reference used for training and validation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the coarse reanalysis fields on which the first-stage downscaling is trained."},{"cited_title":"Priming the downscaling model with statistics and physics-derived rainfall fields [8] improves the physical consistency and alleviates data paucity issues","cited_arxiv_id":null,"evidence_quote":"provides the intermediate high-resolution land reanalysis used as the first-stage reference."},{"cited_title":"The climate hazards infrared precipitation with stations—a new environmental record for monitoring extremes","cited_arxiv_id":null,"evidence_quote":"gives the linear orographic-precipitation theory used to inject physics into the first-guess field."},{"cited_title":"The era5 global reanalysis","cited_arxiv_id":null,"evidence_quote":"supplies the optimal-estimation formalism used for bias correction against observed extremes."}],"review_version":1}