{"id":"d53b372d-74e7-449c-8eb1-269d93eda874","arxiv_id":"2502.02000","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hierarchical Bayesian model estimates nonstationary extreme rainfall distributions and finds robust 10-35% increases in 100-year daily rainfall across the Western Gulf Coast from 1940 to 2022.","lead":"This paper introduces a Bayesian model that combines regional rainfall data and CO2 trends to estimate how extreme daily rainfall in the Western Gulf Coast is changing over time. The model finds that 100-year rainfall amounts have increased by 10 to 35 percent since 1940, and that current engineering guidelines may underestimate future rainfall.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'robust' increase is supported only by posterior-mean maps (Fig. 7) without credible intervals, and the nonstationary model fails to outperform the pooled stationary model in out-of-sample LogS/CRPS (Table 2); the 10–35% claim is not yet statistically substantiated.","rationale":"The stress-test pass should identify the most load-bearing condition for the central claim. The strongest_claim is that the model quantifies robust increases, with 100-year return levels up 10-35%. The paper's own evidence for this is Fig. 7, a set of posterior-mean maps with no uncertainty. In a Bayesian analysis, posterior means alone do not establish that an effect is 'robust'; one needs intervals or probabilities. The MCMC framework makes these computable, so the omission is not a computational limitation but a missing justification. The out-of-sample comparison (Table 2) compounds the concern: the nonstationary model has LogS and CRPS slightly worse than the pooled stationary model (1.9495 vs 1.9322; 0.2574 vs 0.2548) and wins only at QS for p=0.98/0.99. Thus the evidence that the temporal trend is real and useful is mixed. If the posterior intervals for the percentage-change maps are wide, the claim that increases are robust and widespread fails; if they are narrow, the claim is supported despite the model comparison. This is the single, testable condition on which the validity of the central claim rests. The reader's weakest_assumption (linear ln(CO2), fixed shape, no ENSO) is a separate robustness question: it addresses bias in the point estimates, but it is secondary because it presupposes that the point estimates are well determined. The manuscript contains independent support for the spatial pooling (spatial cross-validation, Figures 4-5) and good calibration (Figure 3), which should be credited; but the absence of uncertainty on the headline maps remains the most direct gap between the evidence and the conclusion. Hence the verdict should be unchanged: CONDITIONAL.","tokens_in":23547,"tokens_out":14069,"duration_ms":142889,"concrete_test":"Using the publicly available code and MCMC output, compute the 95% posterior credible interval for the percentage change in 100-year return level from 1940 to 2022 at each station/grid cell in Fig. 7. Report (i) the fraction of cells with the entire credible interval above 0 and (ii) the fraction of cells whose interval falls within [10%, 35%]. If the first fraction is not close to 1 (e.g., <0.9), then 'throughout the study region' is not supported. As a secondary check, report the same calculation for the 10-year return level.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2.2 presents the paper's central quantitative claim: 'return levels have increased by between 10 and 35% over the past 80 years throughout the study region' (Fig. 7). However, Fig. 7 and its gridded version Fig. A4 show only posterior means; no credible intervals or posterior probabilities are reported for the percentage-change maps. Because the MCMC chains are available, this is a reporting gap, not a method limit, but it is load-bearing: the adjectives 'robust' (Abstract, Section 5) require that the posterior mass for the increase lies away from zero across the region. The only displayed uncertainty (Fig. 9) is for return-level time series at a few cities; it does not cover the 'throughout the study area' claim. Independently, Table 2 shows the Spatially Varying Covariates Model has LogS 1.9495 vs 1.9322 for the Pooled Stationary Model, and CRPS 0.2574 vs 0.2548; QS at p=0.9 is also worse (0.4712 vs 0.4671). Thus the data do not clearly favor the nonstationary model despite its additional parameters. The central claim therefore depends on an unstated assumption that the posterior means in Fig. 7 are sufficiently precise and that the temporal signal is real; neither is established in the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hierarchical Bayesian spatial model, the Spatially Varying Covariates Model, for nonstationary frequency analysis of daily extreme precipitation. GEV location and scale parameters are regressed on ln(CO2) with spatially varying coefficients modeled by Gaussian processes, while the shape parameter is constant in space and time. The model is applied to annual maxima from 181 GHCN stations in the Western Gulf Coast. Validation includes temporal and spatial cross-validation, MCMC diagnostics, and comparison with NOAA Atlas 14. The main scientific claim is that 100-year (and 10-year) return levels increased by 10–35% from 1940 to 2022 throughout the study region, with larger increases near Houston and New Orleans, and that future RCP6.0 projections exceed Atlas 14 at several cities.","tokens_in":23886,"tokens_out":2459,"duration_ms":25575,"significance":"If the central claim is correct, the paper has direct practical relevance: it would imply that stationary guidance such as NOAA Atlas 14 underestimates current and future extreme rainfall in parts of the Gulf Coast. Methodologically, the model is a sensible synthesis of regionalization and process-informed nonstationarity, and it is reasonably validated: the authors provide out-of-sample temporal and spatial cross-validation, multiple scoring rules, MCMC trace plots and R-hat checks, and publicly available code. The paper is also honest about limitations (fixed shape, independence conditional on parameters, computational cost). The main weakness is that the headline 'robust increase throughout the study area' is not backed by the uncertainty quantification that the Bayesian framework is designed to provide, and the model comparison does not clearly favor the nonstationary model over a stationary pooled alternative.","major_comments":[{"comment":"The central claim that 'return levels have increased by between 10 and 35% over the past 80 years throughout the study region' is presented only as posterior means on maps, with no credible intervals, no posterior probability that the increase exceeds zero, and no spatial summary of uncertainty. Because the MCMC chains already exist, this is a reporting gap rather than a methodological limitation. Please add, for the percentage-change maps, either (a) maps of posterior standard deviation or 95% credible interval width, (b) a map of the posterior probability that the change is positive, or (c) interval estimates for representative grid cells or regions. Without this, the adjectives 'robust' in the Abstract and Section 5 and 'throughout the study area' in Section 3.2.2 are not supported.","section":"Section 3.2.2, Figs. 7 and A4"},{"comment":"The out-of-sample comparison does not favor the Spatially Varying Covariates Model over the Pooled Stationary Model: the stationary model has lower LogS (1.9322 vs. 1.9495), lower CRPS (0.2548 vs. 0.2574), and lower QS at p=0.9 (0.4671 vs. 0.4712). The nonstationary model improves QS only at p=0.98 and p=0.99 (0.1680 vs. 0.1689 and 0.1018 vs. 0.1027). The text in Section 3.3.1 and Section 5 states that the model 'performs similarly to the stationary framework' and 'outperforms the nonstationary framework at individual stations,' which is fair, but the stronger framing in the Introduction and Conclusions that the model is validated 'through cross-validation and multiple performance metrics' should be calibrated to this result. Please report uncertainty in the score differences (e.g., block bootstrap or per-station score distributions) or a formal model-comparison statistic such as DIC/WAIC, so readers can see whether the observed differences are meaningful rather than noise.","section":"Table 2"},{"comment":"The nonstationary signal is entirely captured by a linear regression on ln(CO2), with no other time-varying covariates and with the shape parameter fixed in space and time. The paper gives reasonable physical and statistical justifications for these choices, but it does not test whether the conclusions are sensitive to them. Given that the pooled stationary model performs comparably out-of-sample, a reader cannot rule out that the estimated trends are a consequence of the linear-in-ln(CO2) assumption rather than a robust feature of the data. Please add a sensitivity analysis: for example, include an ENSO index or a quadratic time term as an additional covariate, or allow the shape parameter to vary slowly in space, and report whether the 10–35% return-level increase persists. A qualitative statement about the plausible direction of bias is not sufficient for the strength of the claim.","section":"Section 2.3.2 and Eqs. (10)–(11)"},{"comment":"The probability integral transform (PIT) histogram in Fig. 3 is described as 'generally displaying a nearly uniform shape,' but no quantitative calibration test is provided. Because this is one of the main pieces of evidence for model adequacy, please report a formal uniformity test (e.g., Anderson–Darling or a chi-square statistic on the PIT values) or, failing that, the number of observations falling in each decile. This is a relatively small issue compared to the previous two, but it would strengthen the validation section.","section":"Section 3.1, Fig. 3"}],"minor_comments":[{"comment":"The kernel subscript in Eq. (21) reads K_{βσ0}; this is likely a typo for K_{βσ}. Please fix.","section":"Section 2.3.3, Eq. (21)"},{"comment":"The heading 'Nonpooled Nonstatioanry Model' contains a typo; it should be 'Nonpooled Nonstationary Model.'","section":"Section 2.3 heading"},{"comment":"The sentence 'The Bayesian framework is built in R and the stan programming language,,' has a double comma and should be rephrased.","section":"Section 2.5"},{"comment":"The color scales for the 'Difference' rows use a map that starts at 10% and ends at 35%; it would be helpful to state explicitly in the caption that values below 10% or above 35% are not present in the posterior mean, or to use a diverging scale that includes zero so the reader can see where changes are near zero.","section":"Figure 7 and A4"},{"comment":"The caption for Fig. 9 refers to 'RCP 6' but the text in Section 3.3.2 says 'RCP6'; please use one consistent abbreviation throughout.","section":"Section 3.3.2, Fig. 9"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the methodological contribution is reasonable. My main concern is that the headline claim of 'robust increases throughout the study area' is not supported by the reported uncertainty, and the model comparison in Table 2 does not show a clear advantage of the nonstationary model. Both issues are fixable in a revision. I also note that the reference list and data/code availability statements are complete and transparent, which is a strength. I would not recommend rejection, but I would require the authors to add uncertainty quantification for the trend maps and a more careful interpretation of the cross-validation results before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. First, the model itself is a genuine, useful extension: it pools GEV location and scale parameters through Gaussian processes, lets the ln-CO2 regression coefficients vary smoothly in space, and validates with both temporal and spatial cross-validation, multiple scoring rules, and a comparison to NOAA Atlas 14. The code and data are public, MCMC diagnostics look reasonable, and the spatial CV maps (Figs. 4-5) are a good check that the GP interpolation is not just chasing noise. Second, the main quantitative claim—\"robust increases\" of 10-35% in 100-year daily rainfall across the whole region—is not actually backed by the reported statistics. Figure 7 and its gridded version A4 show only posterior means, with no credible intervals or posterior probabilities for the percentage-change maps. The only uncertainty shown is for return-level time series at a few cities (Fig. 9), which does not cover the \"throughout the study area\" statement.\n\nI want to give credit where it is due. The SVCM is a sensible middle ground between fully local nonstationary fits and fully stationary pooled models. The comparison with the Nonpooled Nonstationary model is particularly informative: it visually demonstrates that spatial pooling stabilizes the otherwise noisy station-level trend estimates. The calibration histogram is a nice diagnostic, and the authors are transparent about model limitations (conditional independence, fixed shape, computational scaling). The paper is honestly written, not overselling the method.\n\nThe soft spots are real but proportional. The stress-test note is right: the nonstationary model does not clearly outperform the pooled stationary model in out-of-sample scores (Table 2). SVCM has LogS 1.9495 vs 1.9322, CRPS 0.2574 vs 0.2548, and QS at p=0.9 0.4712 vs 0.4671. It only wins at the 0.98 and 0.99 quantile scores. So the data do not strongly favor the nonstationary model overall, and the \"robust\" language in the abstract and conclusions is stronger than the evidence. The simplest fix is reporting posterior intervals for the beta coefficients and for the percentage-change maps; the MCMC chains are already there. Also, the covariate is only ln CO2, shape is fixed in space and time, and ENSO is deliberately excluded; these are defensible simplifying choices, but they mean the 10-35% range is conditional on that specific model, not a model-independent fact.\n\nWho should read it? Hydrologists and engineers thinking about updating IDF curves under a changing climate will find the framework practical and the Atlas 14 comparison directly relevant. Statisticians working on spatial extremes will find the approach familiar but the application useful. It deserves a serious referee, not a desk reject, but the revision needs to confront the uncertainty gap and discuss the validation comparison honestly. I'd engage with it, and I'd recommend the same to you.","headline":"A useful, well-validated hierarchical Bayesian framework for nonstationary extreme rainfall, but the 'robust 10–35% increase' claim is not yet supported by the reported uncertainty or the model-comparison scores.","tokens_in":24414,"tokens_out":2225,"would_cite":true,"duration_ms":24638,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","62M30","62F15","62P12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a hierarchical Bayesian model pooling daily rainfall extremes across space and letting the distribution shift with atmospheric CO2 shows 100-year daily rainfall return levels in the Western Gulf Coast rose 10 to 35…","keywords":["extreme precipitation","nonstationary GEV","Gaussian process","hierarchical Bayesian model","return levels","Western Gulf Coast","climate change","intensity-duration-frequency curves"],"falsifier":"Refit the model on the same 181 stations allowing the GEV shape parameter to vary in space and time and adding an ENSO index as a second covariate; if the posterior distributions for the 1940-to-2022 change in 100-year rainfall then include zero over most of the study region, the paper's central claim of robust 10 to 35 percent increases is falsified.","tokens_in":23340,"feed_emoji":"🌧️","tokens_out":6284,"duration_ms":55084,"temperature":0.7,"pith_summary":"This paper tries to show that nonstationary extreme-rainfall statistics can be estimated reliably from rain-gauge records if the nonstationarity is pooled across space. It proposes the Spatially Varying Covariates Model, which lets the location and scale of a generalized extreme-value distribution depend on $\\ln(\\mathrm{CO}_2)$ at each site, with those dependencies varying smoothly over the region through Gaussian processes. Applied to daily annual maxima at 181 Western Gulf Coast stations, the model finds that extreme rainfall intensity and variability increased throughout the study area, with 100-year daily rainfall return levels 10 to 35 percent higher in 2022 than in 1940. If this is right, stationary guidance such as NOAA Atlas 14 understates today's rainfall hazard in parts of the Gulf Coast and will understate it further under continued emissions, while the model also offers a template for nonstationary frequency analysis elsewhere.","feed_headline":"Western Gulf Coast 100-year rainfall up 10-35% since 1940","feed_subtitle":"A Bayesian spatial model says stationary flood guidance understates today's extreme rain risk.","key_machinery":"The load-bearing object is the Spatially Varying Covariates Model: a hierarchical Bayesian model in which each site's annual maximum follows a GEV distribution whose location $\\mu(s,t)=\\alpha_\\mu(s)+\\beta_\\mu(s)x(t)$ and scale $\\sigma(s,t)=\\exp(\\log\\alpha_\\sigma(s)+\\beta_\\sigma(s)x(t))$ are linear functions of $\\ln(\\mathrm{CO}_2)$, with the four spatially varying fields ($\\alpha_\\mu$, $\\log\\alpha_\\sigma$, $\\beta_\\mu$, $\\beta_\\sigma$) drawn from independent Gaussian processes using an exponential kernel, and with a single shape parameter held constant across space and time. The Gaussian process layer is the regionalization mechanism: it lets nearby stations borrow strength from one another, smoothing away the noisy and physically implausible coefficient maps produced by separate station-by-station fits, while still allowing the climate response to vary in space. The $\\ln(\\mathrm{CO}_2)$ covariate is the nonstationarity driver, chosen as a low-noise proxy for anthropogenic warming, and it is this combination of spatial pooling and a process-informed covariate that carries the argument that nonstationary return levels can be estimated robustly.","core_discovery":"The central discovery is that regionalizing the climate-covariate response, not just the GEV parameters, makes nonstationary extreme-value trends estimable from short and uneven gauge records. On the Western Gulf Coast, the posterior mean of the $\\ln(\\mathrm{CO}_2)$ coefficient is positive for both the GEV location and scale parameters across most of the domain, implying that the daily-extreme distribution is shifting upward and widening, with the strongest trends near Houston and New Orleans. Return levels for both 10-year and 100-year events increase throughout the study area, with 100-year levels up 10 to 35 percent from 1940 to 2022. In cross-validation, the nonstationary model matches a stationary pooled model on overall scores and beats an unpooled station-by-station nonstationary model, especially for the upper quantiles, which supports the claim that the trend signal is robust rather than an artifact of overfitting. Comparison with NOAA Atlas 14 shows the stationary guidance underestimates current 24-hour rainfall in places such as New Orleans, Galveston, and Mobile while overestimating Houston today; under RCP6 emissions the model projects Atlas 14 will understate Houston's 100-year rainfall after about 2025.","pith_inferences":["Because the model fixes the GEV shape parameter in space and time, a natural stress test is to let shape vary; if shape is actually changing, the reported return-level increases could be misallocated between the center and the tail of the distribution.","The choice of $\\ln(\\mathrm{CO}_2)$ as the sole covariate leaves an opening for natural variability such as ENSO; a model that includes such a covariate could separate forced change from internal variability, which the paper explicitly sets aside.","The 10 to 35 percent range is observation-based, and a testable extension is to compare these return-level maps with radar-based or reanalysis-based estimates over the same period, which the paper notes as future work with alternative data sources.","Because the spatial pooling smooths the climate response, the method may understate localized trends driven by urbanization or land-surface change, since the paper does not include elevation or land-cover covariates."],"forward_implications":["If current stationary intensity-duration-frequency guidance is used for design, present-day 100-year rainfall is understated in parts of the Western Gulf Coast, and the gap widens under continued emissions.","The model yields smooth return-level estimates at ungauged locations, because the Gaussian process layer can interpolate the distribution parameters anywhere in the study domain.","The framework can be adapted to other durations and other regions whenever a credible climate covariate is available, and it accommodates stations with uneven record lengths.","Cross-validation shows that pooling nonstationarity across space performs as well as a stationary pooled model on overall scores and better at the 50-year and 100-year quantiles, so the nonstationary estimates are not bought at the cost of predictive skill."],"supporting_citations":[{"why":"Supplies the generalized extreme-value framework used to model annual maximum rainfall.","marker":"[Coles, 2001]"},{"why":"Provides the precedent for Bayesian spatial modeling of precipitation extremes with Gaussian processes and return-level estimation.","marker":"[Cooley et al., 2007]"},{"why":"Motivates nonstationary intensity-duration-frequency curves and the process-informed nonstationary approach.","marker":"[Cheng and AghaKouchak, 2014]"},{"why":"Defines the NOAA Atlas 14 stationary estimates that serve as the baseline for comparison, including the 1.11 conversion to 24-hour rainfall.","marker":"[Perica et al., 2018]"},{"why":"Documents the implausible station-level trend variability in the Gulf Coast that the proposed spatial pooling is designed to fix.","marker":"[Fagnant et al., 2020]"},{"why":"Establishes the precedent of using CO2 as a covariate in nonstationary extreme-precipitation analysis.","marker":"[Risser and Wehner, 2017]"},{"why":"Provides regional Gulf Coast nonstationary return-value estimates whose findings this paper's results corroborate.","marker":"[Jorgensen and Nielsen-Gammon, 2024]"},{"why":"Offers the semi-Bayesian two-step spatial smoothing alternative that the proposed fully spatial model is contrasted against.","marker":"[Ossandón et al., 2021]"},{"why":"Supplies the Gaussian process theory and the exponential kernel used for spatial pooling.","marker":"[Rasmussen and Williams, 2006]"}],"fun_headline_variants":["Bayesian spatial model locks in Gulf Coast rain extremes rise","Gulf Coast 100-year storm rainfall climbs 10-35% since 1940","Robust trend: Gulf Coast daily extreme rain up since 1940","Beyond stationarity: Bayesian model reveals Gulf Coast rain rise","Nonstationary extremes: Gulf Coast 100-year rain up 10-35%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole trend result rides on the assumption that the change in extreme rainfall is fully captured by a straight-line regression of the GEV location and scale on $\\ln(\\mathrm{CO}_2)$, with the shape parameter held fixed in space and time; if the real relationship is nonlinear or shaped by other drivers, the estimated 10 to 35 percent return-level increases could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian spatial model locks in Gulf Coast rain extremes rise","Gulf Coast 100-year storm rainfall climbs 10-35% since 1940","Robust trend: Gulf Coast daily extreme rain up since 1940","Beyond stationarity: Bayesian model reveals Gulf Coast rain rise","Nonstationary extremes: Gulf Coast 100-year rain up 10-35%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000951,"raw_usage":{"total_tokens":4093,"prompt_tokens":1016,"completion_tokens":3077,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":2980}},"tokens_in":632,"tokens_out":3077,"duration_ms":22359,"temperature":1.0,"reasoning_tokens":2980,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:42:45.221568+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Refit the model on the same 181 stations allowing the GEV shape parameter to vary in space and time and adding an ENSO index as a second covariate; if the posterior distributions for the 1940-to-2022 change in 100-year rainfall then include zero over most of the study region, the paper's central claim of robust 10 to 35 percent increases is falsified.","supporting_citations":[{"cited_title":"Bayesian spatial modeling of extreme precipitation return levels","cited_arxiv_id":null,"evidence_quote":"Provides the precedent for Bayesian spatial modeling of precipitation extremes with Gaussian processes and return-level estimation."},{"cited_title":"Nonstationary precipitation intensity-duration-frequency curves for infrastructure design in a changing climate","cited_arxiv_id":null,"evidence_quote":"Motivates nonstationary intensity-duration-frequency curves and the process-informed nonstationary approach."},{"cited_title":"Bedient, and Katherine B","cited_arxiv_id":null,"evidence_quote":"Documents the implausible station-level trend variability in the Gulf Coast that the proposed spatial pooling is designed to fix."}],"review_version":1}