{"id":"8ea25d23-031b-41ea-b880-e83ccbad08ba","arxiv_id":"2505.08591","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new percentile-smoothing method using extreme value theory yields stable basin-level estimates of hurricane-force wind extremes down to the 99.999th percentile in Caribbean and Atlantic ASCAT data.","lead":"This paper develops a statistical method for estimating extremely high ocean wind-speed percentiles, down to the 99.9999th, from 15 years of satellite scatterometer data over tropical basins. The method gives consistent estimates whether the tail is fit with an exponential, a generalised Pareto distribution, or no model, a robustness the authors say is needed before reliable decadal hurricane-trend studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"One-week block bootstrap may understate uncertainty in extreme percentiles because tropical-cyclone activity clusters on intraseasonal timescales longer than a week; the 'within uncertainties' claim at the 99.9999th percentile depends on these variances.","rationale":"The reader's weakest-assumption analysis identified the block-bootstrap resampling scheme and the one-week block length as the key vulnerability; I agree. The paper is explicitly a method-development contribution whose central claim is the robustness of extreme-wind percentiles at basin level to tail-model choice. That claim has two parts: (1) the point estimates from exponential, generalised-Pareto, and raw empirical approaches are close down to low exceedance probabilities, and (2) where they differ, the differences are within the bootstrap uncertainties. The first part is supported directly by the figures and is not strongly threatened by the block-length issue, because point estimates are only weakly affected by resampling-block length. The second part, however, is exactly as strong as the uncertainty estimates, and those estimates are the product of a block-bootstrap whose block length is justified only by the timescale of a cyclone crossing a basin. This is a genuine soft spot: tropical-cyclone activity is known to cluster on intraseasonal timescales, and a one-week block cannot reproduce such clustering. The paper does not provide a block-length sensitivity test, and it explicitly acknowledges the choice is heuristic. The proposed check, varying block length, would settle whether the uncertainty estimates are materially underestimated. If the standard errors are stable, the main claim survives; if they grow substantially, the 'trustworthy down to 99.9999th percentile' statement should be downgraded to 'point estimates are consistent, though the uncertainty quantification is not yet fully validated.' Either way, the appropriate verdict remains CONDITIONAL, as the reader concluded, so no change is needed.","tokens_in":12347,"tokens_out":7279,"duration_ms":79124,"concrete_test":"Rerun the Caribbean basin 2020 ASCAT-A 0.125-degree analysis (the running example in Sections 6-7) with block lengths of 3, 7, 14, and 28 days, using at least 200 bootstrap resamples for each length, and record the bootstrap standard errors of the 99.999th and 99.9999th percentiles for the exponential, generalised-Pareto, and raw estimators. If the standard errors grow by more than about 50% when the block length increases from 7 to 28 days, or if the 95% intervals across the three estimators cease to overlap at the 99.9999th percentile, then the reported uncertainties are sensitive to the block-length choice and the 'within uncertainties' consistency claim needs to be qualified. As a complementary diagnostic, compute the autocorrelation function or extremal index of the daily basin-maximum wind speed; if the correlation length exceeds one week, the one-week block is too short.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 8) that results are 'quite consistent within the uncertainties' down to the 99.9999th percentile rests entirely on the block-bootstrap variance estimates from Section 6a. In that scheme, spatial dependence is preserved by resampling whole days, but temporal dependence is captured only by concatenating blocks of about one week, chosen because a tropical cyclone crosses a basin in roughly that time. This ignores intraseasonal clustering of storm activity, e.g., multiple cyclones or active phases associated with the MJO on 30-60 day timescales. A moving-block bootstrap with one-week blocks breaks any positive autocorrelation beyond one week, so the resampled series has weaker long-range dependence than the original. The estimator of interest, a pooled basin-level extreme percentile, averages over the whole year; its variance depends on the full autocovariance structure, not just the storm-crossing timescale. If storminess clusters over two to four weeks, the bootstrap variance of the 99.999th and 99.9999th percentiles will be systematically too small, and the 'within uncertainties' agreement between exponential, generalised-Pareto, and raw estimates could be an artifact of overconfident error bars. The authors test sensitivity to the number of resamplings (keeping the percentile within about 1 m/s between 25 and 150 resamples), but they do not test sensitivity to block length, which is the more directly relevant tuning parameter for dependence. A longer effective block length would not necessarily change the point estimates, but it could widen the confidence intervals enough to weaken the claim that the method is trustworthy down to 99.9999th percentile.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Payez et al. develop a percentile-smoothing method, based on extreme value theory, for estimating extreme marine wind-speed percentiles as a first step toward studying decadal trends in tropical-cyclone winds. Using ASCAT-A Level-3 scatterometer winds (0.125° and 0.25° products) and collocated ERA5 winds for 2007-2021 over the Caribbean and the North and South Atlantic (0-30°), the authors pool all pixels within a basin and year, choose the 99.99th percentile of the pooled sample as the threshold, and estimate higher percentiles from exponential or generalised-Pareto fits of the tail. Uncertainty is obtained from a month-stratified block bootstrap that resamples whole days in blocks of about one week, preserving spatial correlation and seasonality. The fitted percentiles are compared with 'raw' empirical percentiles of the resampled data; the authors report close agreement down to an exceedance probability of 10^-5 (their main result, the 99.999th percentile) and agreement 'within uncertainties' down to 10^-6 (the 99.9999th percentile). The results behave as expected with resolution (012 exceeds 025 exceeds ERA5), and the South Atlantic tail is essentially exponential. No trend statistics are computed; the authors stress that longer multi-instrument records are needed before decadal trend conclusions can be drawn.","tokens_in":12632,"tokens_out":17470,"duration_ms":175433,"significance":"If the robustness claim holds, the paper delivers a practical, well-tested recipe for estimating basin-scale extreme wind percentiles at levels relevant to tropical cyclones, with a dependence-aware uncertainty scheme that can be applied across scatterometer missions. The systematic three-way comparison (exponential, generalised Pareto, raw) over two products, three basins, and fifteen years is a genuine strength, as are the physical sanity checks (resolution ordering, ERA5 lower than observations, exponential tail where no tropical cyclones occur) and the authors' explicit refusal to overclaim trend results. The main risk is the bootstrap uncertainty estimation: the one-week block length may not capture intraseasonal clustering, which would make the reported error bars too small and would directly affect the 10^-6 'within uncertainties' claim. The validation is entirely internal, and no external comparison against independent extreme-wind observations is provided; no code release is mentioned. These points are addressable with targeted additional analyses and rephrasing.","major_comments":[{"comment":"The one-week bootstrap block length is motivated in Section 6a by the time for a tropical cyclone to cross a basin, which is the timescale of an individual storm rather than the timescale over which extreme-wind activity clusters (successive cyclones in an active phase, MJO or other 10-60-day modulation). A moving block bootstrap preserves autocorrelation only up to the block length; if intraseasonal clustering extends beyond one week, the bootstrap variances of the pooled-basin extreme percentiles will be systematically too small. This is load-bearing because the 'quite consistent within the uncertainties' claim for the 99.9999th percentile in Section 8 (Figure 15), and the error bars in Figures 10-12, rest on these variances. The reported sensitivity of the 99.9999th percentile to the number of resamplings (25-150, about 1 m/s) concerns only the point estimate, not the variance estimate. Please add a sensitivity analysis varying the block length (e.g., 3, 7, 14, and 30 days) or adopt a complementary dependence-robust scheme (stationary bootstrap, subsampling), and report whether the conclusions at 10^-5 and 10^-6 survive.","section":"Section 6a, Section 8"},{"comment":"The agreement between the 'fit' and 'raw' estimates is computed from the same block-bootstrap resamples, so it demonstrates internal consistency between the parametric tail models and the empirical distribution of the resampled data, not agreement with an independent extreme-wind reference. The statement that 'the method seems trustworthy at least down to the 99.9999th percentile' (Section 8) and the significance statement therefore go a step beyond the evidence presented. I recommend either adding a targeted external check (e.g., comparing fitted basin-level percentiles with the dropsonde/SFMR-adjusted winds discussed in Section 2c and Table 1) or explicitly stating that the validation is internal and that absolute calibration against an external reference is deferred to follow-on work.","section":"Section 7, Section 8"}],"minor_comments":[{"comment":"The block-bootstrap procedure is described qualitatively; please specify the exact resampling algorithm (block length in days, number of blocks drawn per resample, overlap handling, and how the month-stratified day counts are reconciled with blocks crossing month boundaries) so that the uncertainty estimates are reproducible.","section":"Section 6a"},{"comment":"The assertion that the results are 'quite robust when changing the threshold by a factor 10 up or down in probability of exceedance' is not supported by any figure, table, or quantitative result, in contrast to the other robustness claims in Figures 13-15; please add the supporting results or clearly label this as a preliminary check.","section":"Section 8"},{"comment":"With 50 bootstrap resamples, the variance estimates themselves carry roughly 20% relative Monte Carlo error for a variance; the reported 25-150 resample check covers only the mean percentile (about 1 m/s), so please also report the stability of the variance estimates or increase the number of resamples used for the error bars in Figures 10-12 and 15.","section":"Section 6b"},{"comment":"The fitted tail parameters (exponential rate; GP scale and shape) and their bootstrap uncertainties are not reported; since the extrapolation from the 10^-4 threshold to 10^-6 is governed by the GP shape parameter, a table or figure of shape-parameter estimates by basin and year would make the extrapolation claim transparent.","section":"Sections 5, 7"},{"comment":"The title advertises 'decadal hurricane wind trends', but no trend statistics are computed or claimed in the paper; a title such as 'An extreme value method for estimating hurricane-force wind percentiles from scatterometer data' would match the content better.","section":"Title"},{"comment":"For 2021 the analysis uses MetOp-B ASCAT data for the whole year, following MetOp-A's retirement in November 2021; please clarify what overlap checks (if any) were performed between MetOp-A and MetOp-B to confirm that the change of instrument does not affect the basin-level percentile comparisons.","section":"Section 3"},{"comment":"The estimation method for the exponential and generalised-Pareto tail fits is not stated; please specify the fitting procedure (e.g., maximum likelihood) and the numerical routine used.","section":"Sections 5, 7"},{"comment":"The reference list entry 'Haan, L., and A. Ferreira' should read 'de Haan, L., and A. Ferreira', and the umlauts in 'Künsch' and 'Rootzén' are garbled in the current rendering.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid methods contribution and the authors are appropriately cautious about trend conclusions. The main issue is the sensitivity of the bootstrap uncertainty estimates to the block-length choice, which bears directly on the 10^-6 'within uncertainties' claim; I would like to see that addressed before the claim is fully credited. The validation is entirely internal; a sentence of caveat or a small external check would strengthen the 'trustworthy' phrasing. The paper fits the journal's scope well and I have no concerns about the citation pattern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThe short version: this is a sensible method paper that does what it says—uses EVT tail fits to push basin-level scatterometer wind percentiles down to exceedance probabilities of 10^-5 (99.999th) and estimates uncertainties with a block bootstrap. It doesn't claim trends; it sets up the tool for a later multi-instrument study. The internal consistency checks are extensive and the main result, that exponential, GP, and raw empirical estimates agree to ~1 m/s at the 99.999th percentile across basins and years, is convincing.\n\nWhat's new relative to the KNMI prior work is reaching 10^-5 rather than 10^-2, and doing it with a bootstrap scheme that keeps spatial dependence (resampled whole days) and seasonality (fixed monthly counts). That's a real step up for scatterometer climatology.\n\nSoft spots, in order of importance. First, the uncertainty estimates depend on a one-week block length, justified by tropical cyclone crossing time. The authors test sensitivity to number of resamples but not to block length. If TC activity clusters on intraseasonal timescales (MJO, 30-60 days), the bootstrap variances could be too small, which would weaken the \"within uncertainties\" agreement at 99.9999th percentile. The main 99.999th result is point-estimate robust, so it likely survives, but the claim of trustworthiness at 10^-6 rests on variance estimates that aren't stress-tested. Second, validation is internal: the raw and fit estimates come from the same resampled data, so their agreement measures smoothing consistency, not accuracy against a ground truth. The paper would be stronger with one external comparison, e.g., vs. hurricane best-track wind radii or a non-ASCAT observing system. I don't consider that fatal, but it should be said.\n\nThe writing is clear and the EVT basis is stated honestly (dependence doesn't change the tail assumption, only the variance). The authors also flag their own limitations—no automatic threshold selection, no trend claim yet, need for multi-decadal records. Citation pattern looks appropriate; self-citations to KNMI prior work are justified.\n\nWho's this for? Anyone working on satellite wind extremes or decadal storm trends, and methodologists interested in extreme percentile estimation under strong spatio-temporal dependence. It deserves a serious referee; the block-length sensitivity and external validation should be requests in revision.\n\nI'd probably cite it if I'm working on scatterometer extremes, and I'd bring it to reading group.\n\nBest,\n[Name]","headline":"A solid method paper with an honest internal-consistency argument; the block-bootstrap uncertainty is the main soft spot, but the 99.999th percentile robustness holds without heavy reliance on those variances.","tokens_in":13225,"tokens_out":3087,"would_cite":true,"duration_ms":27436,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","62F40","62P12","86A10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper develops a percentile-smoothing method, combining extreme value theory with block-bootstrap resampling, that makes basin-level hurricane wind extremes measurable down to the 99.999th percentile without strong tail assumptions.","keywords":["extreme value theory","wind-speed percentiles","hurricane winds","tropical cyclones","ASCAT scatterometer","block bootstrap","decadal trends","ERA5"],"falsifier":"Repeat the entire basin-level analysis with block lengths of two, four, and eight weeks, and with block lengths informed by the autocorrelation of daily basin-maximum winds; if the spread of the resampled percentiles widens substantially and the exponential, generalised Pareto, and raw percentiles no longer agree within those wider uncertainties at the $10^{-5}$ exceedance probability, the central consistency claim fails.","tokens_in":12132,"feed_emoji":"🌪️","tokens_out":12210,"duration_ms":110586,"temperature":0.7,"pith_summary":"This paper develops a practical way to estimate extremely rare wind-speed percentiles from satellite scatterometer data, so that hurricane-force winds can be studied statistically without being drowned out by the noise of scarce tail samples. The method smooths the upper tail with extreme value theory: above a high threshold, the 99.99th percentile at basin level, the tail is fitted with an exponential or generalised Pareto distribution, while below the threshold the data speak for themselves. Uncertainty is estimated by block-bootstrap resampling that keeps whole days, one-week clusters, and the seasonal cycle intact, preserving the spatial and temporal dependence of storm winds. Applied to ASCAT-A winds in the Caribbean and both tropical Atlantic basins over 2007-2021, the fitted and empirical-bootstrap percentiles agree down to a probability of exceedance of $10^{-5}$ (the 99.999th percentile) and remain consistent within uncertainties down to $10^{-6}$. This matters because a stable extreme-wind percentile estimator is the prerequisite for asking whether hurricane winds are changing over decades.","feed_headline":"Hurricane wind extremes now reach the 99.999th percentile","feed_subtitle":"A bootstrap-based smoother turns scarce satellite tail data into stable hurricane-wind percentiles.","key_machinery":"The engine is the extreme-value tail approximation: for a wind-speed distribution $F_X$, the displacement between the quantiles at exceedance probabilities $p$ and $p/\\lambda$, divided by a scale factor $a(p)$, converges as $p\\to 0$ to the functional form $(\\lambda^\\gamma-1)/\\gamma$ (the generalised Pareto/exponential family). The paper uses this to replace the noisy empirical tail above the 99.99th percentile with a smooth fitted tail, producing percentiles down to exceedance probabilities of $10^{-5}$ and $10^{-6}$. The companion mechanism is a block-bootstrap resampling scheme: whole days are resampled in approximately one-week blocks, with the number of days in each calendar month held fixed, so each synthetic dataset reproduces the spatial and temporal dependence of the original storm winds; fitting each resample separately and averaging over resamples yields both the percentile estimates and their uncertainties. This combination is what lets the paper claim consistency across fitted and non-fitted estimates.","core_discovery":"The paper's central claim is that basin-level extreme wind percentiles can be smoothed so effectively that they remain trustworthy far beyond the range where empirical percentiles are useful: down to the 99.999th percentile, corresponding to an exceedance probability of $10^{-5}$, and still consistent within uncertainties at $10^{-6}$. The evidence is that in three tropical basins, over ASCAT-A's 2007-2021 record, exponential tail fits, generalised Pareto fits, and raw block-bootstrap empirical percentiles all give nearly the same wind-speed values at these extreme levels, with the 99.99th percentile serving as the threshold. The same agreement holds for two ASCAT spatial resolutions, and the method is insensitive to the choice between fitting the tail or not fitting it at all. The authors frame this as a method paper: they are not claiming to detect a decadal trend, only to provide the reliable extreme-percentile estimates that trend detection will require over a longer, multi-instrument record.","pith_inferences":["A testable extension would pool several years of data and apply the method at sub-basin or pixel scales to see where the robustness breaks down as the tail sample shrinks; the paper deliberately works at basin level for exactly this reason.","The agreement between exponential, generalised Pareto, and raw bootstrap percentiles suggests the smoother is a general tool for sparse tail estimation, so it could transfer to other geophysical extremes such as extreme precipitation or ocean wave heights.","If the one-week block length is right, the reported uncertainties imply a minimum detectable trend over the 15-year ASCAT record; the paper does not compute it, but it could be estimated directly from the yearly percentile variances.","A quick way to test whether reanalysis input changes contaminate trend signals would be to run the method on a long single-model reanalysis alone: any trend that disappears when only the model is used would point to changing observations rather than changing winds."],"forward_implications":["Basin-level extreme-wind percentiles can be quoted reliably at the 99.999th percentile from a single instrument and a 15-year record, without masking or a strong distributional tail assumption.","The same protocol can be run separately on QuikSCAT and ERS scatterometer records and each compared against ERA5, instead of stitching instruments into one time series, which is the paper's proposed route toward decadal trend studies.","ERA5 model winds are systematically lower than ASCAT winds at these extreme percentiles, so any model-based trend comparison will need quantile mapping or CDF matching; the paper flags this as the planned next step.","The 99.99th percentile emerges as the practical threshold choice: moving it by a factor of ten in exceedance probability leaves the 99.999th percentile estimate stable provided the analysis stays conservative.","With a longer overlapping record of ERS, QuikSCAT, ASCAT, and the future SCA instrument, the method should be able to separate decadal trend signals from ENSO-scale variability."],"supporting_citations":[{"why":"Supplies the Generalised Pareto limit that one of the tail fits relies on.","marker":"Balkema and De Haan (1974)"},{"why":"Supplies the general extreme value theory convergence framework and the functional form of the limit used for tail smoothing.","marker":"Haan and Ferreira (2006)"},{"why":"Introduces the bootstrap resampling principle on which the uncertainty estimation rests.","marker":"Efron (1979)"},{"why":"Extends the bootstrap to dependent stationary data, justifying the block-bootstrap scheme.","marker":"Künsch (1989)"},{"why":"Earlier percentile-based extreme wind study that this method improves by going beyond the 99th percentile.","marker":"Giesen and Stoffelen (2022)"},{"why":"Defines ERA5, the collocated model reference dataset used for comparison.","marker":"Hersbach et al. (2020)"},{"why":"Validates ASCAT scatterometer wind quality in tropical hurricanes, supporting ASCAT as the observation basis.","marker":"Ni et al. (2022)"},{"why":"Documents long-term scatterometer stability, underpinning the proposed multi-instrument trend extension.","marker":"Stoffelen et al. (2021)"}],"fun_headline_variants":["Extreme hurricane winds pinned to 99.999th percentile","Smoothing method tames scarce hurricane wind tail data","Basin-level wind extremes stable at 99.999th percentile","Bootstrap smoother yields solid wind extremes at 10^-5"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method hangs on the block-bootstrap resampling of whole days in one-week blocks with monthly counts fixed capturing the real dependence structure of extreme winds, so the reported uncertainties are honest.","fun_headline_variants_meta":{"raw":{"variants":["Extreme hurricane winds pinned to 99.999th percentile","Smoothing method tames scarce hurricane wind tail data","Basin-level wind extremes stable at 99.999th percentile","Bootstrap smoother yields solid wind extremes at 10^-5"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1360,"prompt_tokens":1024,"completion_tokens":336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":265}},"tokens_in":640,"tokens_out":336,"duration_ms":3660,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:50:45.060292+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the entire basin-level analysis with block lengths of two, four, and eight weeks, and with block lengths informed by the autocorrelation of daily basin-maximum winds; if the spread of the resampled percentiles widens substantially and the exponential, generalised Pareto, and raw percentiles no longer agree within those wider uncertainties at the $10^{-5}$ exceedance probability, the central consistency claim fails.","supporting_citations":[{"cited_title":"A., and L","cited_arxiv_id":null,"evidence_quote":"Supplies the Generalised Pareto limit that one of the tail fits relies on."},{"cited_title":"Ferreira, 2006: Extreme value theory: an introduction, Vol","cited_arxiv_id":null,"evidence_quote":"Supplies the general extreme value theory convergence framework and the functional form of the limit used for tail smoothing."},{"cited_title":"Stoffelen, 2022: Changes in extreme wind speeds over the global ocean","cited_arxiv_id":null,"evidence_quote":"Earlier percentile-based extreme wind study that this method improves by going beyond the 99th percentile."}],"review_version":1}