{"id":"0f471f98-8a54-4b2d-b571-5b725e92637d","arxiv_id":"2411.15819","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper proves weak convergence of a bivariate sequential tail empirical process and its bootstrap, and uses it to derive asymptotically valid tests for equal extreme value indices, equal scedasis functions, and constant tail copulas.","lead":"This paper develops statistical theory for extreme events when the probability of an extreme and the relationship between two variables both shift over time. It provides a bootstrap method, with proven large-sample guarantees, for testing whether tail heaviness, tail-probability patterns, or tail dependence are changing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Null hypothesis H30 in (5.3) is mis-stated: formal H30 uses R' where the intended null uses R, so Proposition 3(c) is false as written and the third test's nominal level is not established.","rationale":"The paper makes a credible contribution: the B-STEP framework and its bootstrap are natural extensions of Einmahl et al. (2014) and Bucher & Dette (2013), and the delta-method derivations in Propositions 1-2 are standard. The reader's weakest assumption, Assumption 2, is a typical uniform-rate condition in EVT; although strong and uncheckable, it is the kind of assumption that can be weakened in specific models and is not internally inconsistent. In contrast, the null hypothesis H30 stated in (5.3) is an internal mathematical error: as written, it asserts R'=R(·,·,1), which is not the intended condition 'R does not depend on z'. Because the test statistic T30 is explicitly constructed to cancel the scedasis factor C(z), the formal statement contradicts the test's design and the derived limit. Proposition 3(c) is therefore false under the literal H30. This is a load-bearing flaw in a central application of the paper, and it is verifiable by direct computation and simulation. The reader's other points (Table 1 mislabeling, tail-independence overreach, missing code) are also valid but less decisive. I therefore recommend keeping the CONDITIONAL verdict, with the requirement that the authors correct the null hypothesis and clarify the intended meaning, or rigorously establish Proposition 3(c) under the stated condition. The central convergence theorem may be sound, but the paper's formal statements about the non-changing tail copula test need revision.","tokens_in":18682,"tokens_out":26323,"duration_ms":217912,"concrete_test":"Re-derive the limit of √k T30(z) under the null as literally stated in (5.3). If the limit is not the centered Gaussian process used to compute critical values, Proposition 3(c) is invalid. Also run a simulation with c1=c2=a3 (non-constant scedasis) and a constant tail copula (C5). If the stated H30 is used to define the rejection region, the Type I error should exceed nominal level because the statistic diverges; if H30 is instead interpreted as 'R constant in z', sizes should be near nominal. This distinguishes the two hypotheses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Under the intended null that the tail copula R(x,y,z) does not depend on z, with c1=c2=c and R homogeneous of order 1 (as in Examples 1-2), R'(x,y,z)=∫_0^z R(c(t)x,c(t)y,1)dt=R(x,y,1)C(z). Hence R' equals R(·,·,1) only when C(z)=z for all z, i.e. c≡1. The formal H30 in (5.3) states R'(x,y,z)=R(x,y,1) for all (x,y,z)∈D_T, which is stronger than non-changing R and additionally imposes c≡1. Under this literal null, T30(z) → 2C(z)-2 in probability, so √k T30(z) diverges except when C(z)=z, and the KS/CVM rejection probability tends to 1. Thus Proposition 3(c) is false if H30 is read literally. The correct null should be R(x,y,z)=R(x,y,1) for all z (or equivalently R'(x,y,z)=R(x,y,1)C(z)). This is not merely typographical: the test statistic's centering depends on the intended null, and a reader following the stated hypothesis would get invalid inference. The asymptotic limit claimed in §5.3 is derived for the intended null, not for the stated one.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a bivariate extension of heteroscedastic extremes for independent but non-identically distributed data, allowing both the marginal tails (via scedasis functions) and the copula (via a changing tail copula R(x,y,z)) to vary across the sample. The main theoretical object is a bivariate sequential tail empirical process (B-STEP) and its weighted bootstrap counterpart. The paper states weak convergence of both processes to the same Gaussian limit (Theorem 1), derives asymptotic distributions for the quasi-tail copula, integrated scedasis functions, and Hill estimators as functionals of the process (Theorems 2 and 3), and constructs bootstrap tests for equal extreme value indices, equal scedasis functions, and a constant tail copula under identical scedasis functions, together with a simulation study. The central proofs are deferred to the supplementary material, and the formal statement of the third null hypothesis in (5.3) is inconsistent with the intended null, which affects Proposition 3(c) as written.","tokens_in":18962,"tokens_out":13016,"duration_ms":114912,"significance":"The proposed framework is a natural and potentially useful unification of two strands of recent EVT literature (heteroscedastic margins and changing dependence), and the process-centric bootstrap approach is attractive because it avoids explicit estimation of complicated covariance structures. The paper also provides readily implementable tests and supplies simulation evidence. Its main strengths are the unifying B-STEP convergence result, the bootstrap construction that jointly resamples margins and copulas, and the explicit treatment of three testing problems. The main limitations are that the proofs of Theorems 1-3 are not in the posted text, so the central derivations could not be independently verified, and that the displayed null hypothesis in (5.3) does not correspond to the null actually used in the test statistic and simulations. These issues are correctable, but they need to be fixed before the paper can be accepted.","major_comments":[{"comment":"The formal null hypothesis in (5.3) is R'(x,y,z) = R(x,y,1) for all (x,y,z) in D_T. Under the intended null (no change in the tail copula) and with c1=c2, the homogeneity of R gives R'(x,y,z) = \\int_0^z R(c(t)x,c(t)y,t) dt = R(x,y,1) C(z), where C(z) is the common integrated scedasis function. Since C(z)<1 for z<1 when c is positive and C(1)=1, the displayed equality R'(x,y,z)=R(x,y,1) cannot hold for all z unless c is identically 1, which is far stronger than a constant tail copula. Under the literal H30, T30(z) converges in probability to 2C(z)-2, so sqrt(k) T30(z) diverges for z<1 and the rejection probability tends to 1, contradicting Proposition 3(c). The simulation evidence in Table 3 is consistent with the intended null (constant R, arbitrary c1=c2) rather than the displayed H30, since the rows with a1,a1,C5 and C6 have non-constant scedasis. The intended null should be stated as R(x,y,z)=R(x,y,1) for all z (equivalently, R'(x,y,z)=R(x,y,1)C(z) when C1=C2). This is not a purely typographical issue: the centering of T30 is derived from the intended null, and a reader implementing the displayed H30 will obtain invalid inference.","section":"§5.3, Eq. (5.3) and Proposition 3(c)"},{"comment":"The main convergence results are stated without proof: Theorem 1 (weak convergence of the B-STEP and its bootstrap version), Theorem 2, and Theorem 3 are all deferred to the supplementary material. Since Proposition 3 and the simulation interpretations rest on these theorems, the posted text is not self-contained. In particular, the bias bound described in Remark 3 depends on Assumption 3 in a way that is not demonstrated in the main text, and the conditional weak convergence in (4.1) is only defined, not established. I could not verify from the posted material that Assumptions 1-5 indeed imply (4.3). Please ensure the supplementary proofs are included in the review package and are complete.","section":"§4, Theorems 1-3 and Remark 3"}],"minor_comments":[{"comment":"The text states that rows 1-4 of Table 1 have equivalent EVIs, but row 4 has lambda1=2.5 and lambda2=2, which differ; the text later correctly refers to row 4 as an alternative case. The null rows are rows 1-3, and this inconsistency should be corrected.","section":"§5.4, Table 1"},{"comment":"The weight function q in (3.1) allows eta=0, but Assumption 5 requires E|1-xi|^{1/eta}<infty, which is undefined for eta=0. Either restrict eta>0 in (3.1) or state that the moment condition is vacuous when eta=0.","section":"Assumption 5, Eq. (3.1)"},{"comment":"The text says 'Weiner process'; this should be 'Wiener process'. Moreover, the process W defined by the covariance in (4.2) is not a classical Wiener process on the product space but a centered Gaussian process with that covariance, so the terminology should be adjusted.","section":"§4, before Eq. (4.2)"},{"comment":"The paragraph discussing R(1,1,z)=0 is outside the scope of Assumption 2, which requires R(1,1,z)>0. If the authors intend to cover asymptotic tail independence, this needs to be stated as a separate assumption or clearly marked as a heuristic discussion.","section":"§5.1"},{"comment":"The limiting statement 'sup_{z in [0,1]} sqrt(k)|T30(z)| -> sup_{z in [0,1]} |{...} B(z)|' is asserted without derivation. Given the covariance structure in (4.2) and the process limits in Theorem 2, this is not immediate and should be proved or explicitly referenced in the supplementary material.","section":"§5.3"}],"recommendation":"major_revision","confidential_remarks":"The H30 misstatement is the main substantive issue; the authors should be asked to restate the null hypothesis in (5.3) and to check that all subsequent formulas, including Proposition 3(c) and the simulations, refer to the corrected null. Because the central theorems are only in the supplementary material, I could not fully check the proofs; please make sure the supplement is included in the review package. The paper's novelty relative to Drees (2023) and Einmahl and Zhou (2024) should also be sharpened in the revision, especially the comparison of convergence rates for the changing-copula test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know upfront. The paper does something real: it builds a bivariate sequential tail empirical process for independent, non-identically distributed extremes, lets the tail copula change across the sample, and proves weak convergence for the process and its bootstrap. That unifies marginal heterogeneity (scedasis) with dependence heterogeneity, which Einmahl-Zhou and Drees treat separately. The bootstrap tests for equal EVI, equal scedasis, and constant tail copula follow as functionals, and the simulations look reasonable.\n\nThe main soft spot is the formal statement of the third null. In (5.3), H30 is written as R'(x,y,z) = R(x,y,1) for all (x,y,z). That cannot be the intended null. Under a tail copula that does not depend on z, R'(x,y,z) = R(x,y,1) C(z) when the two scedasis functions are equal. So the stated H30 forces C(z) = z, i.e., c ≡ 1, and even then it only holds at z = 1. The test statistic itself is centered for the correct null — the ratio R'(k1/k, k2/k, z)/R'(k1/k, k2/k, 1) cancels the C(z) — so the test likely works as intended. But the hypothesis has to be restated as R(x,y,z) = R(x,y,1) for all z. As written, Proposition 3(c) is false because the condition is overstrong; a literal reader would get the wrong limit. That is a fixable but real error.\n\nThe other wart is Table 1: the text says rows 1–4 are equivalent EVI, but row 4 has lambda1 = 2.5 and lambda2 = 2. The same row is later treated as an alternative. That needs cleaning up.\n\nI could not verify the core theorems: the proofs are in the supplementary material, not posted. So my confidence rests on the standard nature of the assumptions and the examples. Assumption 2's uniform polynomial rate is untestable but typical for this area.\n\nThe paper deserves a serious referee. The framework is a genuine step forward for heteroscedastic extremes, and the bootstrap results are useful. A referee should insist on a corrected H30, a fixed Table 1, and ideally a supplement that can be checked. I would send it to review, not desk-reject.","headline":"A solid, genuinely new bivariate tail-process framework; the third null hypothesis is mis-stated and must be corrected, and the proofs are in the supplement.","tokens_in":19477,"tokens_out":7477,"would_cite":true,"duration_ms":62188,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","62G09","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single empirical process makes bootstrap inference valid for bivariate extremes whose marginal tails and tail dependence both change.","keywords":["extreme value theory","heteroscedastic extremes","changing tail copula","bivariate sequential tail empirical process","bootstrap","scedasis function","Hill estimator","functional delta method"],"falsifier":"Simulate a triangular array whose copulas approach the tail limit only logarithmically, for example $t C_{n,i}(x/t,y/t)=R(x,y,i/n)+(\\log t)^{-1}$, and check whether the bootstrap tests still hold their nominal level as the sample size grows; this violates Assumption 2's uniform $t^{-\\alpha}$ rate, so the central convergence should break.","tokens_in":18467,"feed_emoji":"📈","tokens_out":6848,"duration_ms":58788,"temperature":0.7,"pith_summary":"This paper tries to establish that one process, the bivariate sequential tail empirical process (B-STEP), carries valid bootstrap inference for data whose marginal tail heaviness and dependence between extremes both change across observations. The authors prove that the B-STEP and its multiplier-bootstrap version converge weakly to the same Gaussian process, then use the functional delta method to derive asymptotic normality for the quasi-tail copula, integrated scedasis functions, and Hill estimators. From these results, they build bootstrap tests for equal extreme value indices, equal scedasis functions, and an unchanging tail copula, and prove that the rejection probabilities converge to the nominal significance level. The payoff is simultaneous inference on marginal tail risk and changing dependence without needing to know the complicated asymptotic variances, which matters for climate and financial data where both features drift together.","feed_headline":"Bootstrap tames heteroscedastic extremes with shifting tail copulas","feed_subtitle":"A unified empirical process makes bootstrap inference valid when marginal heaviness and tail dependence both drift.","key_machinery":"The B-STEP is the weighted empirical process $F_n(x,y,z) = (\\tilde R'(x,y,z)-R'(x,y,z))/q(x,y)$, where $\\tilde R'$ counts exceedances of high marginal thresholds up to time $\\lfloor nz\\rfloor$, the quasi-tail copula $R'(x,y,z) = \\int_0^z R(c_1(t)x,c_2(t)y,t)\\,dt$ is its expectation limit, and the weight $q(x,y)=(x\\vee y)^\\eta$ with $0\\le\\eta<1/2$ tames the tails. The bootstrap B-STEP replaces the counting indicators with $\\xi_{bi}$-weighted indicators, where the $\\xi_{bi}$ are iid positive multipliers with mean and variance one, so that margins and copula are resampled jointly. The proof routes through the functional delta method, using the Hadamard-differentiable maps $\\Phi$ for the quasi-tail copula and $\\Psi$ for the Hill estimator, making the process convergence transfer to all three estimators and their bootstrap counterparts.","core_discovery":"The central discovery is a functional limit theorem for non-identically distributed bivariate extremes: under Assumptions 1-5, the normalized B-STEP, $\\sqrt{k}F_n$, converges weakly to $W/q$, and its bootstrap version $\\sqrt{k}F_n^b$ converges conditionally given the data to the same $W/q$ in $\\ell^\\infty(D_T)$, where $W$ is a Gaussian process with covariance $R'(x_1\\wedge x_2, y_1\\wedge y_2, z_1\\wedge z_2)$. This single convergence result is the engine of the paper. From it, the quasi-tail copula estimator, the integrated scedasis estimators, and the Hill estimators are shown to be asymptotically normal with explicit Gaussian limits, and the bootstrap versions of these estimators are asymptotically valid. Consequently, bootstrap-based Kolmogorov-Smirnov and Cramér-von Mises tests for equal extreme value indices, equal scedasis functions, and constant tail dependence have rejection probabilities converging to the nominal level as the sample size and number of bootstrap replications diverge.","pith_inferences":["The same multiplier-bootstrap B-STEP could produce simultaneous confidence bands for the quasi-tail copula and for differences of scedasis functions, since the limiting Gaussian covariance is explicit.","A natural next step, which the authors leave open, is extending the process convergence to weakly dependent triangular arrays; the process-centric proof structure suggests block-multiplier versions would be the route.","Because the key convergence-rate condition on the true copula sequence is uncheckable, a practical diagnostic comparing bootstrap-based quantiles across several intermediate orders $k$ could reveal whether that rate assumption is credible for a given dataset."],"forward_implications":["Theorem 1 implies that every Hadamard-differentiable functional of the joint tail, not just the three estimators written out, can be bootstrapped under this model of changing marginal and dependence heterogeneity.","The three bootstrap tests are asymptotically of the nominal level; the simulations show size control at $k=200$ and rejection frequency rising in $k$ under alternatives.","For the equal-scedasis test, the bootstrap is necessary because $\\hat C_1-\\hat C_2$ does not converge to a Brownian bridge when the copula changes across samples; the bootstrap handles the unknown covariance.","For the non-changing-tail-copula test with identical scedasis functions, the proposed statistic converges to a Brownian bridge limit at root-$k$ rate, faster than integrated angular measures, so less data is needed than in the earlier test it generalizes.","When the two margins are asymptotically tail independent, the two Hill estimators become asymptotically independent, so the test for equal extreme value indices remains valid in that boundary case."],"supporting_citations":[{"why":"Introduces the univariate sequential tail empirical process and the scedasis function that the B-STEP extends to two dimensions.","marker":"Einmahl et al. [2014]"},{"why":"Supplies the multiplier-bootstrap construction for tail copulas and the derivative conditions on the tail copula used in Assumption 2.","marker":"Bücher and Dette [2013]"},{"why":"Provides the heteroscedastic-extremes model with a common tail copula that this paper generalizes to a changing tail copula, and the scedasis test that breaks under varying copulas.","marker":"Einmahl and Zhou [2024]"},{"why":"Establishes the baseline test for a changing extreme-value dependence structure; the paper compares its own test statistic and convergence rate to this approach.","marker":"Drees [2023]"},{"why":"Supplies the weak-convergence and Hadamard-differentiability framework used for the functional delta method and conditional bootstrap.","marker":"van der Vaart and Wellner [1996]"},{"why":"Defines the conditional weak convergence used to state the bootstrap validity results.","marker":"Kosorok [2003]"}],"fun_headline_variants":["Bootstrap tames shifting tail copulas in extremes","When tails drift, bootstrap inference stays valid","Functional limit theorem for changing tail dependence","Bootstrap tests pass when marginal heaviness changes","Heteroscedastic extremes: bootstrap still holds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The true dependence between extreme events must converge to its limiting tail copula at a uniform polynomial rate across all observations and sample sizes, a condition that cannot be checked from the data and whose failure would break the central convergence theorem.","fun_headline_variants_meta":{"raw":{"variants":["Bootstrap tames shifting tail copulas in extremes","When tails drift, bootstrap inference stays valid","Functional limit theorem for changing tail dependence","Bootstrap tests pass when marginal heaviness changes","Heteroscedastic extremes: bootstrap still holds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1627,"prompt_tokens":921,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":636}},"tokens_in":537,"tokens_out":706,"duration_ms":6848,"temperature":1.0,"reasoning_tokens":636,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:54:01.646259+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a triangular array whose copulas approach the tail limit only logarithmically, for example $t C_{n,i}(x/t,y/t)=R(x,y,i/n)+(\\log t)^{-1}$, and check whether the bootstrap tests still hold their nominal level as the sample size grows; this violates Assumption 2's uniform $t^{-\\alpha}$ rate, so the central convergence should break.","supporting_citations":[],"review_version":1}