Pith. sign in

REVIEW 3 major objections 5 minor 53 references

Adaptive tail index estimation: minimal assumptions and non-asymptotic guarantees

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Under only the minimal heavy-tail assumption, a calibration-free adaptive rule selects the Hill estimator's k with error within a small constant of the oracle, and with nearly minimax rates under a von Mises condition.

desk verdict A genuinely new constant-free adaptive Hill rule with real finite-sample bounds under plain regular variation, but the unverifiable grid condition and the simulation confidence choice mean the minimal-assumptions claim is not quite as strong as advertised. read the letter →

arxiv 2505.22371 v2 pith:NDENHOYK submitted 2025-05-28 stat.OT math.STstat.TH

classification stat.OTmath.STstat.TH MSC 62G3262G0560G70
keywords tailindexestimationHillestimatoradaptivevalidationextremevaluetheoryregularvariationnon-asymptoticguaranteesminimaxratevonMisescondition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

For heavy-tailed data, the accuracy of the Hill estimator for the tail index $\gamma$ is governed by $k$, the number of largest observations retained, and data-driven rules for choosing $k$ have so far relied on second-order or von Mises assumptions that are hard to verify in practice. This paper introduces an 'Extreme Adaptive Validation' (EAV) rule that selects $k$ through a stopping criterion built from an explicit variance bound alone; the unknown bias term is never estimated or controlled directly. It proves that, under only the minimal assumption that the survival function is regularly varying, the selected $k$ yields an error bounded by a small constant times the error of an oracle choice that balances bias against variance. Under an additional von Mises condition, the rule reaches the nearly minimax rate $\sqrt{\log\log n}\, n^{-|\rho|/(1+2|\rho|)}$ on logarithmic grids of candidate $k$ values, improving on the $(\log\log n / n)^{|\rho|/(1+2|\rho|)}$ rate of earlier non-asymptotic adaptive methods. The only user input is a confidence level and a default grid.

What carries the argument

The load-bearing decomposition, valid for every regularly varying survival function, is the almost-sure error bound of Lemma 2: $|\hat\gamma(k) - \gamma| \le \gamma|Z_k - 1| + 2\bar a\bigl((1-U_{(k+1)})^{-1}\bigr) + \bar b\bigl((1-U_{(k+1)})^{-1}\bigr) Z_k$, obtained from the Karamata representation $Q(t) = A t^\gamma \exp\bigl(\int_{t_0}^{t} b(u)/u\, du\bigr)$ of the quantile function. Because $Z_k$ is the mean of $k$ unit exponentials, $|Z_k - 1|$ has a known sub-gamma quantile that serves as the explicit variance function $V(k,\delta) \le \sqrt{2\log(4/\delta)/k} + \log(4/\delta)/k$; the two remaining terms, built from the monotone Karamata remainders $\bar a$ and $\bar b$, form the bias term $B(k,n,\delta)$ that is only required to be non-decreasing in $k$ and never needs to be computed. The Extreme Adaptive Validation rule then selects the largest grid point $k$ such that all smaller grid points $j$ satisfy $|\hat\gamma(k) - \hat\gamma(j)| \le \hat\gamma(k)\,(1-2V(k,\delta/|\mathcal{K}|))^{-1}\,(V(j,\delta/|\mathcal{K}|)+3V(k,\delta/|\mathcal{K}|))$, a criterion involving only the selected estimate, the explicit variance, and the grid. The grid is logarithmic (size $\ll n$), which is what practitioners already use and what makes the adaptivity cost only a $\sqrt{\log\log n}$ factor.

What would settle it

Two checkable tests: (1) coverage — simulate Fr\'echet samples with known $\gamma$ and $\rho$, run EAV with $\delta = 0.9$ over several hundred replications at $n = 10^4, 10^5, 10^6$, and record how often the Theorem 1 bound (evaluated with the true $k^*$) contains $\gamma$; a coverage rate persistently below $1-\delta$ would refute the main claim. (2) rate — for the same model, the empirical error of $\hat\gamma(\hat k_{\mathrm{EAV}})$ must decay at least as fast as $\sqrt{\log\log n}\, n^{-|\rho|/(1+2|\rho|)}$; a measured exponent closer to $n^{-|\rho|/(2+2|\rho|)}$ would refute Corollary 3.

Watch

Extended reading notes

Core claim

The paper's central claim is that adaptive selection of the extreme sample size $k$ for the Hill estimator reduces to a simple stopping rule built on an explicit variance term, with the bias term absorbed through monotonicity rather than estimated. On an event of probability at least $1-\delta$, the adaptive estimate satisfies $|\hat\gamma(\hat k_{\mathrm{EAV}}) - \gamma| \le \bigl(4\hat\gamma(\hat k_{\mathrm{EAV}})/(1-2V(\hat k_{\mathrm{EAV}}, \delta/|\mathcal{K}|)) + 2\gamma\bigr)\, V(k^*(\delta/|\mathcal{K}|, n), \delta/|\mathcal{K}|)$, where $V(k,\delta)$ is a quantile of $|Z_k - 1|$ with $Z_k$ the mean of $k$ independent unit exponentials, and $k^*$ is the largest grid point at which the bias term still stays below the variance term. Under the minimal regular-variation assumption, this bound is within a small constant of the oracle error $2\gamma V(k^*)$, so the rule is 'good enough' rather than certificate-optimal. Under the von Mises condition the oracle is provably large enough to convert this bound into the nearly minimax rate $\sqrt{\log\log n}\, n^{-|\rho|/(1+2|\rho|)}$. The theory is stated abstractly for any estimator admitting a bias-variance decomposition error $\le \gamma V + B$ with $V$ explicit and non-increasing and $B$ non-decreasing, then instantiated for the Hill estimator.

Load-bearing premise

The whole guarantee rests on the grid of candidate $k$ values strictly straddling the unknown bias-variance balance point (Condition 2); if no grid point has bias above variance while a lower grid point has bias below it, the oracle $k^*$ and every subsequent error bound are undefined, and no data-based test can certify the condition.

Editorial extensions

If this is right

  • Under only the regular-variation assumption (1.1), EAV gives a finite-sample guarantee within a small constant of the oracle error, eliminating the need to estimate second-order parameters.
  • Under the von Mises condition, the selected $k$ attains error of order $\sqrt{\log\log n}\, n^{-|\rho|/(1+2|\rho|)}$, faster than the $(\log\log n / n)^{|\rho|/(1+2|\rho|)}$ rate of earlier non-asymptotic adaptive rules.
  • Only a logarithmic grid of $k$ values is searched, so the computational cost is $O((\log n)^2)$ rather than $O(n^2)$, at no practical cost in performance.
  • The Section 2 framework transfers to any tail estimator whose error admits the decomposition error $\le \gamma V + B$, so EAV-style selection can be reused beyond the Hill estimator.
  • Simulations show lower mean squared error than the two benchmark adaptive rules on ill-behaved distributions, including a new counter-example that is regularly varying but admits no standardized Karamata representation, and comparable performance on well-behaved tails.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The unverifiable Condition 2 suggests a practical stability diagnostic that the paper does not develop: run EAV over several grids and confidence levels and inspect how $\hat k_{\mathrm{EAV}}$ moves, since the experiments show robustness to grid geometry but sensitivity to $\delta$.
  • Because the general framework only needs an explicit non-increasing variance bound plus a monotone bias, EAV selection should transfer directly to high-quantile estimation and to dependence functionals of multivariate extremes, where similar non-asymptotic bounds now exist.
  • The sample-size condition $n_0(\delta)$ is made explicit only under the von Mises condition, yet simulations already work at $n = 1{,}000$; a natural check is to tabulate coverage of the Theorem 1 bound as a function of $n$ to see how much slack the sufficient conditions carry.
  • The new counter-example distribution (regularly varying but failing the standardized Karamata representation) is a sharper stress test than the usual Fr\'echet, stable, and Pareto change-point families, and could serve as a standard benchmark for future adaptive rules.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes an adaptive selection rule, called Extreme Adaptive Validation (EAV), for choosing the number k of upper order statistics used by the Hill tail-index estimator. The rule is defined on a grid of candidate values and its stopping criterion uses only explicit quantiles of a centered Gamma distribution, not an estimated bias term. The main theoretical result is an oracle inequality: under a grid-straddling condition (Condition 2) and a sample-size condition n ≥ n0(δ), the EAV choice has error bounded by a constant multiple of the error of an oracle k* that balances an explicit variance term and an unknown bias term. Under the von Mises Condition 3, the authors derive a lower bound on k*, an explicit bound on n0(δ), and a near-minimax rate of order sqrt(log log n) n^{-|ρ|/(1+2|ρ|)}. The paper also contains simulations comparing EAV with the methods of Boucheron & Thomas (2015) and Drees & Kaufmann (1998) on a range of heavy-tailed distributions.

Significance. If the results are correct, the paper is a useful contribution: it offers a transparent, calibration-free adaptive method for a central problem in extreme value theory, with the core oracle inequality requiring only regular variation. The argument leading to Theorem 1 is careful, the stopping rule is explicit, and the paper provides reproducible code. I do not see a circularity problem: the EAV rule uses only the variance-type term V and does not estimate the bias, so the comparison with an internal oracle is not circular. However, the advertised 'minimal assumptions' claim is currently overstated. The main theorem is conditional on an unverifiable grid-straddling condition and on a sample-size threshold n0(δ) that is not explicit under regular variation alone; moreover, one displayed bias bound appears to be stated incorrectly, and the displayed oracle rate in Theorem 3 does not match the proof. These issues are substantial but local, and I believe they can be fixed within the manuscript's scope.

major comments (3)
  1. [Theorem 2 (Section 3.3), Eq. (3.14)] The bias term B(k,n,δ) in Eq. (3.14) is defined with R(1,δ/2), while the derivation in Appendix A.2, Eq. (A.1), contains R(k+1,δ/2). Since R(k,δ) = sqrt(3 log(1/δ)/k) + 3 log(1/δ)/k is decreasing in k, one has R(1,δ/2) ≥ R(k+1,δ/2). Consequently the argument n/(k+1)(1+R(1,δ/2)) is larger than n/(k+1)(1+R(k+1,δ/2)), and because ā and b̄ are non-increasing, the displayed B is smaller than the actual upper bound on the bias derived in the appendix. As stated, Theorem 2 does not establish the bias-variance decomposition required by Condition 1, and the oracle in (2.2) could be defined from a bias bound that is too small. Please replace R(1,δ/2) by R(k+1,δ/2) in (3.14), or otherwise provide a valid upper bound on the bias.
  2. [Condition 2 (Section 2.1)] The statement that Condition 2 is automatically satisfied for grids with kmin = 1 and kmax = n for large n is not correct in general. For an exact Pareto distribution, the Karamata functions in (3.1) can be chosen as a(t) ≡ 1 and b(t) ≡ 0, so the bias term in Theorem 2 is identically zero and the requirement B(kmax,n,δ) > γV(kmax,δ) fails for every n. Furthermore, the second inequality of Condition 2 is not used in the proofs of Lemma 1, Proposition 3, or Theorem 1; those proofs only need the set in (2.2) to be non-empty. In addition, outside the von Mises Condition 3, the threshold n0(δ) in (2.5) is defined through the unknown bias and no data-checkable certificate for n ≥ n0(δ) is provided. The abstract's claim that the analysis is valid 'for all heavy-tailed distributions' should therefore be qualified, and Condition 2 should either be weakened or its role in the main theorem should be clarified.
  3. [Theorem 3 (Section 4.3) and Appendix A.5] The displayed oracle bound in Theorem 3 contains the factor sqrt(C2(ρ)), whereas the proof in Appendix A.5 obtains C2(ρ)^{-1/2}. The intermediate line in the proof, '≤ 2(1+√2) γ √β sqrt(1+log(4/δ)) γ^{-1/(1-2ρ)} C2(ρ)^{-1/2} n^{ρ/(1-2ρ)}', is the correct expression. This is not a purely cosmetic mismatch: if C2(ρ) < 1, the printed bound is smaller than what the proof supports, making the oracle error bound overly optimistic. Please correct Theorem 3 and check that Corollary 3 uses the same convention.
minor comments (5)
  1. [Proposition 3 (Section 2.2)] The statement quantifies over all (j,k) in {1,...,k*(δK,n)}, but the proof and the EAV rule (2.6) require j,k ∈ K. Please add the restriction j,k ∈ K to the statement.
  2. [Remark 4 (Section 3.3)] The chain '22(1+√2)^2 log(4|K|/δ) ≤ 36 log(4|K|/δ)' is arithmetically false, since 22(1+√2)^2 ≈ 128.2. Either the constant in the derivation or the displayed inequality should be corrected.
  3. [Appendix A.2] The sentence saying that R is a non-decreasing function of k is inconsistent with the definition R(k,δ) = sqrt(3 log(1/δ)/k) + 3 log(1/δ)/k, which decreases with k; this is related to the first major comment and should be fixed in the same revision.
  4. [Section 5] The text refers to 'Drees et al. (1998)' and to 'Family 7', while the reference list has Drees & Kaufmann (1998) and the enumeration describes six distribution families; please unify these references.
  5. [Theorem 2 (Section 3.3), Eq. (3.14)] The formula for B(k,n,δ) has an unmatched parenthesis in the b̄ term; it should read b̄( n/(k+1)(1+R(1,δ/2)) ).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the EAV guarantee is an oracle inequality derived directly from the algorithm's explicit stopping threshold and a bias-variance decomposition with an explicit Gamma-quantile variance; self-citations are provenance only.

full rationale

The derivation chain is self-contained. Theorem 2 verifies Condition 1 for the Hill estimator under (1.1) with an explicit variance term V (a quantile of |Z_k - 1|) and a non-explicit bias term B bounded via Karamata functions. The oracle k* is defined as the largest grid point with B(k) <= gamma V(k); Propositions 1 and 2 are immediate consequences of this definition and monotonicity, not fitted claims. Theorem 1's adaptive bound follows algebraically from the stopping rule (2.6): because S(k_hat) = 0 and k_hat >= k* on the favorable event, |gamma_hat(k_hat) - gamma_hat(k*)| is bounded by the algorithm's own threshold, which is at most 4 gamma_hat V(k*)/(1-2V(k_hat)) + 2 gamma V(k*). This is a direct oracle inequality, not a prediction of an independently fitted quantity. The near-minimax rate under Condition 3 uses external lower bounds (Carpentier-Kim; Boucheron-Thomas) and standard Karamata/sub-gamma inequalities; no uniqueness theorem or ansatz is imported from the authors' prior work. Self-citations (Lederer 2022; Chichignoud et al. 2016; Li & Lederer 2019; Taheri et al. 2023) are provenance for the AV idea, with the proofs reproduced in this paper. A genuine limitation, which is a correctness concern rather than circularity, is that Condition 2 is not data-checkable under (1.1) and n0(delta) in (2.5) is explicit only under Condition 3; the headline guarantee for all regularly varying distributions must be read conditionally on the existence of the oracle k*. But no step equates an output with an input by construction.

Assumptions & free parameters 2 free parameters · 8 assumptions · 0 invented entities

The method itself introduces no invented entities and fits no parameters to the data; the only user-chosen constants are δ and the grid ratio. The theoretical guarantees rely on regular variation as the domain assumption, on unverifiable Condition 2 in the general case, and, for the sharper rates, on von Mises Condition 3. The bias function B is never estimated, so no free parameter is hidden in the stopping rule. The constants in Section 4 depend on unknown distributional quantities ρ and C, but these enter only the theoretical rates, not the algorithm.

free parameters (2)
  • confidence level δ = 0.9 in simulations
    Sole tuning parameter of EAV; affects k0 and the size of V. Hand-chosen, not estimated from data.
  • geometric grid ratio β = 1.1 in simulations
    Sets grid K and |K|; enters Condition 4 and constants in Section 4. User-chosen default.
assumptions (8)
  • domain assumption Regular variation of the survival function: F̄(tx)/F̄(t) → x^{-1/γ} as t→∞ (Eq. 1.1).
    Maintains the heavy-tailed domain; Theorem 2 and Corollary 1 are proved under this condition.
  • domain assumption Condition 1: a bias-variance decomposition |γ̂(k)-γ| ≤ γV(k,δ)+B(k,n,δ) with monotone V and B.
    Section 2 states the framework for any estimator satisfying Condition 1. For the Hill estimator it is proved as Theorem 2, so in the main application it is a theorem rather than an assumption.
  • domain assumption Condition 2: sufficiently wide grid, B(kmin)≤γV(kmin) and B(kmax)>γV(kmax).
    Defines the oracle k* in (2.2); cannot be verified from data in general and is needed for all guarantees.
  • domain assumption n ≥ n0(δ), the minimal sample size such that k0(δ) ≤ k*(δK,n).
    Used in Corollary 1; n0 is unknown under minimal assumptions and only explicitly bounded under Condition 3 in Corollary 2.
  • domain assumption Condition 3 (von Mises): Q(t)=At^γ exp(∫_{t0}^t b(u)/u du) with b̄(t) ≤ C t^ρ, ρ<0.
    Used in Section 4 to lower bound k* and derive near-minimax oracle and adaptive rates.
  • domain assumption Condition 4: grid ratio k_{m+1}/k_m ≤ β for some β>1.
    Used in Proposition 4 to relate k* to the next grid point; geometric grids with fixed β satisfy it.
  • standard math Karamata representation (3.1), Renyi representation of exponential spacings, sub-gamma concentration bounds.
    Standard tools invoked in Section 3 and Appendix A; accepted without proof.
  • standard math Zhang and Zhou (2020) lower tail bounds for Gamma variables.
    External result used in Lemma 5 to lower-bound the variance quantile V(k,δ).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adaptive tail index estimation: minimal assumptions and non-asymptotic guarantees." pith.science (2026). https://pith.science/paper/NDENHOYK

@misc{pith2026250522371,
  author       = {Pith},
  title        = {Pith review of: Adaptive tail index estimation: minimal assumptions and non-asymptotic guarantees},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NDENHOYK}},
  note         = {Machine review of arXiv:2505.22371}
}
abstract

A notoriously difficult challenge in extreme value theory is the choice of the number $k\ll n$, where $n$ is the total sample size, of extreme data points to consider for inference of tail quantities. Existing theoretical guarantees for adaptive methods typically require second-order assumptions or von Mises assumptions that are difficult to verify and often come with tuning parameters that are challenging to calibrate. This paper revisits the problem of adaptive selection of $k$ for the Hill estimator. Our goal is not an `optimal' $k$ but one that is `good enough', in the sense that we strive for non-asymptotic guarantees that might be sub-optimal but are explicit and require minimal conditions. We propose a transparent adaptive rule that does not require preliminary calibration of constants, inspired by `adaptive validation' developed in high-dimensional statistics. A key feature of our approach is the consideration of a grid for $k$ of size $ \ll n $, which aligns with common practice among practitioners but has remained unexplored in theoretical analysis. Our rule only involves an explicit expression of a variance-type term; in particular, it does not require controlling or estimating a biasterm. Our theoretical analysis is valid for all heavy-tailed distributions, specifically for all regularly varying survival functions. Furthermore, when von Mises conditions hold, our method achieves `almost' minimax optimality with a rate of $\sqrt{\log \log n}~ n^{-|\rho|/(1+2|\rho|)}$ when the grid size is of order $\log n$, in contrast to the $ (\log \log (n)/n)^{|\rho|/(1+2|\rho|)} $ rate in existing work. Our simulations show that our approach performs particularly well for ill-behaved distributions.

Figures

Figures reproduced from arXiv: 2505.22371 by the authors.

Figure 1
Figure 1. Monte-Carlo estimates of the standardised RMSE of Hill estimators as a function of the [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

53 extracted references · 49 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sentence := #2 '...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    , Bertail, P

    Aghbalou, A. , Bertail, P. , Portier, F. & Sabourin, A. (2024). Cross-validation on extreme regions. Extremes , 1--51

  4. [4]

    , Goegebeur, Y

    Beirlant, J. , Goegebeur, Y. , Segers, J. & Teugels, J. L. (2006). Statistics of extremes: theory and applications. John Wiley & Sons

  5. [5]

    , Vynckier, P

    Beirlant, J. , Vynckier, P. & Teugels, J. L. (1996). Tail index estimation, pareto quantile plots regression diagnostics. Journal of the American Statistical Association 91, 1659--1667

  6. [6]

    , Goldie, C

    Bingham, N. , Goldie, C. & Teugels, J. (1987). Regular Variation. Cambridge University Press

  7. [7]

    , Lugosi, G

    Boucheron, S. , Lugosi, G. & Massart, P. (2013). Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press

  8. [8]

    & Thomas, M

    Boucheron, S. & Thomas, M. (2015). Tail index estimation, concentration and adaptivity. Electronic Journal of Statistics 9, 2751--2792

Show all 53 references
  1. [9]

    & Kim, A

    Carpentier, A. & Kim, A. K. (2015). Adaptive and minimax optimal estimation of the tail coefficient. Statistica Sinica , 1133--1144

  2. [10]

    , Lederer, J

    Chichignoud, M. , Lederer, J. & Wainwright, M. J. (2016). A practical scheme and fast algorithm to tune the lasso with optimality guarantees. Journal of Machine Learning Research 17, 1--20

  3. [11]

    , Jalalzai, H

    Cl \'e men c on, S. , Jalalzai, H. , Lhaut, S. , Sabourin, A. & Segers, J. (2023). Concentration bounds for the empirical angular measure with statistical learning applications. Bernoulli 29, 2797--2827

  4. [12]

    & Sabourin, A

    Cl \'e men c on, S. & Sabourin, A. (2025). Weak signals and heavy tails: Machine-learning meets extreme value theory. arXiv:2504.06984

  5. [13]

    & Lacour, C

    Comte, F. & Lacour, C. (2013). Anisotropic adaptive kernel deconvolution. Annales de l'IHP Probabilit \'e s et statistiques 49, 569--609

  6. [14]

    , Deheuvels, P

    Csorgo, S. , Deheuvels, P. & Mason, D. (1985). Kernel estimates of the tail index of a distribution. The Annals of Statistics , 1050--1077

  7. [15]

    & Resnick, S

    De Haan, L. & Resnick, S. (1987). On regular variation of probability densities. Stochastic processes and their applications 25, 83--93

  8. [16]

    de Haan, L. F. M. (1970). On regular variation and its application to the weak convergence of sample extremes, vol. 32. Mathematisch Centrum

  9. [17]

    Drees, H. (2001). Minimax risk bounds in extreme value theory. The Annals of Statistics 29, 266--294

  10. [18]

    & Kaufmann, E

    Drees, H. & Kaufmann, E. (1998). Selecting the optimal sample fraction in univariate extreme value estimation. Stochastic Processes and their applications 75, 149--172

  11. [19]

    & Sabourin, A

    Drees, H. & Sabourin, A. (2021). Principal component analysis for multivariate extremes. Electronic Journal of Statistics 15, 908--943

  12. [20]

    , Kl \"u ppelberg, C

    Embrechts, P. , Kl \"u ppelberg, C. & Mikosch, T. (2013). Modelling extremal events: for insurance and finance, vol. 33. Springer Science & Business Media

  13. [21]

    , Lalancette, M

    Engelke, S. , Lalancette, M. & Volgushev, S. (2021). Learning extremal graphical structures in high dimensions. arXiv:2111.00840

  14. [22]

    Fedotenkov, I. (2020). A review of more than one hundred pareto-tail index estimators. Statistica 80, 245--299

  15. [23]

    , Sabourin, A

    Goix, N. , Sabourin, A. & Cl \'e men c on, S. (2015). Learning the dependence structure of rare events: a non-asymptotic study. In Conference on Learning Theory

  16. [24]

    & Lepski, O

    Goldenshluger, A. & Lepski, O. (2011). Bandwidth selection in kernel density estimation: Oracle inequalities and adaptive minimax optimality. The Annals of Statistics 39, 1608--1632

  17. [25]

    , De Haan, L

    Gomes, I. , De Haan, L. & Rodrigues, L. H. (2008). Tail index estimation for heavy-tailed models: accommodation of bias in weighted log-excesses. Journal of the Royal Statistical Society Series B: Statistical Methodology 70, 31--52

  18. [26]

    Gomes, M. I. , Caeiro, F. , Henriques-Rodrigues, L. & Manjunath, B. (2016). Bootstrap methods in statistics of extremes. Handbook of Extreme Value Theory and Its Applications to Finance and Insurance. Handbook Series in Financial Engineering and Econometrics (Ruey Tsay Adv. Ed...

  19. [27]

    Gomes, M. I. , Figueiredo, F. & Neves, M. M. (2012). Adaptive estimation of heavy right tails: resampling-based methods in action. Extremes 15, 463--489

  20. [28]

    Gomes, M. I. & Oliveira, O. (2001). The bootstrap methodology in statistics of extremes—choice of the optimal sample fraction. Extremes 4, 331--358

  21. [29]

    & Spokoiny, V

    Grama, I. & Spokoiny, V. (2008). Statistics of extremes by oracle estimation. The Annals of Statistics 36, 1619--1648

  22. [30]

    Hall, P. (1982). On some simple estimates of an exponent of regular variation. Journal of the Royal Statistical Society: Series B (Methodological) 44, 37--42

  23. [31]

    & Welsh, A

    Hall, P. & Welsh, A. H. (1984). Best attainable rates of convergence for estimates of parameters of regular variation. The Annals of Statistics , 1079--1084

  24. [32]

    & Welsh, A

    Hall, P. & Welsh, A. H. (1985). Adaptive estimates of parameters of regular variation. The Annals of Statistics , 331--341

  25. [33]

    Hill, B. M. (1975). A simple general approach to inference about the tail of a distribution. The Annals of Statistics 3, 1163--1174

  26. [34]

    Ibragimov, I. (1975). Independent and stationary sequences of random variables. Wolters, Noordhoff Pub

  27. [35]

    & Massart, P

    Lacour, C. & Massart, P. (2016). Minimal penalty for goldenshluger--lepski method. Stochastic Processes and their Applications 126, 3774--3789

  28. [36]

    , Fischer, A

    Laszkiewicz, M. , Fischer, A. & Lederer, J. (2021). Thresholded adaptive validation: Tuning the graphical lasso for graph recovery. In International Conference on Artificial Intelligence and Statistics. PMLR

  29. [37]

    Lederer, J. (2022). Fundamentals of High-Dimensional Statistics: with exercises and R labs. Springer Texts in Statistics

  30. [38]

    Lepski, O. V. (1990). A problem of adaptive estimation in gaussian white noise. Teoriya Veroyatnostei i ee Primeneniya 35-3, 459--470

  31. [39]

    Lepski, O. V. , Mammen, E. & Spokoiny, V. G. (1997). Optimal spatial adaptation to inhomogeneous smoothness: an approach based on kernel estimates with variable bandwidth selectors. The Annals of Statistics , 929--947

  32. [40]

    , Sabourin, A

    Lhaut, S. , Sabourin, A. & Segers, J. (2021). Uniform concentration bounds for frequencies of rare events. arXiv:2110.05826

  33. [41]

    & Lederer, J

    Li, W. & Lederer, J. (2019). Tuning parameter calibration for 1-regularized logistic regression. Journal of Statistical Planning and Inference 202, 80--98

  34. [42]

    Mason, D. M. (1982). Laws of large numbers for sums of extreme values. The Annals of Probability , 754--764

  35. [43]

    Nolan, J. P. (2020). Univariate stable distributions. Springer Series in Operations Research and Financial Engineering 10, 978--3

  36. [44]

    Pickands III, J. (1975). Statistical inference using extreme order statistics. The Annals of Statistics , 119--131

  37. [45]

    Reiss, R.-D. (2012). Approximate distributions of order statistics: with applications to nonparametric statistics. Springer science & business media

  38. [46]

    (2007 a )

    Resnick, S. (2007 a ). Heavy-tail phenomena: probabilistic and statistical modeling. Springer Science & Business Media

  39. [47]

    (2007 b )

    Resnick, S. (2007 b ). Heavy-tail phenomena: probabilistic and statistical modeling, vol. 10. Springer Science & Business Media

  40. [48]

    Resnick, S. I. (2008). Extreme values, regular variation, and point processes, vol. 4. Springer Science & Business Media

  41. [49]

    & MacDonald, A

    Scarrott, C. & MacDonald, A. (2012). A review of extreme value threshold estimation and uncertainty quantification. REVSTAT-Statistical journal 10, 33--60

  42. [50]

    Segers, J. (2002). Abelian and tauberian theorems on the bias of the hill estimator. Scandinavian Journal of Statistics 29, 461--483

  43. [51]

    , Lim, N

    Taheri, M. , Lim, N. & Lederer, J. (2023). Balancing statistical and computational precision: A general theory and applications to sparse regression. IEEE Transactions on Information Theory 69, 316--333

  44. [52]

    Weissman, I. (1978). Estimation of parameters and large quantiles based on the k largest observations. Journal of the American Statistical Association 73, 812--815

  45. [53]

    Zhang, A. R. & Zhou, Y. (2020). On the non-asymptotic and sharp lower tail bounds of random variables. Stat 9, e314

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.