REVIEW 3 major objections 5 minor 4 references
High-Dimensional Extreme Quantile Regression
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Extreme conditional quantiles can be estimated consistently when the number of predictors grows with the sample size, via penalized intermediate quantile regression, a refined Hill tail-index step, and extreme-value extrapolation.
desk verdict A genuinely new high-dimensional extreme quantile method with a load-bearing common-tail-index assumption; the practical tuning rule does not fall under the stated theory and the supplement is missing, but the core is worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the tail linear quantile model $Q_Y(\tau|X)=Z^\top\beta(\tau)$ for $\tau\in[\tau_{ln},1)$ with only $s$ active slopes. On it the paper stacks three mechanisms: an $\ell_1$-penalized quantile regression at intermediate levels, with penalty scale $\lambda\asymp \sqrt{n\log p}/\sqrt{1-\tau_{0n}}$; a refined Hill estimator $\hat\gamma=\big(\sum_{j=1}^J \phi(\ell_j)\log(1/\ell_j)\big)^{-1}\sum_{j=1}^J \phi(\ell_j)\log\big(z^\top\hat\beta(\tau_j)/z^\top\hat\beta(\tau_1)\big)$ that combines a fixed number of intermediate quantile lines; and the extrapolation $\hat Q_Y(\tau_n|x)=\big((1-\tau_{0n})/(1-\tau_n)\big)^{\hat\gamma}z^\top\hat\beta(\tau_{0n})$. The argument runs on Condition C5, which assumes an auxiliary linear response $Z^\top\theta_r$ such that the tail of $U=Y-Z^\top\theta_r$ is uniformly close to one heavy-tailed distribution $F_0$ with a single extreme value index $\gamma$ and controlled second-order error; that assumption supplies the density-decay rate that replaces the usual central-quantile density lower bound and justifies using one EVI estimate at every covariate value.
What would settle it
Run the method on simulated data from a sparse location-scale model with a covariate-dependent tail index — for example $Y=X_1+(1+0.5X_1)\varepsilon$ with $\varepsilon\sim t(\nu_1)$ for $X_2\le 0$ and $\varepsilon\sim t(\nu_2)$ for $X_2>0$, $\nu_1\neq\nu_2$, with $p$ growing like $n$ — and check whether the refined Hill estimator stabilizes around a single γ and whether the relative error at $\tau_n=0.999$ converges at the claimed rate. This violates Condition C5's common-EVI requirement, so the relative error should fail to vanish.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that a sparse tail linear quantile model $Q_Y(\tau|X)=Z^\top\beta(\tau)$ for $\tau$ close to 1 makes arbitrarily extreme conditional quantiles estimable in high dimensions. The estimator first runs $\ell_1$-penalized quantile regression at intermediate levels $\tau_j=1-\ell_j k/n$, then estimates the extreme value index $\gamma$ with a weighted refined Hill estimator applied to the fitted quantile lines, and finally extrapolates through the tail-quantile relation $U_Y(tz|x)/U_Y(t|x)\to z^\gamma$. Theorem 3 states that the relative error of the extrapolated estimator is $O_p\big(\log[k/\{n(1-\tau_n)\}]\,(ns/k)\sqrt{\log(p\vee n)/n}\big)$, which vanishes under the paper's conditions; the rate also requires the intermediate level $k$ to exceed $\sqrt n$ by a factor and stay below $n$. Supporting this, the paper proves a uniform bound for the penalized intermediate-quantile estimator whose main inflation term, $d_n^\gamma(1-\tau_{0n})^{-\gamma-1}$, measures how fast the conditional density decays at intermediate quantiles.
Load-bearing premise
The load-bearing premise is that, after subtracting an auxiliary linear component, the tail of the response at every covariate value is uniformly close to one heavy-tailed distribution with a single extreme value index γ; if the tail index in fact varies across covariate values, or the uniform approximation error is too large, the extrapolation claim in Theorem 3 does not follow.
Editorial extensions
If this is right
- If the central claim is right, extreme conditional quantiles become estimable when $p$ is comparable to or larger than $n$, where the fixed-dimension EQR method is inapplicable and direct high-dimensional quantile regression has large bias.
- The relative-error rate makes the bias–variance tradeoff explicit: consistency requires a window with $k$ above $\sqrt n$, $k=o(n)$, and $\log[k/\{n(1-\tau_n)\}](ns/k)\sqrt{\log(p\vee n)/n}\to 0$, so the intermediate level cannot be chosen as in the fixed-dimension regime.
- Because the $\ell_1$ penalty produces sparse coefficient paths at intermediate quantiles, the same framework performs tail-relevant variable selection, not just quantile estimation.
- The refined Hill estimator has a rate that does not depend on $\gamma$, so the method remains stable across heavier-tailed responses where direct quantile estimators show much larger errors.
Reading between the lines
- Editorial extension: the common-γ assumption is the main boundary; a grouped or local version of the refined Hill step would extend the method to data where the tail index varies with covariates, at the cost of pooling fewer observations per group.
- Editorial extension: the paper's rule-of-thumb choice $k=\lfloor c_0 n^{0.5+\delta_1}(\log p)^{0.5+\delta_2}\rfloor$ is heuristic; a data-driven selector that minimizes a plug-in estimate of the relative-error bound in Theorem 3 is a natural next test.
- Editorial extension: the same three-step logic should transfer to other tail functionals such as extreme expectiles or conditional tail expectations by replacing the quantile loss and the extrapolation relation, though the density-decay argument would need re-derivation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-step estimator for extreme conditional quantiles in high-dimensional linear quantile regression with heavy tails: l1-penalized quantile regression at intermediate levels, a refined Hill estimator based on a fixed number of intermediate quantile fits, and Weissman-type extrapolation from an intermediate quantile to the extreme level. Theorems 1–3 state uniform error bounds for the intermediate quantile estimator, a convergence rate for the refined Hill estimator, and a relative-error rate for the extreme quantile estimator, all under conditions C1–C6 that are mostly deferred to a supplementary file. The paper also reports simulations and an auto-insurance application, and it discusses practical tuning for k, λ, and the refined Hill weights.
Significance. If the results are correct, the paper makes a useful contribution: it is among the first to derive uniform rates for penalized quantile regression at intermediate levels and to combine them with a refined Hill estimator in a high-dimensional setting, and it explicitly quantifies how the decay of the conditional density in the tail inflates the estimation rate. The proposed method is simple to implement, and the simulation study indicates that it can outperform both the fixed-dimensional extremal method of Wang et al. (2012) and the high-dimensional central-quantile method of Belloni and Chernozhukov (2011). However, the practical significance is tempered by two issues: the main results rest on a strong homogeneous-tail-index assumption (Condition C5), and the simulation settings violate the theoretical lower bound on k, so the reported numerical success is not covered by the stated theory.
major comments (3)
- [Section 2.1, Theorems 1–3; Section 2.4] The conditions C1–C4 and C6 and all proofs of Theorems 1–3 are in a supplementary file that is not included with the manuscript. The main text states Theorem 1 and Theorem 2 as consequences of these omitted conditions plus Lemma 1 and Remark S.2, which are also only referenced. Without the supplement, the central consistency claims cannot be verified; for example, the rate in Theorem 1 depends on unstated conditions on the design, the sparsity of θ_r, and the tail density, and the proof of Theorem 2 relies on a second-order expansion that is not shown. The supplement must be made available to reviewers before the technical content can be assessed.
- [Section 2.1, Condition C5; Section 2.3, Remark 3] The assumption of a single extreme value index γ for all covariate values is load-bearing: the refined Hill estimator is computed at one point x̄ and then used in the extrapolation factor for every x. If the conditional tail index varies with X, the relative error of \Q_Y(τ_n|x) contains a term ((1-τ0n)/(1-τn))^{γ(x)-γ(x̄)}-1 that diverges as τ_n→1 and dominates the rate in Theorem 3. The paper's title and abstract suggest a general method, but the theoretical result applies only to a homothetic-tail family. The simulations in Section 3 use a location-scale model with constant γ by construction, so they provide no evidence on robustness to heterogeneous tail indices. The paper should either narrow its claims, add a diagnostic for the common-tail assumption, or provide a sensitivity analysis.
- [Section 2.4, 'Selection of k'; Example 1(II)] The rule of thumb k = ⌊c0 n^{0.5+δ1} (log p)^{0.5+δ2}⌋ with (c0,δ1,δ2)=(0.8,0.01,0.05) yields k ≈ 54 for n=1000, p≈32 and k ≈ 137 for n=5000, p≈71, while the lower bound stated in Example 1(II) is √(log p)√n ≈ 59 and ≈ 146, respectively. Thus the simulation settings do not satisfy the assumptions under which Theorems 2 and 3 are proven. Although asymptotically the n^{δ1} factor eventually makes the rule exceed the lower bound, the sample sizes used in the simulations are not large enough for this to occur. Since the paper's simulation claims are a central part of the demonstration, the tuning rule should be revised (e.g., a larger c0) or the discrepancy should be discussed explicitly.
minor comments (5)
- [Section 2.4, 'Selection of λ'] The theorems require λ ≍ √n log p / √(1-τ0n), but all numerical results use 10-fold cross-validation. The cross-validated λ is not shown to satisfy the theoretical scaling, so the finite-sample results are not covered by the proof. This is a secondary gap, but it would help to comment on how CV-based λ relates to the theoretical choice.
- [Introduction] The symbol τ_n is used in the introduction for both the intermediate and the extreme quantile level, which conflicts with the later distinction between τ0n and τn. Please clarify the notation in the first occurrence.
- [Tables 1 and 2] The caption for EQR reads 'the method in W ang, Li and He (2012)' with a typo in 'Wang'.
- [Table 3] Some standard errors for HQR are large relative to the MISE values (e.g., 0.33 when MISE is 0.44), so the claim of 'significantly less efficient' is not supported by a formal comparison. Consider reporting confidence intervals or a paired comparison.
- [Section 2.2] The sentence 'Suppose there exist a positive sequence d1n such that, K1(Z) ≍ d1n hold for all Z ∈ Z.' repeats part of Condition C5; please rephrase or remove.
Circularity Check
No significant circularity: the extreme quantile target is outside the fitted inputs, and the extrapolation is a standard EVT plug-in built on explicit assumptions.
full rationale
Walking the derivation chain, the proposed estimator bQY(tau_n|x) = ((1 - tau_0n)/(1 - tau_n))^{gamma_hat} z^T beta_hat(tau_0n) is a plug-in combination of an intermediate-level quantile estimate and a refined Hill estimate of gamma; the target QY(tau_n|x) is never used to construct beta_hat or gamma_hat. The levels used for fitting satisfy n(1 - tau_0n) = k -> infinity, while the claimed extreme level satisfies n(1 - tau_n) = o(k), so the prediction is made at a level outside the fitted sample. Condition C5 is an explicit domain-of-attraction and uniform-tail assumption that ensures a single extreme value index gamma; if gamma varies with the covariates, Theorem 3 need not hold, but that is a scope/robustness limitation rather than a circular definition. The paper's citations to the authors' earlier extreme-quantile work supply standard EVT constructions and weight choices, but the main consistency result rests on the l1-penalized intermediate quantile bound of Theorem 1 and the second-order tail condition in C5, not on a citation to a result equivalent to Theorem 3. No fitted parameter is renamed as a prediction, and no uniqueness claim is imported from prior work by the same authors. The main text defers full conditions and proofs to the supplementary material, so complete verification is not possible from the supplied text, but no circular reduction is evident from the material provided.
Assumptions & free parameters
free parameters (3)
- k (number of upper order statistics, τ0n = 1 - k/n) =
rule of thumb: k = ⌊0.8 n^{0.51} (log p)^{0.55}⌋ in numerical studies
- λ (L1 penalty parameter) =
selected by 10-fold cross-validation
- J, s, a (refined Hill weight parameters) =
recommended by He et al. (2022): ϕ(ℓj)=ℓj^a, ℓj=s^{j-1}, with a,s∈(0,1), J fixed
assumptions (5)
- domain assumption FY(·|X) ∈ D(Gγ) with γ > 0 for all X in the support
- domain assumption Tail linear quantile model QY(τ|X) = Z^T β(τ) for all τ ∈ [τ_ln, 1)
- domain assumption Sparsity: s = o(n) tail-relevant covariates
- domain assumption Condition C5: existence of auxiliary line and uniform tail approximations with rates
- domain assumption Conditions C1-C4 and C6 (design, moment, and second-order conditions)
Cite this review
Pith. "Pith review of High-Dimensional Extreme Quantile Regression." pith.science (2026). https://pith.science/paper/WREZ6YHM
@misc{pith2026241113822,
author = {Pith},
title = {Pith review of: High-Dimensional Extreme Quantile Regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/WREZ6YHM}},
note = {Machine review of arXiv:2411.13822}
}
read the original abstract
The estimation of conditional quantiles at extreme tails is of great interest in numerous applications. Various methods that integrate regression analysis with an extrapolation strategy derived from extreme value theory have been proposed to estimate extreme conditional quantiles in scenarios with a fixed number of covariates. However, these methods prove ineffective in high-dimensional settings, where the number of covariates increases with the sample size. In this article, we develop new estimation methods tailored for extreme conditional quantiles with high-dimensional covariates. We establish the asymptotic properties of the proposed estimators and demonstrate their superior performance through simulation studies, particularly in scenarios of growing dimension and high dimension where existing methods may fail. Furthermore, the analysis of auto insurance data validates the efficacy of our methods in estimating extreme conditional insurance claims and selecting important variables.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Belloni, A. and Chernozhukov, V. (2011), ‘ ℓ1-penalized quantile regression in high- dimensional sparse models’, Annals of Statistics 39(1), 82–130. Belloni, A., Chernozhukov, V., Chetverikov, D. and Fern´ andez-Val, I. (2019), ‘Condi- tional quantile processes based on series or many regressors’, Journal of Econometrics 213(1), 4–29. Bickel, P. J., Ritov...
arXiv 2011
-
[222]
(1989), ‘On m-processes and m-estimation’, Annals of Statistics 17(1), 337–361
31 Welsh, A. (1989), ‘On m-processes and m-estimation’, Annals of Statistics 17(1), 337–361. Xu, W., Hou, Y. and Li, D. (2022), ‘Prediction of extremal expectile based on regres- sion models with heteroscedastic extremes’, Journal of Business & Economic Statistics 40(2), 522–536. Youngman, B. D. (2019), ‘Generalized additive models for exceedances of high...
arXiv 1989
-
[839]
Chetverikov, D., Liao, Z. and Chernozhukov, V. (2021), ‘On cross-validated lasso in high dimensions’, Annals of Statistics 49(3), 1300–1317. Clemente, C., Guerreiro, G. R. and Bravo, J. M. (2023), ‘Modelling motor insurance claim frequency and severity using gradient boosting’, Risks 11(9), 1–20. Daouia, A., Gardes, L. and Girard, S. (2013), ‘On kernel sm...
arXiv 2021
-
[1464]
Wang, H. and Tsai, C.-L. (2009), ‘Tail index regression’,Journal of the American Statistical Association 104(487), 1233–1240. Wang, L., Wu, Y. and Li, R. (2012), ‘Quantile regression for analyzing heterogeneity in ultra-high dimension’, Journal of the American Statistical Association 107(497), 214–
work page 2009
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.