REVIEW 4 major objections 4 minor 50 references
On rank estimators in increasing dimensions
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper establishes the first high-dimensional theory for rank estimators such as Han’s maximum rank correlation estimator, showing the estimator stays accurate at the minimax rate but its normal approximation requires a much stronger…
desk verdict The qualitative message is solid, but the advertised Bahadur-type bounds are off by a square root — the proof only supports the square root of the stated rate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the U-process objective $\Gamma_n(\theta) = \frac{1}{n(n-1)}\sum_{i\neq j} f(Z_i,Z_j;\theta)$ and its Hoeffding decomposition $\Gamma_n(\theta) = \Gamma(\theta) + P_n g(\cdot;\theta) + U_n h(\cdot,\cdot;\theta)$. The key technical contribution is a new maximal inequality for degenerate U-processes in increasing dimensions, which controls the uniform decay of the remainder $\sup_{\theta\in B(\theta_0,r_n)}|U_n h(\cdot,\cdot;\theta)|$ in terms of the VC dimension $\nu_n$, the radius $r_n$, and a variance proxy $\tilde\epsilon_n$. This inequality makes possible a Bahadur-type representation for $\hat\theta_n$ by transferring the analysis to the smoothed objective $\tilde\Gamma_n(\theta) = \Gamma(\theta) + P_n \tau(\cdot;\theta)$, whose theoretical properties are handled via Assumption 3 and exponential-moment bounds. The machinery is what turns the discontinuous loss into a tractable smooth one while keeping track of how $p_n$ affects all rates.
What would settle it
Simulate Han's MRC estimator with heavy-tailed covariates (for example, $t$-distributed with few degrees of freedom) while keeping $p_n$ and $n$ within the paper's scaling regime, and examine whether the coverage probability of the normal confidence interval for a fixed projection deteriorates substantially faster than in the Gaussian-design simulations, or whether the Bahadur expansion's error term grows at a rate larger than the paper's bound.
Extended reading notes
Core claim
The paper's central claim is that for M-estimators whose objective functions are U-processes and possibly discontinuous, in the increasing-dimension regime, estimation still achieves the minimax-optimal $(p_n/n)^{1/2}$ rate, but normal approximation demands far more stringent scaling. Specifically, for Han's maximum rank correlation estimator under Assumptions 4–6, the paper establishes $\|\hat\theta_n^H - \theta_0 + (V^H)^{-1}P_n\nabla_1\tau^H(\cdot;\theta_0)\|_2 = O_P\bigl(\log(n/p_n^2)p_n^{3/2}/n^{5/4}\bigr)$ whenever $p_n^2/n = o(1)$ and $\log(n/p_n^2)p_n^{3/2}/n^{5/4} = o(1)$. It further shows that $\sqrt{n}\gamma^T(\hat\theta_n^H - \theta_0) / (\gamma^T (V^H)^{-1}\Delta^H (V^H)^{-1} \gamma)^{1/2} \Rightarrow N(0,1)$ under the stronger condition $\log(n/p_n^2)p_n^{3/2}/n^{1/4} = o(1)$. The same pattern holds for the other three rank estimators: a minimax estimation rate, followed by a much more demanding condition for valid normal inference. The paper also proves that the numerical-derivative covariance estimator is consistent only if the step size is tuned with respect to $p_n$, not just $n$.
Load-bearing premise
The results for Han's estimator require an exponential moment bound (Assumption 3(v) under Conditions 1–3), which is verified only for subgaussian designs with smooth conditional densities and bounded second derivatives of the link function; if the covariates are heavy-tailed or the conditional density is not smooth, the Bahadur bound and normal approximation may fail.
Editorial extensions
If this is right
- For semiparametric index models with many covariates, rank estimators remain usable for point estimation under only $p_n/n \to 0$, matching the minimax-optimal $(p_n/n)^{1/2}$ rate.
- The usual scaling condition $p_n^2/n \to 0$ used for smooth M-estimators is generally insufficient for normal approximation of rank estimators; inference requires the stronger $\log(n/p_n^2)p_n^{3/2}/n^{1/4} = o(1)$ in the Han example.
- To obtain consistent covariance matrices by numerical differentiation, the step size must shrink with $p_n$: the paper shows consistency under $\varepsilon_n\sqrt{p_n} = o(1)$ and $\varepsilon_n^{-2}p_n/\sqrt{n} = o(1)$, so the recommended $\varepsilon_n \asymp (p_n/n)^{1/6}$ depends on the dimension.
- Confidence intervals based on normal approximation deteriorate quickly as $p_n$ increases for fixed $n$, as confirmed by the paper's simulations where even $p_n=3$ or $4$ shows severe coverage distortion at moderate sample sizes.
Reading between the lines
- A likely consequence is that practitioners who want valid confidence intervals with many regressors should shift to alternative inferential methods for rank estimators, such as resampling or bootstrap calibrations, though the paper does not analyze those procedures here.
- The maximal inequality for degenerate U-processes is a general tool that could be applied to other non-smooth econometric estimators beyond rank correlations, such as maximum score or other pairwise-comparison estimators in increasing dimensions.
- A natural testable extension is to check whether the conditions imply that the bootstrap, or subsampling, can restore valid coverage under weaker scaling than the normal approximation; the paper leaves this unexplored.
- Because the Assumption 3(v) exponential-moment bound is verified for Han's estimator only under subgaussian designs and smooth conditional densities, one might expect the normal approximation to break down even earlier for heavy-tailed designs; this is not tested in the paper.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops asymptotic theory for M-estimators whose objective functions are U-processes, possibly discontinuous, in an increasing-dimension setting where both the data dimension m_n and parameter dimension p_n grow with n. The main results are: a maximal inequality for degenerate U-processes in increasing dimensions (Theorem 2.1); consistency under ν_n/n → 0 (Theorem 2.2); a (ν_n∨p_n)/n rate of convergence (Theorem 2.3); a Bahadur-type bound and normal approximation under stronger scaling (Theorem 2.4); and consistency of numerical-derivative covariance estimators with step-size calibration (Theorem 2.6). The general results are applied to four rank estimators: Han's maximum rank correlation, Cavanagh–Sherman, Khan–Tamer, and Abrevaya–Shin, with corollaries giving p_n/n^{1/2} estimation rates, Bahadur-type bounds of order log(n/p_n^2)p_n^{3/2}/n^{5/4}, and normal approximation under log(n/p_n^2)p_n^{3/2}/n^{1/4}=o(1). Simulations for Han's MRC illustrate coverage deterioration with p_n.
Significance. The paper addresses a genuine gap: existing increasing-dimension M-estimation theory excludes discontinuous U-process objectives, and rank estimators have only been analyzed for fixed p. The new maximal inequality for degenerate U-processes (Theorem 2.1) and the careful tracking of ν_n, p_n, and m_n are valuable technical contributions, and the conclusion that normal inference requires a much stronger scaling condition than estimation is economically important. The extensive proofs for the general M-estimator and for Han's MRC are detailed and represent serious work. However, the advertised Bahadur-type rates are not supported by the proof as written (missing square root; see major comments), and the minimax optimality claim is not backed by a lower bound. With those corrected, the framework would be a solid contribution to the nonparametric and semiparametric econometrics literature.
major comments (4)
- [§A.3.5, Theorem 2.4(i)] The displayed Bahadur bound is not implied by the proof. After (A.28) one has 0 ≤ −(1/2)(t̂_n − t*_n)^T V(t̂_n − t*_n) ≤ 2nφε, and with Assumption 3(ii) this yields ‖t̂_n − t*_n‖ = O_P(√(nφε)), so ‖θ̂_n − θ0 + V^{-1}P_n∇_1τ(·;θ0)‖_2 = O_P(√φε), not O_P(φε). The rate displayed in Theorem 2.4(i) and the rates in Corollaries 3.1(iii), 3.2(iii), 3.3(iii), and 3.4(iii) must be replaced by square-root versions; for Corollary 3.1(iii), for example, the proof supports O_P((log(n/p_n^2))^{1/2}p_n^{3/4}/n^{5/8} + p_n^{5/4}/n^{3/4}) under the stated scaling, not the displayed log(n/p_n^2)p_n^{3/2}/n^{5/4}. The asymptotic-normality scaling condition log(n/p_n^2)p_n^{3/2}/n^{1/4}=o(1) is unchanged because it is equivalent to √φε = o(1/√n), but the quantitative claims in the abstract and corollaries need revision.
- [Abstract and §1.1] The abstract and Section 1.1 call (p_n/n)^{1/2} the 'minimax optimal' rate, but no minimax lower bound is proved for the increasing-dimension problems studied. The reference (Yu, 1997) supplies lower bounds only in fixed-dimensional settings; the triangular-array framework with changing parameter spaces requires a new lower-bound argument. As it stands, Theorems 2.3 and Corollaries 3.1–3.4 only establish upper bounds, so the optimality claim should either be proved or softened.
- [Appendix A.4.1, Lemmas A.7–A.9] Lemmas A.7–A.9, which bound sup_{θ∈B(θ0,r)} E{h^C(·,·;θ)}^2, E{h^K(·,·;θ)}^2, and E{h^A(·,·;θ)}^2, are omitted with the statement that their proofs are similar to Lemma A.6. These lemmas are load-bearing for Corollaries 3.2–3.4, and the similarity is not immediate: the censored-duration objective includes R_i and V_i, and the Abrevaya–Shin objective includes a kernel K_b(W_i−W_j) with bandwidth, which changes the differentiation and moment arguments. The proofs should be supplied or at least the differences from Lemma A.6 detailed.
- [§3.2–§3.4, Assumption 3(v)] Assumption 3(v), the exponential moment condition on the smoothed Hessian, is verified only for Han's MRC in Theorem 3.1 under Conditions 1–3. For the Cavanagh–Sherman, Khan–Tamer, and Abrevaya–Shin estimators, no primitive conditions are given under which Assumption 3(v) (or Assumption 3(iii)) holds; the corresponding corollaries therefore rely on an unverified high-level condition. The authors should state sufficient design and smoothness conditions for these estimators or explicitly flag Assumption 3 as a high-level condition that must be checked case by case.
minor comments (4)
- [§A.3.5] The quantity φε is used in (A.26) before it is defined in (A.27); reorder the display or define φε earlier.
- [§1.4] The notation 'P− →' in Section 1.4 appears garbled; standard notations for convergence in probability should be used.
- [Corollary 3.4(iii)] In Corollary 3.4(iii), the term n^{-δJ} should be made explicit as O_P(n^{-δJ}), and the dependence of the constant on J and on the kernel K(·) should be stated.
- [Section 4] Tables 1–3 report coverage probabilities for three projection directions, and Figures 1–3 are not explicitly cross-referenced to those directions in the text; a sentence stating which figure corresponds to which projection would improve readability.
Circularity Check
No circularity: the derivation is self-contained and no prediction reduces to a fitted input or self-citation chain.
full rationale
The paper's central results are derived under explicit assumptions (Assumptions 1–6, Conditions 1–3) rather than fitted to the data, and the main theorems are proved from external empirical-process tools (Nolan and Pollard 1987, Kosorok 2007, Spokoiny 2012a/b, 2013, Sherman 1993). The only self-citation is Han et al. (2017), mentioned in the concluding remarks as a direction for future smoothing work ('if no further smoothing (cf. Han et al. (2017)) is made'); it is not used to establish consistency, rates, Bahadur bounds, or normality, so it is not load-bearing. The claimed sqrt(p/n) estimation rate follows from Theorem 2.3 after bounding the VC dimension of the rank-estimator function classes, and the Bahadur-type bound in Corollary 3.1(iii) is obtained by plugging the moment bound from Lemma A.6 into Theorem 2.4, with no parameter renamed as a prediction. The asymptotic covariance estimator consistency is also proved from the assumed scaling of the numerical-derivative step size rather than chosen to match the conclusions. The reader-submitted concern about a missing square root in Theorem 2.4's displayed rate is a potential correctness or proof-error issue, not circularity, because the theorem's conclusion does not reproduce an input by construction. Accordingly, no specific circular step can be exhibited, and the score is 0.
Assumptions & free parameters
free parameters (2)
- step size epsilon_n in numerical derivative covariance estimation =
epsilon_n ~ (p_n/n)^{1/6} (recommended)
- bandwidth b in Abrevaya-Shin estimator =
b = c n^{-delta} with 1/J < delta < 1/5
assumptions (7)
- standard math Hoeffding decomposition of the U-process objective into a smooth empirical process plus a degenerate U-process
- standard math Nolan and Pollard (1987) maximal inequality for degenerate U-processes
- standard math Covering number bounds for VC-subgraph classes (Theorem 9.3 in Kosorok 2007)
- standard math Spokoiny's finite-sample theory for differentiable M-estimators with increasing dimension (Spokoiny 2012a, 2013)
- domain assumption Assumption 1: identifiability with a uniform gap xi0 between the maximum and the boundary
- domain assumption Assumption 3: local strong convexity, smoothness of E tau, and subgaussian exponential moment condition on the Hessian of zeta
- domain assumption The function class F is a uniformly bounded VC-subgraph class with VC dimension nu_n
Cite this review
Pith. "Pith review of On rank estimators in increasing dimensions." pith.science (2026). https://pith.science/paper/LVOWM5JZ
@misc{pith2026190805255,
author = {Pith},
title = {Pith review of: On rank estimators in increasing dimensions},
year = {2026},
howpublished = {\url{https://pith.science/paper/LVOWM5JZ}},
note = {Machine review of arXiv:1908.05255}
}
abstract
The family of rank estimators, including Han's maximum rank correlation (Han, 1987) as a notable example, has been widely exploited in studying regression problems. For these estimators, although the linear index is introduced for alleviating the impact of dimensionality, the effect of large dimension on inference is rarely studied. This paper fills this gap via studying the statistical properties of a larger family of M-estimators, whose objective functions are formulated as U-processes and may be discontinuous in increasing dimension set-up where the number of parameters, $p_{n}$, in the model is allowed to increase with the sample size, $n$. First, we find that often in estimation, as $p_{n}/n\rightarrow 0$, $(p_{n}/n)^{1/2}$ rate of convergence is obtainable. Second, we establish Bahadur-type bounds and study the validity of normal approximation, which we find often requires a much stronger scaling requirement than $p_{n}^{2}/n\rightarrow 0.$ Third, we state conditions under which the numerical derivative estimator of asymptotic covariance matrix is consistent, and show that the step size in implementing the covariance estimator has to be adjusted with respect to $p_{n}$. All theoretical results are further backed up by simulation studies.
Figures
Reference graph
Works this paper leans on
-
[1]
Abrevaya, J. and Shin, Y. (2011). Rank estimation of partially linear index models. The Econometrics Journal , 14(3):409--437
work page 2011
-
[2]
Bahadur, R. R. (1966). A note on quantiles in large samples. The Annals of Mathematical Statistics , 37(3):577--580
work page 1966
-
[3]
Belloni, A., Chernozhukov, V., Chetverikov, D., and Wei, Y. (2018). Uniformly valid post-regularization confidence regions for many functional parameters in Z -estimation framework. The Annals of Statistics , 46(6B):3643--3675
work page 2018
-
[4]
Belloni, A., Chernozhukov, V., and Kato, K. (2014). Uniform post-selection inference for least absolute deviation regression and other Z -estimation problems. Biometrika , 102(1):77--94
work page 2014
-
[5]
Caner, M. (2014). Near exogeneity and weak identification in generalized empirical likelihood estimators: Many moment asymptotics. Journal of Econometrics , 182(2):247--268
work page 2014
-
[6]
Cattaneo, M. D., Jansson, M., and Newey, W. K. (2018a). Alternative asymptotics and the partially linear model with many regressors. Econometric Theory , 34:277--301
work page 2018
-
[7]
Cattaneo, M. D., Jansson, M., and Newey, W. K. (2018b). Inference in linear regression models with many covariates and heteroskedasticity. Journal of the American Statistical Association , 113(523):1350--1361
work page 2018
-
[8]
Cavanagh, C. and Sherman, R. P. (1998). Rank estimators for monotonic index models. Journal of Econometrics , 84(2):351--381
work page 1998
Show all 50 references
-
[9]
Chernozhukov, V., Chetverikov, D., and Kato, K. (2017). Central limit theorems and bootstrap in high dimensions. The Annals of Probability , 45(4):2309--2352
2017
-
[10]
Chernozhukov, V., Hansen, C., and Spindler, M. (2015). Valid post-selection and post-regularization inference: An elementary, general approach. Annual Review of Economics , 7:649--688
2015
-
[11]
and Gin \'e , E
de la Pena, V. and Gin \'e , E. (2012). Decoupling: From Dependence to Independence . New York: Springer
2012
-
[12]
Dudley, R. M. (1999). Uniform Central Limit Theorems . Cambridge University Press
1999
-
[13]
Fan, J., Liao, Y., and Yao, J. (2015). Power enhancement in high-dimensional cross-sectional tests. Econometrica , 83(4):1497--1541
2015
-
[14]
Han, A. K. (1987). Non-parametric analysis of a generalized regression model: the maximum rank correlation estimator. Journal of Econometrics , 35(2-3):303--316
1987
-
[15]
and Phillips, P
Han, C. and Phillips, P. C. (2006). GMM with many moment conditions. Econometrica , 74(1):147--192
2006
-
[16]
Han, F., Ji, H., Ji, Z., and Wang, H. (2017). A provable smoothing approach for high dimensional generalized regression with applications in genomics. Electronic Journal of Statistics , 11(2):4347--4403
2017
-
[17]
and Shao, Q.-M
He, X. and Shao, Q.-M. (1996). A general B ahadur representation of M -estimators and its application to linear regression with nonstochastic designs. The Annals of Statistics , 24(6):2608--2630
1996
-
[18]
and Shao, Q.-M
He, X. and Shao, Q.-M. (2000). On parameters of increasing dimensions. Journal of Multivariate Analysis , 73(1):120--135
2000
-
[19]
Hoeffding, W. (1948). A class of statistics with asymptotically normal distribution. The Annals of Mathematical Statistics , 19(3):293--325
1948
-
[20]
Honor \'e , B. E. and Powell, J. (2005). Pairwise difference estimators for nonlinear models. In Andrews, D.W.K., Stock, J.H. (Eds.) Identification and Inference in Econometric Models. Essays in Honor of Thomas Rothenberg , pages 520--553. Cambridge University Press
2005
-
[21]
Huber, P. J. (1967). The behavior of maximum likelihood estimates under nonstandard conditions. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability , pages 221--233. Berkeley, CA
1967
-
[22]
Huber, P. J. (1973). Robust regression: Asymptotics, conjectures and M onte C arlo. The Annals of Statistics , 1(5):799--821
1973
-
[23]
and Montanari, A
Javanmard, A. and Montanari, A. (2018). De-biasing the lasso: Optimal sample size for G aussian designs. The Annals of Statistics , 46(6A):2593--2622
2018
-
[24]
K., and Picek, J
Jure c kov \'a , J., Sen, P. K., and Picek, J. (2012). Methodology in Robust and Nonparametric Statistics . CRC Press
2012
-
[25]
and Tamer, E
Khan, S. and Tamer, E. (2007). Partial rank estimation of duration models with general forms of censoring. Journal of Econometrics , 136(1):251--280
2007
-
[26]
Kiefer, J. (1967). On B ahadur's representation of sample quantiles. The Annals of Mathematical Statistics , 38(5):1323--1342
1967
-
[27]
Kosorok, M. R. (2007). Introduction to Empirical Processes and Semiparametric Inference . Springer
2007
-
[28]
D., Sun, D
Lee, J. D., Sun, D. L., Sun, Y., and Taylor, J. E. (2016). Exact post-selection inference, with application to the lasso. The Annals of Statistics , 44(3):907--927
2016
-
[29]
J., and Karoui, N
Lei, L., Bickel, P. J., and Karoui, N. E. (2018). Asymptotics for high dimensional regression m-estimates: Fixed design results. Probability Theory and Related Fields , 172(3-4):983---1079
2018
-
[30]
Mammen, E. (1989). Asymptotics with increasing dimension for robust regression with applications to the bootstrap. The Annals of Statistics , 17(1):382--400
1989
-
[31]
Mammen, E. (1993). Bootstrap and wild bootstrap for high dimensional linear models. The Annals of Statistics , 21(1):255--285
1993
-
[32]
N., Ravikumar, P., Wainwright, M
Negahban, S. N., Ravikumar, P., Wainwright, M. J., and Yu, B. (2012). A unified framework for high-dimensional analysis of M -estimators with decomposable regularizers. Statistical Science , 27(4):538--557
2012
-
[33]
Newey, W. K. and Windmeijer, F. (2009). Generalized method of moments with many weak moment conditions. Econometrica , 77(3):687--719
2009
-
[34]
and Pollard, D
Nolan, D. and Pollard, D. (1987). U-processes: rates of convergence. The Annals of Statistics , 15(2):780--799
1987
-
[35]
and Pollard, D
Pakes, A. and Pollard, D. (1989). Simulation and the asymptotics of optimization estimators. Econometrica , 57(5):1027--1057
1989
-
[36]
Pollard, D. (1984). Convergence of Stochastic Processes . Springer
1984
-
[37]
Portnoy, S. (1984). Asymptotic behavior of M -estimators of p regression parameters when p^2/n is large. I. Consistency . The Annals of Statistics , 12(4):1298--1309
1984
-
[38]
Portnoy, S. (1985). Asymptotic behavior of M estimators of p regression parameters when p^2/n is large; II. Normal approximation . The Annals of Statistics , 13(4):1403--1417
1985
-
[39]
Portnoy, S. (1988). Asymptotic behavior of likelihood methods for exponential families when the number of parameters tends to infinity. The Annals of Statistics , 16(1):356--366
1988
-
[40]
Sherman, R. P. (1993). The limiting distribution of the maximum rank correlation estimator. Econometrica , 61(1):123--137
1993
-
[41]
Sherman, R. P. (1994). Maximal inequalities for degenerate U -processes with applications to optimization estimators. The Annals of Statistics , 22(1):439--459
1994
-
[42]
Spokoiny, V. (2012a). Parametric estimation. F inite sample theory. The Annals of Statistics , 40(6):2877--2909
2012
-
[43]
Spokoiny, V. (2012b). Supplement to `` P arametric estimation. F inite sample theory". The Annals of Statistics
2012
-
[44]
Spokoiny, V. (2013). Bernstein-von M ises T heorem for growing parameter dimension. arXiv preprint arXiv:1302.3430
2013 arXiv
-
[45]
Subbotin, V. Y. (2008). Essays on the Econometric Theory of Rank Regressions . PhD thesis, Northwestern University
2008
-
[46]
Van de Geer, S., B \"u hlmann, P., Ritov, Y., and Dezeure, R. (2014). On asymptotically optimal confidence regions and tests for high-dimensional models. The Annals of Statistics , 42(3):1166--1202
2014
-
[47]
and Wellner, J
van der Vaart, A. and Wellner, J. (1996). Weak Convergence and Empirical Processes . Springer
1996
-
[48]
Wang, H. (2007). A note on iterative marginal optimization: a simple algorithm for maximum rank correlation estimation. Computational Statistics and Data Analysis , 51(6):2803--2812
2007
-
[49]
Yu, B. (1997). Assouad, F ano, and L e C am. In Festschrift for Lucien Le Cam, 423--435 . Springer, New York
1997
-
[50]
and Zhang, S
Zhang, C.-H. and Zhang, S. S. (2014). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society: Series B , 76(1):217--242
2014
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.