REVIEW 1 major objections 5 minor 17 references
On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins
T0 review · 1 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read This paper proves that in random-design nonparametric regression with s-Hölder regression functions and dimension d>4s, the minimax root-mean-square risk for estimating a constant conditional variance is at least n^{-β} with β=[d(3s+1)+8s]/
desk verdict A convincing negative answer to Robins' question, with an intricate construction whose main risk is verification depth rather than a found flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The proof uses a standard two-prior testing argument with two priors whose error variances differ by Δ but whose n-observation predictive laws are nearly indistinguishable. The central object is a three-point response distribution on {−a,0,a} whose probabilities are linear in q=f(x) and q²+σ², enabling exact cancellation of all one-observation differences through three moment identities. Random design perturbations are realized through a contraction fixed point: a provisional pair-density kernel R(x,y) is mapped to densities p_T satisfying E[p_T(x)p_T(y)]=R(x,y), via a projective decomposition into rank-one kernels and independent categorical variables. The load-bearing matrix identity is a
What would settle it
Check the contraction proof at the heart of the construction: on a discretized torus with small h and N=2, compute the map T in Proposition 3.8 at a small c0; if it fails to have a fixed point satisfying R(x,y)=E[p_T(x)p_T(y)], the cancellation identities fail and the lower bound collapses. Alternatively, any explicit estimator whose worst-case root-mean-square error is o(n^{-β}) on the stated class would refute the theorem.
Extended reading notes
Core claim
The central claim is that the minimax root-mean-square risk of estimating σ² in this random-design regression model is at least c n^{-β} for every n, with β=[d(3s+1)+8s]/[(d+2s)(d+4)]. Equivalently, the conjectured rate n^{-4s/(d+4s)} is uniformly unattainable: the gap between the two exponents is d(d−4s)(s−1)/[(d+4s)(d+2s)(d+4)], which is positive for every s>1 and d>4s. The paper does not determine the exact minimax rate, only rules out the conjectured one, leaving the true rate somewhere between the fixed-grid benchmark n^{-2s/d} and the conjecture.
Load-bearing premise
The whole argument hinges on allowing the noise distribution to be chosen afresh at every covariate value; if the conditional error laws had to be identical across the design, the indistinguishability construction is not known to work, and the paper explicitly does not claim the lower bound in that case.
Editorial extensions
If this is right
- The conjectured n^{-4s/(d+4s)} rate is false for this entire model class: no estimator, however adaptive, attains it uniformly.
- The true minimax rate lies strictly between n^{-2s/d} and n^{-4s/(d+4s)}; the paper's β gives a new lower bound that improves on the fixed-grid benchmark.
- The obstruction is not an artifact of difference-based estimators; it applies to all measurable estimators of the variance.
- The lower-bound construction does not apply when conditional error laws are forced to be identical across x; under a common error law independent of X, the question remains open.
- For s≤1 or d≤4s, the argument gives no information, so the conjectured rate may still be attainable in those regimes.
Reading between the lines
- If the x-dependent error-law freedom is essential, then datasets with homogeneous noise may still admit the conjectured rate; the hardness may be driven by heteroscedasticity rather than by random design alone.
- The pair-density fixed-point device is general: any exact cancellation requiring E[p_T(x)p_T(y)]=R(x,y) can be produced by the same contraction, so the technique may transfer to other quadratic semiparametric functionals.
- The exponent β suggests a transition as d approaches 4s; a natural next step is to test numerically whether the lower bound is tight in moderate dimensions or to look for matching upper bounds.
- Because the gap between the conjectured rate and β shrinks as s approaches 1, the phenomenon may be specific to genuine Hölder smoothness with s>1 rather than to Lipschitz-scale smoothness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the minimax rate for estimating a constant conditional variance σ² in nonparametric regression under random design. The class Θ allows an unknown design density bounded above and away from zero, an s-Hölder regression function, and conditional error laws that may depend on x but have mean zero, common variance σ², and uniformly bounded fourth moments. The main result, Theorem 1.1, states that for every s>1 and integer d>4s, the minimax root-mean-square risk satisfies R_n(Θ)^{1/2} ≥ c n^{-β} with β = [d(3s+1)+8s]/[(d+2s)(d+4)]. Since 4s/(d+4s) − β > 0, the conjectured rate n^{-4s/(d+4s)} is not uniformly attainable, and the minimax rate lies strictly between the regular-grid benchmark n^{-2s/d} and the conjectured rate. The proof uses Le Cam's two-prior method, constructing priors with variance gap Δ and nearly indistinguishable n-observation predictive laws. The construction combines periodic feature systems, a localized covariance dual, random design perturbations realized through a contraction fixed point in a Wiener-norm ball, and bounded coefficient vectors satisfying Gaussian fourth-moment identities. The full-sample Hellinger bound is obtained via a geometric-component decomposition and spanning-tree counting.
Significance. If the proof is correct, this is a substantial contribution: it resolves an open problem of Robins and shows that the random-design minimax rate is strictly worse than the conjectured close-pair rate. The construction is novel and technically deep, combining harmonic analysis, random design perturbations, and high-dimensional probability. The paper is careful and honest about its scope: it explicitly states in §1.2 that the lower-bound construction uses error laws Q_x that vary with x, and therefore does not settle the common-error-law submodel. This transparency is a strength. The exponent arithmetic and the cancellation structure in the proof are internally consistent; my own checks of (1.3), (1.4), (2.3)–(2.4), (2.8), the variance-gap matching Δ = κ²δ, the V₀-term cancellation in Lemma 3.12, and the dyadic-shell bound in Lemma 3.4 did not reveal an error. The paper does not ship machine-checked proofs or code, but the argument is self-contained and the constants are tracked with explicit dependence on c₀, h, and N.
major comments (1)
- [§3.3.5, Proposition 3.8 and Lemma 3.4] The contraction fixed point is the linchpin of the whole construction: it alone realizes the pair-density identity R(x,y) = E[p_T(x)p_T(y)] (Eq. (3.74)). The proof of Proposition 3.8 derives the O(c₀)-Lipschitz estimates (3.81)–(3.82) by saying that summing (3.79) and (3.80) over indices j, taking the supremum in k, and applying (3.70)–(3.71) yields the result. This step is compressed: it implicitly uses the bounded-overlap constant D₀ from (3.28) and the fact that the number of j whose U_j¹ meets a fixed U_k² is uniformly bounded, but the constants and the summation are not shown. A hidden dependence on h, N, or c₀ in these estimates would invalidate the fixed point and hence the lower bound. I have not found a concrete error, but this is a load-bearing point and the proof should be expanded to make the Lipschitz constant explicit and independent of the chart index and scale.
minor comments (5)
- [§1.2] The sentence that the construction uses error laws Q_x that vary with x is important. I recommend adding a remark in the introduction after Theorem 1.1 stating explicitly that the common-error-law submodel remains open; currently this is only in the related-literature section.
- [§3.3.2, Lemma 3.4] The dyadic-shell estimate (3.46) is stated with a brief justification. In particular, the behavior of the cutoff ϑ(N/t ‖R‖) in the intermediate range N/2 < t < 4N is only discussed verbally. Expanding this derivation would help the reader verify the uniform derivative bounds.
- [§2, proof of Theorem 1.1] In (2.2), the exponent (4s−d)/(2(d+2s)) is negative exactly because d>4s. It would be helpful to point out this connection when introducing d>4s, since it also makes N_n → ∞ in (1.23).
- [Notation] The notation p and p̄ for the lower and upper density bounds is easy to confuse, especially in displayed inequalities. Consider using p_min and p_max or adding a one-line reminder after (1.5).
- [§3.3.5, Proposition 3.8] The completeness argument for the ball B_{M*} is written in a long paragraph. It is correct but could be shortened or made into a separate lemma for readability.
Circularity Check
No significant circularity: the lower bound is derived by an explicit two-prior construction; no fitted input is renamed as a prediction and no load-bearing self-citation is used.
full rationale
The paper proves Theorem 1.1 by Le Cam's two-prior method. The two priors are constructed explicitly: the variance gap Δ = c0^3 h^{2s} N^{-2} is a deterministic function of the construction constants, and the Hellinger bound is a theorem (Proposition 2.1) obtained from local moment identities, a contraction fixed point (Proposition 3.8), and component counting. No parameter is fitted to data and then called a prediction; the lower bound is the output of the construction, not an input. The one- and two-observation matching identities (1.12)–(1.13), (3.91), (3.112) are derived, not assumed as the target result. The paper contains no self-citations: reference [11] is the original problem source by Richardson and Rotnitzky, and the other references are external classical work. The construction's reliance on error laws Q_x varying with x is explicitly stated as a limitation, not hidden. The only residual risk is the unverified fixed-point estimates in Proposition 3.8, which is a correctness or verification concern, not circularity. Thus the appropriate score is 0.
Assumptions & free parameters
free parameters (4)
- c0 =
c0 ∈ (0, c*], chosen small enough that C1(c0^4 + c0^6) ≤ 1/16 (Eq. 2.5)
- h (localization scale) =
h ≍ n^{-3/[2(d+2s)]} (Eq. 1.22)
- N (thin-scale cutoff) =
N ≍ n^{(d-4s)/[2(d+2s)(d+4)]} (Eq. 1.23)
- V0 and a (base variance and response amplitude) =
V0 in interior of I ⊂ [v, min{v̄, √C4}]; V* < a^2 < C4/V*
assumptions (6)
- standard math Le Cam's two-prior method: minimax risk ≥ Bayes risk, and two priors with TV ≤ 1/4 and variance separation Δ force RMSE ≥ (√3/4)Δ (§2).
- standard math Banach fixed-point theorem applied to T in a complete local-Wiener-norm ball B_{M*} (Proposition 3.8).
- standard math Wiener-algebra estimates: ||KL||_W ≤ ||K||_W ||L||_W, completeness of W, and (3.33) bounding W-norms by smooth derivatives (Lemmas 3.3–3.5).
- standard math Wick/fourth-moment identity for the bounded coefficient vectors Wj with coordinates in {0, ±√3} (Lemma 3.9).
- domain assumption The design-density band contains 1: 0 < p < 1 < p̄ (1.5).
- domain assumption Conditional error laws Q_x may depend arbitrarily on x, subject only to mean-zero, common variance, and uniformly bounded fourth moment (condition (iii), §1.3).
Cite this review
Pith. "Pith review of On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins." pith.science (2026). https://pith.science/paper/IFGXYZ3T
@misc{pith2026260713170,
author = {Pith},
title = {Pith review of: On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins},
year = {2026},
howpublished = {\url{https://pith.science/paper/IFGXYZ3T}},
note = {Machine review of arXiv:2607.13170}
}
abstract
We give a negative answer to a problem posed by James Robins on estimating a constant conditional variance in nonparametric regression under random design. For every $s>1$ and integer $d>4s$, when the regression function is $s$-H\"older, the unknown design density is bounded above and away from zero, and the conditional error laws may depend on the design but have mean zero, a common variance, and uniformly bounded fourth moments, we show that the minimax root-mean-square risk is bounded below by $n^{-\beta}$ with $\beta=\frac{d(3s+1)+8s}{(d+2s)(d+4)}$. Hence the conjectured rate $n^{-4s/(d+4s)}$ is not uniformly attainable.
Reference graph
Works this paper leans on
-
[1]
Brown and Michael Levine,Variance estimation in nonparametric regression via the difference sequence method, The Annals of Statistics35(2007), no
Lawrence D. Brown and Michael Levine,Variance estimation in nonparametric regression via the difference sequence method, The Annals of Statistics35(2007), no. 5, 2219–2232
2007
-
[2]
Tony Cai, Michael Levine, and Lie Wang,Variance function estimation in multivariate nonparametric regression with fixed design, Journal of Multivariate Analysis100(2009), no
T. Tony Cai, Michael Levine, and Lie Wang,Variance function estimation in multivariate nonparametric regression with fixed design, Journal of Multivariate Analysis100(2009), no. 1, 126–136
2009
-
[3]
Donoho and Michael Nussbaum,Minimax quadratic estimation of a quadratic functional, Journal of Complexity6(1990), no
David L. Donoho and Michael Nussbaum,Minimax quadratic estimation of a quadratic functional, Journal of Complexity6(1990), no. 3, 290–323
1990
-
[4]
3, 645–660
Jianqing Fan and Qiwei Yao,Efficient estimation of conditional variance functions in stochastic regression, Biometrika85(1998), no. 3, 645–660
1998
-
[5]
Carroll,Variance function estimation in regression: The effect of estimating the mean, Journal of the Royal Statistical Society: Series B (Methodological)51(1989), no
Peter Hall and Raymond J. Carroll,Variance function estimation in regression: The effect of estimating the mean, Journal of the Royal Statistical Society: Series B (Methodological)51(1989), no. 1, 3–14
1989
-
[6]
Peter Hall, J. W. Kay, and D. M. Titterington,Asymptotically optimal difference-based estimation of variance in nonparametric regression, Biometrika77(1990), no. 3, 521–528
1990
-
[7]
5, 927–949
Li-Shan Huang and Jianqing Fan,Nonparametric estimation of quadratic regression functionals, Bernoulli5 (1999), no. 5, 927–949
1999
-
[8]
Robins, and Eric J
Lin Liu, Rajarshi Mukherjee, James M. Robins, and Eric J. Tchetgen Tchetgen,Adaptive estimation of nonparametric functionals, Journal of Machine Learning Research22(2021), no. 99, 1–66
2021
Show all 17 references
-
[9]
1, 19–41
Axel Munk, Nicolai Bissantz, Thorsten Wagner, and Gudrun Freitag,On difference-based variance estimation in nonparametric regression when the covariate is high dimensional, Journal of the Royal Statistical Society: Series B (Statistical Methodology)67(2005), no. 1, 19–41
2005
-
[10]
4, 1215–1230
John Rice,Bandwidth choice for nonparametric regression, The Annals of Statistics12(1984), no. 4, 1215–1230
1984
-
[11]
Richardson and Andrea Rotnitzky,Causal etiology of the research of James M
Thomas S. Richardson and Andrea Rotnitzky,Causal etiology of the research of James M. Robins, Statistical Science29(2014), no. 4, 459–484
2014
-
[12]
James Robins, Eric Tchetgen Tchetgen, Lingling Li, and Aad van der Vaart,Semiparametric minimax rates, Electronic Journal of Statistics3(2009), 1305–1321
2009
-
[13]
Robins, Lingling Li, Rajarshi Mukherjee, Eric J
James M. Robins, Lingling Li, Rajarshi Mukherjee, Eric J. Tchetgen Tchetgen, and Aad W. van der Vaart, Minimax estimation of a functional on a structured high-dimensional model, The Annals of Statistics45 (2017), no. 5, 1951–1987
2017
-
[14]
Robins, Lingling Li, Eric J
James M. Robins, Lingling Li, Eric J. Tchetgen Tchetgen, and Aad W. van der Vaart,Higher order influence functions and minimax estimation of nonlinear functionals, Probability and Statistics: Essays in Honor of David A. Freedman (Deborah Nolan and Terry Speed, eds.), Institute...
2008
-
[15]
6, 3589–3618
Yandi Shen, Chao Gao, Daniela Witten, and Fang Han,Optimal estimation of variance in nonparametric regression with random design, The Annals of Statistics48(2020), no. 6, 3589–3618
2020
-
[16]
Tsybakov,Introduction to nonparametric estimation, Springer Series in Statistics, Springer, New York, 2009
Alexandre B. Tsybakov,Introduction to nonparametric estimation, Springer Series in Statistics, Springer, New York, 2009
2009
-
[17]
Brown, T
Lie Wang, Lawrence D. Brown, T. Tony Cai, and Michael Levine,Effect of mean on variance function estimation in nonparametric regression, The Annals of Statistics36(2008), no. 2, 646–664
2008
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.