Pith. sign in

REVIEW 1 major objections 5 minor 17 references

On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins

T0 review · 1 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read This paper proves that in random-design nonparametric regression with s-Hölder regression functions and dimension d>4s, the minimax root-mean-square risk for estimating a constant conditional variance is at least n^{-β} with β=[d(3s+1)+8s]/

desk verdict A convincing negative answer to Robins' question, with an intricate construction whose main risk is verification depth rather than a found flaw. read the letter →

arxiv 2607.13170 v1 pith:IFGXYZ3T submitted 2026-07-14 math.ST stat.TH

classification math.STstat.TH MSC 62G0862C20
keywords varianceestimationnonparametricregressionrandomdesignminimaxlowerboundsHöldersmoothnessquadraticfunctionaltwo-priortestingsemiparametric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper settles a question about estimating a constant conditional variance when the covariate distribution is random and unknown. It proves that for regression functions with s Hölder smoothness and dimension d>4s, no estimator can achieve the previously conjectured root-mean-square rate n^{-4s/(d+4s)} uniformly over the model class. Instead the minimax risk is bounded below by n^{-β} with β=[d(3s+1)+8s]/[(d+2s)(d+4)], which lies strictly between the fixed-grid benchmark n^{-2s/d} and the conjecture. The proof builds two nearly indistinguishable data-generating processes with different error variances, hiding the variance change through specially tuned conditional error distributions and random design perturbations. The result matters because it closes a gap in the theory of semiparametric functional estimation: extra smoothness beyond one derivative is less useful under random design than was hoped.

What carries the argument

The proof uses a standard two-prior testing argument with two priors whose error variances differ by Δ but whose n-observation predictive laws are nearly indistinguishable. The central object is a three-point response distribution on {−a,0,a} whose probabilities are linear in q=f(x) and q²+σ², enabling exact cancellation of all one-observation differences through three moment identities. Random design perturbations are realized through a contraction fixed point: a provisional pair-density kernel R(x,y) is mapped to densities p_T satisfying E[p_T(x)p_T(y)]=R(x,y), via a projective decomposition into rank-one kernels and independent categorical variables. The load-bearing matrix identity is a

What would settle it

Check the contraction proof at the heart of the construction: on a discretized torus with small h and N=2, compute the map T in Proposition 3.8 at a small c0; if it fails to have a fixed point satisfying R(x,y)=E[p_T(x)p_T(y)], the cancellation identities fail and the lower bound collapses. Alternatively, any explicit estimator whose worst-case root-mean-square error is o(n^{-β}) on the stated class would refute the theorem.

Watch

Extended reading notes

Core claim

The central claim is that the minimax root-mean-square risk of estimating σ² in this random-design regression model is at least c n^{-β} for every n, with β=[d(3s+1)+8s]/[(d+2s)(d+4)]. Equivalently, the conjectured rate n^{-4s/(d+4s)} is uniformly unattainable: the gap between the two exponents is d(d−4s)(s−1)/[(d+4s)(d+2s)(d+4)], which is positive for every s>1 and d>4s. The paper does not determine the exact minimax rate, only rules out the conjectured one, leaving the true rate somewhere between the fixed-grid benchmark n^{-2s/d} and the conjecture.

Load-bearing premise

The whole argument hinges on allowing the noise distribution to be chosen afresh at every covariate value; if the conditional error laws had to be identical across the design, the indistinguishability construction is not known to work, and the paper explicitly does not claim the lower bound in that case.

Editorial extensions

If this is right

  • The conjectured n^{-4s/(d+4s)} rate is false for this entire model class: no estimator, however adaptive, attains it uniformly.
  • The true minimax rate lies strictly between n^{-2s/d} and n^{-4s/(d+4s)}; the paper's β gives a new lower bound that improves on the fixed-grid benchmark.
  • The obstruction is not an artifact of difference-based estimators; it applies to all measurable estimators of the variance.
  • The lower-bound construction does not apply when conditional error laws are forced to be identical across x; under a common error law independent of X, the question remains open.
  • For s≤1 or d≤4s, the argument gives no information, so the conjectured rate may still be attainable in those regimes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the x-dependent error-law freedom is essential, then datasets with homogeneous noise may still admit the conjectured rate; the hardness may be driven by heteroscedasticity rather than by random design alone.
  • The pair-density fixed-point device is general: any exact cancellation requiring E[p_T(x)p_T(y)]=R(x,y) can be produced by the same contraction, so the technique may transfer to other quadratic semiparametric functionals.
  • The exponent β suggests a transition as d approaches 4s; a natural next step is to test numerically whether the lower bound is tight in moderate dimensions or to look for matching upper bounds.
  • Because the gap between the conjectured rate and β shrinks as s approaches 1, the phenomenon may be specific to genuine Hölder smoothness with s>1 rather than to Lipschitz-scale smoothness.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper studies the minimax rate for estimating a constant conditional variance σ² in nonparametric regression under random design. The class Θ allows an unknown design density bounded above and away from zero, an s-Hölder regression function, and conditional error laws that may depend on x but have mean zero, common variance σ², and uniformly bounded fourth moments. The main result, Theorem 1.1, states that for every s>1 and integer d>4s, the minimax root-mean-square risk satisfies R_n(Θ)^{1/2} ≥ c n^{-β} with β = [d(3s+1)+8s]/[(d+2s)(d+4)]. Since 4s/(d+4s) − β > 0, the conjectured rate n^{-4s/(d+4s)} is not uniformly attainable, and the minimax rate lies strictly between the regular-grid benchmark n^{-2s/d} and the conjectured rate. The proof uses Le Cam's two-prior method, constructing priors with variance gap Δ and nearly indistinguishable n-observation predictive laws. The construction combines periodic feature systems, a localized covariance dual, random design perturbations realized through a contraction fixed point in a Wiener-norm ball, and bounded coefficient vectors satisfying Gaussian fourth-moment identities. The full-sample Hellinger bound is obtained via a geometric-component decomposition and spanning-tree counting.

Significance. If the proof is correct, this is a substantial contribution: it resolves an open problem of Robins and shows that the random-design minimax rate is strictly worse than the conjectured close-pair rate. The construction is novel and technically deep, combining harmonic analysis, random design perturbations, and high-dimensional probability. The paper is careful and honest about its scope: it explicitly states in §1.2 that the lower-bound construction uses error laws Q_x that vary with x, and therefore does not settle the common-error-law submodel. This transparency is a strength. The exponent arithmetic and the cancellation structure in the proof are internally consistent; my own checks of (1.3), (1.4), (2.3)–(2.4), (2.8), the variance-gap matching Δ = κ²δ, the V₀-term cancellation in Lemma 3.12, and the dyadic-shell bound in Lemma 3.4 did not reveal an error. The paper does not ship machine-checked proofs or code, but the argument is self-contained and the constants are tracked with explicit dependence on c₀, h, and N.

major comments (1)
  1. [§3.3.5, Proposition 3.8 and Lemma 3.4] The contraction fixed point is the linchpin of the whole construction: it alone realizes the pair-density identity R(x,y) = E[p_T(x)p_T(y)] (Eq. (3.74)). The proof of Proposition 3.8 derives the O(c₀)-Lipschitz estimates (3.81)–(3.82) by saying that summing (3.79) and (3.80) over indices j, taking the supremum in k, and applying (3.70)–(3.71) yields the result. This step is compressed: it implicitly uses the bounded-overlap constant D₀ from (3.28) and the fact that the number of j whose U_j¹ meets a fixed U_k² is uniformly bounded, but the constants and the summation are not shown. A hidden dependence on h, N, or c₀ in these estimates would invalidate the fixed point and hence the lower bound. I have not found a concrete error, but this is a load-bearing point and the proof should be expanded to make the Lipschitz constant explicit and independent of the chart index and scale.
minor comments (5)
  1. [§1.2] The sentence that the construction uses error laws Q_x that vary with x is important. I recommend adding a remark in the introduction after Theorem 1.1 stating explicitly that the common-error-law submodel remains open; currently this is only in the related-literature section.
  2. [§3.3.2, Lemma 3.4] The dyadic-shell estimate (3.46) is stated with a brief justification. In particular, the behavior of the cutoff ϑ(N/t ‖R‖) in the intermediate range N/2 < t < 4N is only discussed verbally. Expanding this derivation would help the reader verify the uniform derivative bounds.
  3. [§2, proof of Theorem 1.1] In (2.2), the exponent (4s−d)/(2(d+2s)) is negative exactly because d>4s. It would be helpful to point out this connection when introducing d>4s, since it also makes N_n → ∞ in (1.23).
  4. [Notation] The notation p and p̄ for the lower and upper density bounds is easy to confuse, especially in displayed inequalities. Consider using p_min and p_max or adding a one-line reminder after (1.5).
  5. [§3.3.5, Proposition 3.8] The completeness argument for the ball B_{M*} is written in a long paragraph. It is correct but could be shortened or made into a separate lemma for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the lower bound is derived by an explicit two-prior construction; no fitted input is renamed as a prediction and no load-bearing self-citation is used.

full rationale

The paper proves Theorem 1.1 by Le Cam's two-prior method. The two priors are constructed explicitly: the variance gap Δ = c0^3 h^{2s} N^{-2} is a deterministic function of the construction constants, and the Hellinger bound is a theorem (Proposition 2.1) obtained from local moment identities, a contraction fixed point (Proposition 3.8), and component counting. No parameter is fitted to data and then called a prediction; the lower bound is the output of the construction, not an input. The one- and two-observation matching identities (1.12)–(1.13), (3.91), (3.112) are derived, not assumed as the target result. The paper contains no self-citations: reference [11] is the original problem source by Richardson and Rotnitzky, and the other references are external classical work. The construction's reliance on error laws Q_x varying with x is explicitly stated as a limitation, not hidden. The only residual risk is the unverified fixed-point estimates in Proposition 3.8, which is a correctness or verification concern, not circularity. Thus the appropriate score is 0.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

All parameters listed are minimax-construction choices, not fits to data: they tune the two adversarial hypotheses so that the Hellinger distance is small while the variance gap Δ is large. That is the standard role of scale parameters in Le Cam lower bounds and creates no circularity: the theorem's bound is a consequence of the explicit construction, not an input to it. The domain assumptions are the theorem's hypotheses; the most consequential is the allowance of x-dependent error laws, which the paper honestly disclaims for the common-error-law variant. No new physical or model entities are postulated; the priors Π0, Π1 and latent variables (Tj, Wj) are proof devices inside a lower-bound construction.

free parameters (4)
  • c0 = c0 ∈ (0, c*], chosen small enough that C1(c0^4 + c0^6) ≤ 1/16 (Eq. 2.5)
    Scales the regression amplitude κ = c0 h^s and the variance gap Δ = c0^3 h^{2s} N^{-2}; chosen to make the Hellinger distance between the two priors small. A lower-bound construction constant, not a data fit.
  • h (localization scale) = h ≍ n^{-3/[2(d+2s)]} (Eq. 1.22)
    Grid spacing for the feature system; chosen as the largest h allowed by the constraint n^3 h^{2d+4s} ≲ 1 (Eq. 1.20). Standard hardest-hypothesis scale choice.
  • N (thin-scale cutoff) = N ≍ n^{(d-4s)/[2(d+2s)(d+4)]} (Eq. 1.23)
    Sets the thickness h/N of the two-point exceptional region; chosen as the smallest N satisfying the second Hellinger constraint, maximizing Δ. The condition d>4s makes the exponent positive.
  • V0 and a (base variance and response amplitude) = V0 in interior of I ⊂ [v, min{v̄, √C4}]; V* < a^2 < C4/V*
    Chosen so the three-point channel probabilities are uniformly positive and the conditional fourth moment stays below C4 (§3.5). Construction tuning, not data fitting.
assumptions (6)
  • standard math Le Cam's two-prior method: minimax risk ≥ Bayes risk, and two priors with TV ≤ 1/4 and variance separation Δ force RMSE ≥ (√3/4)Δ (§2).
    Invoked in the proof of Theorem 1.1; standard lower-bound machinery from Tsybakov (2009), cited as [16].
  • standard math Banach fixed-point theorem applied to T in a complete local-Wiener-norm ball B_{M*} (Proposition 3.8).
    Produces the pair-density fixed point R(x,y) = E[p_T(x)p_T(y)]; requires the O(c0)-Lipschitz estimates (3.81)–(3.82).
  • standard math Wiener-algebra estimates: ||KL||_W ≤ ||K||_W ||L||_W, completeness of W, and (3.33) bounding W-norms by smooth derivatives (Lemmas 3.3–3.5).
    The engine of the N²-singularity control in Lemma 3.4 and of the projective decomposition in Lemma 3.5.
  • standard math Wick/fourth-moment identity for the bounded coefficient vectors Wj with coordinates in {0, ±√3} (Lemma 3.9).
    Ensures the two-point fourth-order difference reduces to the covariance displacement (3.100)–(3.101); verified algebraically in Lemma 3.10.
  • domain assumption The design-density band contains 1: 0 < p < 1 < p̄ (1.5).
    The perturbed densities are p_T = 1 + ε Σ A_j with zero-integral A_j ((3.29), Prop 3.8); the uniform baseline and mass preservation require 1 ∈ [p, p̄]. The theorem is only stated for such bands.
  • domain assumption Conditional error laws Q_x may depend arbitrarily on x, subject only to mean-zero, common variance, and uniformly bounded fourth moment (condition (iii), §1.3).
    The pointwise matching (1.10)–(1.13) exploits this freedom; the paper explicitly does not cover a common error law independent of X (§1.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins." pith.science (2026). https://pith.science/paper/IFGXYZ3T

@misc{pith2026260713170,
  author       = {Pith},
  title        = {Pith review of: On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFGXYZ3T}},
  note         = {Machine review of arXiv:2607.13170}
}
abstract

We give a negative answer to a problem posed by James Robins on estimating a constant conditional variance in nonparametric regression under random design. For every $s>1$ and integer $d>4s$, when the regression function is $s$-H\"older, the unknown design density is bounded above and away from zero, and the conditional error laws may depend on the design but have mean zero, a common variance, and uniformly bounded fourth moments, we show that the minimax root-mean-square risk is bounded below by $n^{-\beta}$ with $\beta=\frac{d(3s+1)+8s}{(d+2s)(d+4)}$. Hence the conjectured rate $n^{-4s/(d+4s)}$ is not uniformly attainable.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

17 extracted references

  1. [1]

    Brown and Michael Levine,Variance estimation in nonparametric regression via the difference sequence method, The Annals of Statistics35(2007), no

    Lawrence D. Brown and Michael Levine,Variance estimation in nonparametric regression via the difference sequence method, The Annals of Statistics35(2007), no. 5, 2219–2232

  2. [2]

    Tony Cai, Michael Levine, and Lie Wang,Variance function estimation in multivariate nonparametric regression with fixed design, Journal of Multivariate Analysis100(2009), no

    T. Tony Cai, Michael Levine, and Lie Wang,Variance function estimation in multivariate nonparametric regression with fixed design, Journal of Multivariate Analysis100(2009), no. 1, 126–136

  3. [3]

    Donoho and Michael Nussbaum,Minimax quadratic estimation of a quadratic functional, Journal of Complexity6(1990), no

    David L. Donoho and Michael Nussbaum,Minimax quadratic estimation of a quadratic functional, Journal of Complexity6(1990), no. 3, 290–323

  4. [4]

    3, 645–660

    Jianqing Fan and Qiwei Yao,Efficient estimation of conditional variance functions in stochastic regression, Biometrika85(1998), no. 3, 645–660

  5. [5]

    Carroll,Variance function estimation in regression: The effect of estimating the mean, Journal of the Royal Statistical Society: Series B (Methodological)51(1989), no

    Peter Hall and Raymond J. Carroll,Variance function estimation in regression: The effect of estimating the mean, Journal of the Royal Statistical Society: Series B (Methodological)51(1989), no. 1, 3–14

  6. [6]

    Peter Hall, J. W. Kay, and D. M. Titterington,Asymptotically optimal difference-based estimation of variance in nonparametric regression, Biometrika77(1990), no. 3, 521–528

  7. [7]

    5, 927–949

    Li-Shan Huang and Jianqing Fan,Nonparametric estimation of quadratic regression functionals, Bernoulli5 (1999), no. 5, 927–949

  8. [8]

    Robins, and Eric J

    Lin Liu, Rajarshi Mukherjee, James M. Robins, and Eric J. Tchetgen Tchetgen,Adaptive estimation of nonparametric functionals, Journal of Machine Learning Research22(2021), no. 99, 1–66

Show all 17 references
  1. [9]

    1, 19–41

    Axel Munk, Nicolai Bissantz, Thorsten Wagner, and Gudrun Freitag,On difference-based variance estimation in nonparametric regression when the covariate is high dimensional, Journal of the Royal Statistical Society: Series B (Statistical Methodology)67(2005), no. 1, 19–41

  2. [10]

    4, 1215–1230

    John Rice,Bandwidth choice for nonparametric regression, The Annals of Statistics12(1984), no. 4, 1215–1230

  3. [11]

    Richardson and Andrea Rotnitzky,Causal etiology of the research of James M

    Thomas S. Richardson and Andrea Rotnitzky,Causal etiology of the research of James M. Robins, Statistical Science29(2014), no. 4, 459–484

  4. [12]

    James Robins, Eric Tchetgen Tchetgen, Lingling Li, and Aad van der Vaart,Semiparametric minimax rates, Electronic Journal of Statistics3(2009), 1305–1321

  5. [13]

    Robins, Lingling Li, Rajarshi Mukherjee, Eric J

    James M. Robins, Lingling Li, Rajarshi Mukherjee, Eric J. Tchetgen Tchetgen, and Aad W. van der Vaart, Minimax estimation of a functional on a structured high-dimensional model, The Annals of Statistics45 (2017), no. 5, 1951–1987

  6. [14]

    Robins, Lingling Li, Eric J

    James M. Robins, Lingling Li, Eric J. Tchetgen Tchetgen, and Aad W. van der Vaart,Higher order influence functions and minimax estimation of nonlinear functionals, Probability and Statistics: Essays in Honor of David A. Freedman (Deborah Nolan and Terry Speed, eds.), Institute...

  7. [15]

    6, 3589–3618

    Yandi Shen, Chao Gao, Daniela Witten, and Fang Han,Optimal estimation of variance in nonparametric regression with random design, The Annals of Statistics48(2020), no. 6, 3589–3618

  8. [16]

    Tsybakov,Introduction to nonparametric estimation, Springer Series in Statistics, Springer, New York, 2009

    Alexandre B. Tsybakov,Introduction to nonparametric estimation, Springer Series in Statistics, Springer, New York, 2009

  9. [17]

    Brown, T

    Lie Wang, Lawrence D. Brown, T. Tony Cai, and Michael Levine,Effect of mean on variance function estimation in nonparametric regression, The Annals of Statistics36(2008), no. 2, 646–664

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.