REVIEW 1 major objections 4 minor 20 references
Estimating a regression function under possible heteroscedastic and heavy-tailed errors. Application to shape-restricted regression
T0 review · 1 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read New estimator stays consistent where least squares fails
desk verdict A substantial new oracle inequality for robust regression, but the main theorem's universal 'every F' claim needs a permissibility condition or outer expectations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the test statistic $T(Z,f,g)=\sum_{i=1}^n (f(x_i)-Y_i)\operatorname{sgn}(f(x_i),g(x_i))$, whose expectation is controlled by the $\ell_1$ distances between $f^\star$, $f$ and $g$. The estimator is an approximate minimiser of $T(Z,f)=\sup_{g\in F} T(Z,f,g)$ over $F$. The argument is carried by the sets $O(D)$ of functions whose upper and lower level sets form VC classes of dimension at most $D$, together with a maximal inequality (Lemma 3) bounding the expected supremum of an empirical process over such a class by $\kappa n \sigma_p(x)(D/n)^{1-1/p}$ using symmetrisation, VC covering numbers and truncation.
What would settle it
Simulate the paper's Section 2.3 example, with one observation having variance $n\log^2 n$ and all others noiseless: the new estimator's $\ell_1$ risk should scale as $(\log n)/\sqrt{n}$, while the LSE risk stays bounded away from zero. If the observed risk instead scales with the maximal variance, the oracle inequality is wrong. Alternatively, a direct counterexample to inequality (55) for a non-measurable VC-subgraph class would falsify Lemma 3.
Extended reading notes
Core claim
The paper establishes that the minimiser of the sup-norm of a signed test statistic, placed in a model class $F$, achieves for every $D$ a risk bound $E[\ell(f^\star, \hat f)] \le 3\ell(f^\star, O(D)) + \kappa \inf_{p\in[1,2]} \sigma_p(x)(D/n)^{1-1/p} + c/n$, where $O(D)$ consists of functions whose upper and lower level sets form VC classes of dimension at most $D$. This shows the estimator automatically balances approximation and stochastic error, and that the stochastic error depends only on the average $p$-th absolute moment of the errors, not their maximum, which is what makes it stable under heteroscedasticity. For shape-restricted classes such as piecewise monotone or piecewise convex-concave functions, the paper derives explicit rates by combining the oracle inequality with approximation results by step functions and splines, and by bounding the degrees of extremal functions via VC arguments.
Load-bearing premise
The theorem silently assumes that the class of candidate functions is 'permissible' in the empirical-process sense, so that the suprema appearing in the proof are measurable; the paper does not state this condition.
Editorial extensions
If this is right
- For monotone or convex regression, the estimator matches the known LSE adaptation rates while requiring only a finite $p$-th moment of the errors for some $p\in[1,2]$.
- In heteroscedastic settings where the average variance is small but the maximum is large, the estimator remains consistent and can achieve rate $(\log n)/\sqrt{n}$ where the LSE is inconsistent.
- The oracle inequality yields risk bounds for piecewise monotone and piecewise convex/concave models, including S-shaped functions, with explicit approximation rates for splines of degree 0 and 1.
- The same method applies to logistic and Poisson regression, which are heteroscedastic by nature, because the bound depends on the average of the variances.
- The approximation results and VC dimension bounds for level sets may be usable in other nonparametric estimation problems.
Reading between the lines
- The dependence on the average moment suggests that a single large outlier inflates the risk far less than for the LSE; one could test this by designing a regression with one high-variance point and comparing the two estimators.
- The automatic adaptation to $p$ through the infimum over $p$ hints at a general principle: test-based estimators with bounded-expectation statistics can adapt to the heaviest existing moment without tuning, which may extend to density estimation and other inverse problems.
- The measurability caveat in Lemma 3 could be resolved by explicitly restricting $F$ to permissible classes; the practical impact is likely negligible for the shape-restricted examples, but a fully rigorous statement would need to add this condition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies fixed-design regression Y_i=f*(x_i)+ξ_i under independent errors with only a p-th moment for some unknown p∈[1,2], allowing heteroscedasticity and heavy tails. The authors define an estimator as an approximate minimizer of T(Z,f)=sup_{g∈F} T(Z,f,g), where T(Z,f,g)=∑(f(x_i)-Y_i)sgn(f-g). The main result (Theorem 1, Eqs. (14)–(15)) is a non-asymptotic oracle inequality for the ℓ_1 loss: for any f∈O(D), Eℓ(bf,f)≤2ℓ(f*,f)+κ inf_{p∈[1,2]} σ_p(x)(D/n)^{1-1/p}+c/n, where O(D) is the set of functions whose subgraph comparison classes have VC dimension at most D. The bound yields adaptation to unknown error integrability and to shape constraints. Applications are developed for k-piecewise monotone, k-piecewise convex-concave, and monotone single-index classes, with supporting extremal-degree and spline approximation results (Theorems 2–3, Propositions 3–8).
Significance. If correct, the paper offers a genuinely new alternative to least squares in shape-restricted and heteroscedastic regression. The risk bound depends on the average of p-th absolute moments rather than their maximum, which is what makes the estimator consistent in examples where the LSE is not, and the infimum over p provides adaptation to unknown heaviness of the tails. The proof strategy—combining a testing-based surrogate loss with VC-dimension bounds for extremal subgraph classes—is coherent, and the paper contains complete proofs of the supporting approximation and VC results. The principal weakness is the unstated measurability condition discussed below; once that is fixed, the results are strong and publishable.
major comments (1)
- [Section 6, Lemma 3; Section 2.2, Eqs. (6)–(7); Theorem 1, Eq. (14)] The proof of Lemma 3 and Theorem 1 uses ordinary expectations of the quantities sup_{g∈F} eT+(Z,f,g) and sup_{g∈F} eT-(Z,g,f) for an arbitrary class F, without any measurability or permissibility condition. For arbitrary F the map g↦T(Z,f,g) is not necessarily jointly measurable, so the suprema in (6) and in Lemma 3 (display near (54)) may fail to be measurable, and the random set E_c(Z,F) in (7) need not admit a measurable selection bf; consequently E[ℓ(f*,bf)] in (14) need not be defined. This is not merely a request for outer expectations: the theorem claims a bound for every class F and every f∈O(D), so the proof as written has a genuine gap. The fix is to state explicitly that F is permissible in the sense of Pollard (1984), or to formulate the result with outer expectations as in van der Vaart and Wellner (1996), Section 1.2. The concrete classes in Section 4 are likely permissible, but the universal statements in Theorem 1 and Lemma 3 need this qualification.
minor comments (4)
- [Section 2.2, Eq. (7)] The definition of E_c(Z,F) does not guarantee existence of an element when c=0 and the infimum over F is not attained; the theorem should either require c>0 or explicitly assume attainment or an approximate minimizer.
- [Throughout] There are small typographical errors and OCR artifacts (e.g., 'A interesting consequence' near Eq. (15), 'wihout' in the proof of Theorem 2, and '2β' for 2^β in Example 6); these should be corrected in the final version.
- [Section 4.3, Corollary 5] In the displayed consequence for p=2, the symbol σ_p appears in the second term where σ_2 is meant; the notation should be made consistent.
- [Section 2.2] The paper does not discuss the computational implementation of the approximate minimizer in (7) for the infinite-dimensional shape-constrained classes; a brief comment on algorithmic feasibility would be useful for practitioners.
Circularity Check
No significant circularity; the oracle inequality and adaptation claims are derived self-contained from the stated construction.
full rationale
After walking the derivation chain from the test statistic T(Z,f,g) in Section 2.2 through Lemma 1, the key inequality (53), and Lemma 3 in Section 6, I find no step in which a claimed prediction or oracle inequality is equivalent by construction to a fitted input or to an author-supplied ansatz. The estimator is explicitly defined by (6)-(7), and Theorem 1 is proved rather than assumed. The maximal inequality (55) is obtained by standard symmetrization and VC chaining, with the dimension D entering only through the definition of O(D) and the covering-number bound; D is not a parameter fitted to the data or to the target. Adaptation to p and sigma_p is a genuine infimum over bounds that hold simultaneously for every p in [1,2], not a fitted quantity. The comparisons to the LSE in Sections 2.3 and 4 are external illustrations, not ingredients of the proof. The only caveat worth flagging is a technical rigor issue, not a circularity: Lemma 3 and Theorem 1 state ordinary expectations of suprema over an arbitrary class F without an explicit measurability or permissibility condition (Section 6, Lemma 3, display near eq. (55); see also the supremum in eq. (6)-(7)). This can be repaired by standard outer-expectation or permissibility assumptions and does not make the theorem's conclusion identical to its hypotheses. Self-citations to Baraud-Birgé and Baraud-Halconruy-Maillard appear as methodological lineage only; the main proof does not load-bearingly rely on those papers for any unproved step.
Assumptions & free parameters
free parameters (1)
- c =
unspecified small constant >= 0
assumptions (5)
- domain assumption The errors ξ_i,n are independent and centered with finite p-th absolute moment for some p∈[1,2].
- ad hoc to paper The class F is such that suprema over F of the test statistic are measurable (or a permissible class).
- standard math Standard empirical process inequalities: symmetrization, Dudley entropy bound, and VC covering number bounds.
- domain assumption The approximate minimizer set E_c(Z,F) is nonempty and contains the estimator f̂.
- domain assumption The design points are deterministic.
Cite this review
Pith. "Pith review of Estimating a regression function under possible heteroscedastic and heavy-tailed errors. Application to shape-restricted regression." pith.science (2026). https://pith.science/paper/FIQ56IBV
@misc{pith2026250600852,
author = {Pith},
title = {Pith review of: Estimating a regression function under possible heteroscedastic and heavy-tailed errors. Application to shape-restricted regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/FIQ56IBV}},
note = {Machine review of arXiv:2506.00852}
}
abstract
We consider a regression framework where the design points are deterministic and the errors possibly non-i.i.d. and heavy-tailed (with a moment of order $p$ in $[1,2]$). Given a class of candidate regression functions, we propose a surrogate for the classical least squares estimator (LSE). For this new estimator, we establish a nonasymptotic risk bound with respect to the absolute loss which takes the form of an oracle type inequality. This inequality shows that our estimator possesses natural adaptation properties with respect to some elements of the class. When this class consists of monotone functions or convex functions on an interval, these adaptation properties are similar to those established in the literature for the LSE. However, unlike the LSE, we prove that our estimator remains stable with respect to a possible heteroscedasticity of the errors and may even converge at a parametric rate (up to a logarithmic factor) when the LSE is not even consistent. We illustrate the performance of this new estimator over classes of regression functions that satisfy a shape constraint: piecewise monotone, piecewise convex/concave, among other examples. The paper also contains some approximation results by splines with degrees in $\{0,1\}$ and VC bounds for the dimensions of classes of level sets. These results may be of independent interest.
Reference graph
Works this paper leans on
-
[1]
and Birg\'e, L
Baraud, Y. and Birg\'e, L. (2016). Rho-estimators for shape restricted density estimation. Stochastic Process. Appl. , 126(12):3888--3912
2016
-
[2]
and Chen, J
Baraud, Y. and Chen, J. (2024). Robust estimation of a regression function in exponential families. J. Statist. Plann. Inference , 233:Paper No. 106167, 25
2024
-
[3]
Baraud, Y., Halconruy, H., and Maillard, G. (2022). Robust density estimation with the L _1 -loss. A pplications to the estimation of a density on the line satisfying a shape constraint. to appear in Ann. Inst. H. Poincar \'e Probab. Statist
work page 2022
-
[4]
Baum, L. E. and Katz, M. (1965). Convergence rates in the law of large numbers. Trans. Amer. Math. Soc. , 120:108--123
work page 1965
-
[5]
Bellec, P. C. (2018). Sharp oracle inequalities for Least Squares estimators in shape restricted regression . The Annals of Statistics , 46(2):745 -- 780
work page 2018
-
[6]
Birg \'e , L. and Massart, P. (2001). Gaussian model selection. J. Eur. Math. Soc. (JEMS) , 3(3):203--268
work page 2001
-
[7]
Chatterjee, S. (2014). A new perspective on least squares under convex constraint. Ann. Statist. , 42(6):2340--2381
work page 2014
-
[8]
Chatterjee, S. (2016). An improved global risk bound in concave regression . Electronic Journal of Statistics , 10(1):1608 -- 1629
work page 2016
Show all 20 references
-
[9]
Chatterjee, S., Guntuboyina, A., and Sen, B. (2015). On risk bounds in isotonic and other shape restricted regression problems. Ann. Statist. , 43(4):1774--1800
2015
-
[10]
and Lafferty, J
Chatterjee, S. and Lafferty, J. (2019). Adaptive risk bounds in unimodal regression . Bernoulli , 25(1):1 -- 25
2019
-
[11]
Durrett, R. (2019). Probability---theory and examples , volume 49 of Cambridge Series in Statistical and Probabilistic Mathematics . Cambridge University Press, Cambridge
2019
-
[12]
and Pinsker, M
Efromovich, S. and Pinsker, M. (1996). Sharp-optimal and adaptive estimation for heteroscedastic nonparametric regression . Statistica Sinica , 6:925--942
1996
-
[13]
Y., Chen, Y., Han, Q., Carroll, R
Feng, O. Y., Chen, Y., Han, Q., Carroll, R. J., and Samworth, R. J. (2022). Nonparametric, tuning-free estimation of s-shaped functions. Journal of the Royal Statistical Society Series B: Statistical Methodology , 84(4):1324--1352
2022
-
[14]
and Sen, B
Guntuboyina, A. and Sen, B. (2018). Nonparametric Shape-Restricted Regression . Statistical Science , 33(4):568 -- 594
2018
-
[15]
Koltchinskii, V. (2011). Oracle Inequalities in Empirical Risk minimization and Sparse Recovery Problems . Lectures from the 38th Summer School on Probability Theory held in Saint-Flour, 2008. Springer
2011
-
[16]
and Talagrand, M
Ledoux, M. and Talagrand, M. (2011). Probability in B anach spaces . Classics in Mathematics. Springer-Verlag, Berlin. Isoperimetry and processes, Reprint of the 1991 edition
2011
-
[17]
Minami, K. (2020). Estimating piecewise monotone signals . Electronic Journal of Statistics , 14(1):1508 -- 1576
2020
-
[18]
Pollard, D. (1984). Convergence of stochastic processes . Springer Series in Statistics. Springer-Verlag, New York
1984
-
[19]
van der Vaart, A. W. and Wellner, J. A. (1996). Weak Convergence and Empirical Processes. With Applications to Statistics . Springer Series in Statistics. Springer-Verlag, New York
1996
-
[20]
Zhang, C.-H. (2002). Risk bounds in isotonic regression . The Annals of Statistics , 30(2):528 -- 555
2002
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.