REVIEW 1 major objections 15 references
Leveraging tails for adaptation
T0 review · 1 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Priors with p-exponential tails on coefficients achieve improving contraction rates and full adaptation to smoothness as p goes to zero.
desk verdict The paper shows p-exponential tails with p<1 improve contraction rates over Laplace and yield adaptation in the p to 0 limit, including a claim for overparametrized shallow ReLU nets up to beta=2. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
p-exponential tail priors placed independently on the coefficients of a fixed basis or dictionary
What would settle it
A concrete counterexample would be a function in a Sobolev ball whose posterior fails to contract at the predicted adaptive rate when p is taken small, or where an overparametrized shallow ReLU network posterior does not adapt beyond regularity level 2.
Extended reading notes
Core claim
Placing independent priors with p-exponential tails on the coefficients of a fixed basis or dictionary produces posterior contraction rates that improve as p decreases. In an appropriate regime as p tends to zero, this setup yields full adaptation to the unknown smoothness parameter, up to logarithmic factors. The same mechanism is applied to shallow ReLU neural networks, showing that overparametrized versions adapt to any regularity 0 ≤ β ≤ 2 in random design regression.
Load-bearing premise
The true function belongs to a Sobolev or Besov-type ball of unknown radius, and the priors are placed independently on coefficients of a fixed basis or dictionary.
Editorial extensions
If this is right
- Posterior contraction rates improve as the tail parameter p is decreased.
- Full adaptation to unknown smoothness holds up to logarithmic factors when p approaches zero.
- Overparametrized shallow ReLU networks adapt to any regularity 0 ≤ β ≤ 2.
- Simulation studies show agreement between observed behavior and the predicted rates.
Reading between the lines
- The tail-based mechanism could be examined in related nonparametric problems such as density estimation.
- One might test whether the adaptation extends when the basis or dictionary is allowed to vary with sample size.
- Practitioners could explore very small p values in software implementations to check empirical adaptation in moderate sample sizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies Bayesian posterior contraction in nonparametric regression using priors with p-exponential tails (including Laplace and heavier tails for p<1) placed independently on coefficients of a fixed basis or dictionary. It claims that contraction rates improve as p decreases and that an appropriate p→0 regime yields full adaptation to unknown smoothness up to logarithmic factors. Applications are given to series priors in white noise regression and to shallow ReLU networks in random design regression, with the specific claim that overparametrized shallow ReLU networks adapt to any regularity 0≤β≤2. A simulation study is included to support the theory.
Significance. If the results hold, the work would contribute to Bayesian nonparametrics by showing how tail choice alone can deliver adaptation without hyperpriors on smoothness. The neural-network application, if rigorously justified, would be notable for linking tail-based priors to overparametrized ReLU models. The empirical component provides concrete support for the predicted behavior.
major comments (1)
- [Application to shallow ReLU networks (abstract and corresponding theorems)] The core contraction theorems are proved under independent coefficient priors on a fixed countable dictionary with controlled metric entropy. The claim that overparametrized shallow ReLU networks adapt to 0≤β≤2 requires an explicit reduction showing that the continuous-parameter prior on hidden weights satisfies the same tail and entropy conditions uniformly in the unknown β; without this reduction the adaptation statement for networks rests on an unverified modeling equivalence between the fixed-dictionary setting and the continuous-atom setting.
Simulated Author's Rebuttal
We thank the referee for their careful reading and for identifying a key point regarding the rigor of the ReLU network application. We address the concern directly below and commit to a targeted revision that supplies the missing explicit reduction.
read point-by-point responses
-
Referee: [Application to shallow ReLU networks (abstract and corresponding theorems)] The core contraction theorems are proved under independent coefficient priors on a fixed countable dictionary with controlled metric entropy. The claim that overparametrized shallow ReLU networks adapt to 0≤β≤2 requires an explicit reduction showing that the continuous-parameter prior on hidden weights satisfies the same tail and entropy conditions uniformly in the unknown β; without this reduction the adaptation statement for networks rests on an unverified modeling equivalence between the fixed-dictionary setting and the continuous-atom setting.
Authors: We agree that the general theorems are stated for a fixed countable dictionary and that the network claim requires a separate verification step. In the revised manuscript we will insert a new subsection (and supporting appendix) that explicitly reduces the continuous-weight prior to the dictionary setting. Concretely, we will (i) derive the marginal prior on the effective coefficients obtained by integrating the p-exponential prior over the hidden weights, (ii) verify that this marginal satisfies the required p-exponential tail bound uniformly in β, and (iii) establish a uniform bound on the metric entropy of the resulting function class for β ∈ [0,2]. With these steps the adaptation statement will rest on a verified reduction rather than an implicit modeling equivalence. revision: yes
Circularity Check
No circularity; derivation applies standard contraction theory to new tail family
full rationale
The paper derives posterior contraction rates for p-exponential tail priors on fixed-dictionary coefficients via standard arguments, then applies the same framework to series priors and ReLU networks. No quoted step reduces a claimed prediction to a fitted parameter by construction, invokes a self-citation as the sole justification for a uniqueness claim, or renames an input as an output. The modeling assumption that overparametrized ReLU networks can be treated as linear combinations over a countable dictionary is an external modeling choice rather than a circular reduction inside the derivation. The result is therefore self-contained against the stated assumptions.
Assumptions & free parameters
assumptions (2)
- domain assumption The true function lies in a Sobolev or Besov ball of unknown radius and smoothness.
- domain assumption Coefficients receive independent p-exponential priors.
Cite this review
Pith. "Pith review of Leveraging tails for adaptation." pith.science (2026). https://pith.science/paper/Z2H4STXJ
@misc{pith2026260620480,
author = {Pith},
title = {Pith review of: Leveraging tails for adaptation},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z2H4STXJ}},
note = {Machine review of arXiv:2606.20480}
}
abstract
We consider contraction of Bayesian posterior distributions in nonparametric settings where coefficients of a function over a basis or dictionary are given priors with $p$--exponential tails, including Laplace tails $(p=1)$ and heavier tails $(p<1)$. It is shown that contraction rates improve as $p$ decreases and that full adaptation to smoothness, up to logarithmic factors, is obtained in an appropriate $p\to 0$ regime. As applications, we consider both series priors in white noise regression and shallow ReLU neural networks in random design regression. In particular, we show that overparametrised shallow ReLU networks can adapt to any regularity $0\le \beta\le 2$. Through a simulation study, we show strong empirical agreement with the behavior predicted by our theory.
Reference graph
Works this paper leans on
-
[1]
Usingulog 1 u ≤1 +u, available for all u∈(0,1)we obtainψ(u)≤4 +e <7
= 0, where θ∗ 2 := nσp k 1−p ! 1 p−2 .(40) 26 Denoteψ(u) := 4u 1/(1+u) + 4−uu−u/(1+u), foru∈(0,1). Usingulog 1 u ≤1 +u, available for all u∈(0,1)we obtainψ(u)≤4 +e <7. We can then compute h′ k(θ∗ 2)≥h ′ k(4θ∗
-
[2]
SinceM >7 2−p in (35) and thanks to our choice ofd >0, by (36) we havenX 2 k >7 2−p(Xk/σk)p and thush ′ k(θ∗ 2)≥h ′ k(4θ∗ 2)>0
=n Xk −4 nσp k 1−p ! 1 p−2 − 4p−1 nσp k nσp k 1−p ! p−1 p−2 =n Xk −(nσ p k) 1 p−2 ψ(1−p) ≥n(X k −7(nσ p k) 1 p−2 ). SinceM >7 2−p in (35) and thanks to our choice ofd >0, by (36) we havenX 2 k >7 2−p(Xk/σk)p and thush ′ k(θ∗ 2)≥h ′ k(4θ∗ 2)>0. Thus,h ′ k starts from−∞at0, ends at−∞at+∞and takes a positive value atθ ∗ 2, thereforeh ′ k vanishes at e...
-
[3]
First notice thatlim θ→∞ hk(θ) =−∞
= 0). First notice thatlim θ→∞ hk(θ) =−∞. Second, recalling (40) the definition ofθ ∗ 2, we have hk(2θ∗ 2)> h k(0), this shows that the global maximum ofhk is attained atθ ∗ M and not at the boundary. Indeed hk(2θ∗ 2)−h k(0) = 2nXkθ∗ 2 −2n(θ ∗ 2)2 − 2p pσp k (θ∗ 2)p = 2nθ∗ 2 Xk −(1 + 2p 2p(1−p) )θ∗ 2 >0, where the last inequality comes from choice ofMlarg...
-
[4]
=−n(1−2 p−2)≤ −n/2and thus I2 ≤e hk(θ∗ M) Z ∞ 2θ∗ 2 (θ−θ ∗ M)2e− n 4 (θ−θ∗ M)2 dθ≤e hk(θ∗ M) 4√π n√n . Now forI 1, we knowh k decreases from0toθ ∗m and then increases up untilθ ∗ M ≥2θ ∗ 2, thus I1 ≤ ehk(0) ∨e hk(2θ∗ 2) Z 2θ∗ 2 0 (θ−θ ∗ M)2 dθ≤4θ ∗ 2[(θ∗ 2)2 + (θ∗ M)2](ehk(0) ∨e hk(2θ∗ 2)). Considering the maximum on the right hand side, we showed in the ...
-
[5]
forβ∈(1,2], we have1 +p−βp≥(2−β)p≥0, hence the second term dominates the first and overall in the right hand side of the bound
-
[6]
forβ∈(0,1], we have1 +p−βp > p(1−β) +, hence again the second term dominates the first and overall in the right hand side of the bound. For anyβ∈(0,2]we thus get that I≥exp(−N β(c2 + log(Nασn εn ))−c 3 N1+(1−β)p β σp n ) and combining with the bounds for the previous terms, we obtain that under the assumption εn/(σnNα)≳log 1/q n,(52) Π(||f−f 0||∞ ≤ε n)≥ex...
-
[7]
it holds σ−pn n pn N1+(1−β)pn β ≲nε 2 n or equivalently εn ≳ε ∗ n n( 1−β 1+2β +t)pn/2 √pn . The rateε n ≥ε ∗n √lognin the statement trivially satisfies the first condition, while for the second one, under our assumptionsn ( 1−β 1+2β +t)pn/2 is bounded, so it is again satisfied. 40 Remark 3.It is easy to verify that the proof of Theorem 5 goes through as w...
2024
-
[8]
, n γ .(61) Then, for a small enough constantd >0above, one hasP f0[An] = 1 +o(1)asn→ ∞
LetA n be the event defined by, forµ k as in(59)andn γ =dN γ, An = µk ≥f 0,k/4,for allk= 1, . . . , n γ .(61) Then, for a small enough constantd >0above, one hasP f0[An] = 1 +o(1)asn→ ∞
Show all 15 references
-
[9]
Proof.One first notes that for small enoughd, for anyk≤n γ =dN γ, one hasf 0,k/2≥(nσ k)−1, by definition off 0,k
There exist constantsc 1, c2 >0such that, forw + k as in(60), on the eventA n as in(61), max 1≤k≤nγ (1−w + k )≤c 1e−c2n(α−β)/(α+β+1) . Proof.One first notes that for small enoughd, for anyk≤n γ =dN γ, one hasf 0,k/2≥(nσ k)−1, by definition off 0,k. Since underP f0 we haveµ k =...
-
[10]
It now suffices to show that the expectation underP (n) f0 of the last display goes to0. Starting first with the indicator, and denoting bye(·)the function with coefficientse k =ε k for≤K n and0otherwise, applying the triangle inequality gives P (n) f0 [∥f0 −h∥ Kn > M εn/2]≤P[...
-
[11]
forβ∈(0,1], we have1 +p−βp > p(1−β) +, hence again the second term dominates the first in (51) and overall in the right hand side of the bound
-
[12]
forβ∈(1,2], the second term dominates the first forp≤(β−1) −1, otherwise the first term dominates in the right hand side of the bound. Forβ∈(0,1 + 1 p], the remaining of the proof is still identical to the proof of Theorem 4, and we only need to deal with the caseβ∈(1 + 1 p ,2...
2023
-
[13]
either by Equation(7), in which case for anyρ∈(0,1), {f∈ F:||f−f 0||2 ≤ε} ⊂ B n(f0, ε)and 1 n Dρ(P (n) f , P(n) f0 ) = ρ 2(1−ρ) ||f−f 0||2 2
-
[14]
Further assuming||f|| ∞ ∨ ||f0||∞ ≤F, we have, for anyρ∈(0,1), 1 n Dρ(P (n) f , P(n) f0 )≥ ρ 2 e−2F 2ρ(1−ρ)||f−f 0||2 2,PX
or by Equation(20), and then there exists a constantC >0such that {f∈ F:||f−f 0||∞ ≤ε} ⊂ B n(f0, Cε). Further assuming||f|| ∞ ∨ ||f0||∞ ≤F, we have, for anyρ∈(0,1), 1 n Dρ(P (n) f , P(n) f0 )≥ ρ 2 e−2F 2ρ(1−ρ)||f−f 0||2 2,PX . D.2. Lemmas for series priors In this Section we r...
-
[15]
Proof.A union bound directly shows thatP (n) f0 [Bcn] =o(1)
for any constantM=M(p), one can choosed >0small enough inn γ :=dN γ, such that nX2 k ≥M Xk σk p . Proof.A union bound directly shows thatP (n) f0 [Bcn] =o(1). Also, by definition off 0,k andσ k the second pointLσ k ≤2X k immediately follows from the first. Let us check the lat...
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.