REVIEW 4 minor 21 references
Stochastic Approximation in Banach Spaces Without Geometric Constraints
T0 review · 0 major / 4 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Stochastic approximation converges almost surely on every Banach space once the noise is mean-zero and i.i.d. (or independent and tight), with no geometric restrictions on the space.
desk verdict Clean extension of SA to arbitrary Banach spaces by swapping geometry for tightness; the deterministic scaffolding is the real contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The deterministic class L({γ_n}) of Banach-valued sequences whose weighted averages can be split into a convergent part and a uniformly small residual; once the noise is shown to lie in this class almost surely, a pure contraction argument yields convergence of the recursion.
What would settle it
Construct a Banach space, a map G satisfying the contraction condition, and an independent mean-zero tight noise sequence with finite second moment such that the recursion with square-summable steps fails to converge almost surely; any such counter-example would refute the claim.
Extended reading notes
Core claim
On every Banach space the stochastic-approximation sequence X_{n+1}=X_n-β_n(G(X_n)+noise) converges almost surely to the unique root of G whenever the noise is independent, mean-zero and either i.i.d. with a finite α-moment or tight with a uniform α-moment, under the usual step-size conditions and a structural contraction assumption on G.
Load-bearing premise
The map G must satisfy a uniform contraction condition: after a fixed multiple of G is subtracted, every point moves strictly closer to the root by a factor less than one.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proves almost-sure convergence of the stochastic approximation recursion X_{n+1}=X_n-eta_n(G(X_n)+ heta_n W_{n+1}) to the unique root x* of G, for an arbitrary Banach space B, without any geometric assumptions (type, cotype, Radon–Nikodym, etc.). The standing hypotheses are a uniform contraction condition (2) on G, the usual step-size requirements (3), a linear growth bound (4) on the multiplicative noise factor, and one of three noise regimes (J1)–(J3): i.i.d. mean-zero with finite second moment and square-summable steps; i.i.d. mean-zero with finite heta-moment ( heta∈[1,2)) and steps O(n^{-1/ heta}); or independent mean-zero tight noise with uniform heta-moment ( heta∈(1,2]) and summable heta-powers of the steps. The argument proceeds by introducing a deterministic class L({ heta_•}) of sequences that admit heta-weighted averages converging after an heta-small residual, proving a comparison theorem (Theorem 5) that reduces the recursion to membership of the noise in this class, and then verifying that membership via a tightness decomposition (Theorem 6) plus classical martingale convergence under each of (J1)–(J3).
Significance. The result removes the geometric hypotheses that have been standard in infinite-dimensional stochastic approximation since the 1980s, while recovering the classical finite-dimensional theorems as special cases. The only extra probabilistic price for non-i.i.d. noise is tightness, which is natural and checkable. The deterministic comparison lemmas (Lemmas 2–4, Theorem 5) and the finite-dimensional approximation under tightness (Theorem 6) are of independent interest and cleanly separate the analytic and probabilistic ingredients. The proofs are self-contained and elementary (weighted averages + martingale convergence), so the paper is likely to be usable by researchers working in C[0,1], L^1, or other spaces that fail the usual geometric conditions.
minor comments (4)
- Page 3, line after (5): “x_0∈R^d” is a leftover from the finite-dimensional setting; it should read x_0∈B.
- In the definition of L({ heta_•}) the sequences are indexed from n≥1 while the weighted sums begin at k=0; a uniform convention (or an explicit empty-product convention) would avoid occasional index shifts later.
- Theorem 6 constructs the approximating random variables Y_n and heta_{n,j} but does not explicitly record that they remain independent of the past filtration F_n; a one-line remark would make the subsequent martingale arguments completely transparent.
- A few typographical slips: “Mutatis Mutandis”, “thestep size”, “istight”, and the arXiv date “11 Jul 2026” should be corrected.
Circularity Check
No circularity: self-contained almost-sure argument from contraction (2), deterministic class L, tightness decomposition and martingale convergence.
full rationale
The paper proves Theorem 1 by an explicit reduction: the recursive scheme (5) is rewritten as a weighted average (40) whose noise term belongs almost surely to the deterministic class L({β•}) under any of (J1)–(J3). Membership in L is obtained from the tightness decomposition of Theorem 6 (which produces a finite-rank mean-zero part plus a small remainder) together with classical L^{2}-bounded martingale convergence for the scalar coefficients; the deterministic comparison Theorem 5 then extracts ||X_n - x*|| o 0 from the structural contraction (2). All steps are proved in full inside the paper; the self-citations [13,16] are used only for historical comparison and are not load-bearing inputs. There are no fitted parameters, no uniqueness theorems imported from the authors, and no renaming of known results. The derivation is therefore independent of its own conclusions.
Assumptions & free parameters
assumptions (5)
- standard math Standard real-valued martingale convergence: an L^2-bounded martingale converges almost surely.
- standard math Hoffmann-Jørgensen SLLN: i.i.d. Banach-valued random variables with finite first moment and mean zero satisfy the strong law on every Banach space.
- domain assumption Condition (2): ∃τ>0, ρ<1 such that ||x-x*-τG(x)||≤ρ||x-x*|| for all x.
- domain assumption Step-size conditions (3): β_n→0 and ∑β_n=∞, together with the moment/summability requirements (8),(10) or (14).
- domain assumption Tightness of the noise sequence when the variables are not identically distributed (condition (12)).
invented entities (1)
-
The linear space L({γ_•}) of Banach-valued sequences that admit an ε-approximation by a convergent weighted average plus a small residual.
Cite this review
Pith. "Pith review of Stochastic Approximation in Banach Spaces Without Geometric Constraints." pith.science (2026). https://pith.science/paper/S64VIJLE
@misc{pith2026260710356,
author = {Pith},
title = {Pith review of: Stochastic Approximation in Banach Spaces Without Geometric Constraints},
year = {2026},
howpublished = {\url{https://pith.science/paper/S64VIJLE}},
note = {Machine review of arXiv:2607.10356}
}
read the original abstract
The thrust of this article is to show that on all Banach spaces, stochastic approximation holds when the noise sequence is an i.i.d. sequence with mean 0, without imposing any condition on the geometry of the space. Also, the same is true when the noise is a sequence of independent random variables under appropriate conditions on the moment. In this case, we need to require that the noise sequence is tight.
Reference graph
Works this paper leans on
-
[1]
and Priouret, P.Adaptive Algorithms and Stochastic Approxima- tion
Benveniste, A, Metivier, M. and Priouret, P.Adaptive Algorithms and Stochastic Approxima- tion. Springer-Verlag, 1990
1990
-
[2]
Bertsekas, D. P. Reinforcement Learning and Optimal Control.Athena Scientific, 2019
2019
-
[3]
Blum, J. R. Approximation methods which converge with probability one.Ann. Math. Statist., 25: 382–386, 1954
1954
-
[4]
Blum, J. R. Multidimensional stochastic approximation .Annals of Mathematical Statistics, 25, 737–744, 1954
1954
-
[5]
Borkar, V. S. Asynchronous stochastic approximations.SIAM Journal on Control and Opti- mization, 36(3), 840–851, 1998
1998
-
[6]
Borkar, V. S. Stochastic Approximation: A Dynamical Systems Viewpoint.Hindustan Book Agency, New Delhi, India and Cambridge University Press, Cambridge, UK, 2008. 18
2008
-
[7]
Borkar, V. S. and Meyn, S.P. The O.D.E. method for convergence of stochastic approximation and reinforcement learning.SIAM Journal on Control and Optimization, 38, 447–469, 2000
2000
-
[8]
On stochastic approximation.Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, 39–55, 1956
Dvoretzky A. On stochastic approximation.Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, 39–55, 1956
1956
Show all 21 references
-
[9]
Gladyshev, E. G. On stochastic approximation.Theory Probab Appl10: 275–278, 1965
1965
-
[10]
Probability in Banach SpaceEcole d’Ete de Probabilites de Saint-Flour VI, 1976Lecture Notes in Mathematics, 598, 1977
Hoffmann-Jørgensen, J. Probability in Banach SpaceEcole d’Ete de Probabilites de Saint-Flour VI, 1976Lecture Notes in Mathematics, 598, 1977
1977
-
[11]
Jaakkola, T, Jordan, M. I. and Singh, S. P. Convergence of stochastic iterative dynamic programming algorithms.Neural Computation, 6, 1185–1201, 1994
1994
-
[12]
and Wolfowitz, J
Kiefer, J. and Wolfowitz, J. Stochastic estimation of the maximum of a regression function. Ann. Math. Statist., 23 462–466, 1952
1952
-
[13]
and Rao, B
Karandikar, R.L. and Rao, B. V. Stochastic approximation in infinite dimensionsInfinite Dimensional Analysis, Quantum Probability and Related Topics27, 2024
2024
-
[14]
and Vidyasagar, M
Karandikar, R.L. and Vidyasagar, M. Convergence of batch asynchronous stochastic approximation with applications to reinforcement learning. (2021) https://arxiv.org/pdf/2109.03445.pdf
2021 arXiv
-
[15]
Karandikar, R.L. and Vidyasagar, M.: Convergence Rates for Stochastic Approximation: Biased Noise with Unbounded Variance, and Applications.Journal of Optimization Theory and Applications, 203, 2412—2450, 2024.https://arxiv.org/pdf/2312.02828v2.pdf
2024 arXiv
-
[16]
Karandikar, R.L., Rao, B. V. and Vidyasagar, M.: Revisiting Stochastic Approximation and Stochastic Gradient Descent To appear in Pure and Applied Functional Analysis: Memory of Professor Allen Tannenbaum, 2026.https://arxiv.org/abs/2505.11343
2026
-
[17]
Kushner, H. J. and Shwartz, A. Stochastic Approximation in Hilbert Space: Identification and Optimization of Linear Continuous Parameter Systems.SIAM Journal on Control and Optimization, 23(5), 774–793, 1985
1985
-
[18]
Lai, T. L. Stochastic approximation (invited paper).The Annals of Statistics, 31, 391– 406, 2003. 19
2003
-
[19]
and Monro, S
Robbins, H. and Monro, S. A stochastic approximation method.Annals of Mathematical Statistics, 22, 400–407, 1951
1951
-
[20]
and Siegmund, D
Robbins, H. and Siegmund, D. A convergence theorem for nonnegative almost supermartin- gales and some applications. InOptimizing Methods in Statistics(J. S. Rustagi, ed.) 233–257, Academic Press, New York,1971
1971
-
[21]
Stochastic approximation.Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability.587–609, 1960
Schmetterer, L. Stochastic approximation.Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability.587–609, 1960. 20
1960
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.