Pith. sign in

REVIEW 4 minor 21 references

Stochastic Approximation in Banach Spaces Without Geometric Constraints

T0 review · 0 major / 4 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Stochastic approximation converges almost surely on every Banach space once the noise is mean-zero and i.i.d. (or independent and tight), with no geometric restrictions on the space.

desk verdict Clean extension of SA to arbitrary Banach spaces by swapping geometry for tightness; the deterministic scaffolding is the real contribution. read the letter →

arxiv 2607.10356 v1 pith:S64VIJLE submitted 2026-07-11 math.PR

classification math.PR MSC 60F1562L2046B09
keywords stochasticapproximationBanachspacesalmost-sureconvergencetightnoisei.i.d.step-sizeconditionsstronglawoflargenumbers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves that the classical stochastic-approximation recursion finds the unique root of a map G on an arbitrary Banach space, without any assumption on the geometry of that space. Earlier infinite-dimensional results required the space to have special properties such as the Radon–Nikodym property or uniform smoothness; here those restrictions are removed. When the observation noise is i.i.d. with mean zero and a finite moment of order α between 1 and 2, ordinary step-size conditions already guarantee almost-sure convergence. For merely independent mean-zero noise the same conclusion holds once the noise sequence is tight and has a uniform moment. The argument works by reducing the recursion to a deterministic comparison that absorbs the noise into a weighted average, then showing that this average vanishes almost surely under the stated moment and tightness hypotheses. Because many applied spaces (for example continuous functions on an interval) fail the classical geometric hypotheses, the result substantially enlarges the class of problems to which stochastic approximation can be applied with certainty.

What carries the argument

The deterministic class L({γ_n}) of Banach-valued sequences whose weighted averages can be split into a convergent part and a uniformly small residual; once the noise is shown to lie in this class almost surely, a pure contraction argument yields convergence of the recursion.

What would settle it

Construct a Banach space, a map G satisfying the contraction condition, and an independent mean-zero tight noise sequence with finite second moment such that the recursion with square-summable steps fails to converge almost surely; any such counter-example would refute the claim.

Watch

Extended reading notes

Core claim

On every Banach space the stochastic-approximation sequence X_{n+1}=X_n-β_n(G(X_n)+noise) converges almost surely to the unique root of G whenever the noise is independent, mean-zero and either i.i.d. with a finite α-moment or tight with a uniform α-moment, under the usual step-size conditions and a structural contraction assumption on G.

Load-bearing premise

The map G must satisfy a uniform contraction condition: after a fixed multiple of G is subtracted, every point moves strictly closer to the root by a factor less than one.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper proves almost-sure convergence of the stochastic approximation recursion X_{n+1}=X_n-eta_n(G(X_n)+ heta_n W_{n+1}) to the unique root x* of G, for an arbitrary Banach space B, without any geometric assumptions (type, cotype, Radon–Nikodym, etc.). The standing hypotheses are a uniform contraction condition (2) on G, the usual step-size requirements (3), a linear growth bound (4) on the multiplicative noise factor, and one of three noise regimes (J1)–(J3): i.i.d. mean-zero with finite second moment and square-summable steps; i.i.d. mean-zero with finite heta-moment ( heta∈[1,2)) and steps O(n^{-1/ heta}); or independent mean-zero tight noise with uniform heta-moment ( heta∈(1,2]) and summable heta-powers of the steps. The argument proceeds by introducing a deterministic class L({ heta_•}) of sequences that admit heta-weighted averages converging after an heta-small residual, proving a comparison theorem (Theorem 5) that reduces the recursion to membership of the noise in this class, and then verifying that membership via a tightness decomposition (Theorem 6) plus classical martingale convergence under each of (J1)–(J3).

Significance. The result removes the geometric hypotheses that have been standard in infinite-dimensional stochastic approximation since the 1980s, while recovering the classical finite-dimensional theorems as special cases. The only extra probabilistic price for non-i.i.d. noise is tightness, which is natural and checkable. The deterministic comparison lemmas (Lemmas 2–4, Theorem 5) and the finite-dimensional approximation under tightness (Theorem 6) are of independent interest and cleanly separate the analytic and probabilistic ingredients. The proofs are self-contained and elementary (weighted averages + martingale convergence), so the paper is likely to be usable by researchers working in C[0,1], L^1, or other spaces that fail the usual geometric conditions.

minor comments (4)
  1. Page 3, line after (5): “x_0∈R^d” is a leftover from the finite-dimensional setting; it should read x_0∈B.
  2. In the definition of L({ heta_•}) the sequences are indexed from n≥1 while the weighted sums begin at k=0; a uniform convention (or an explicit empty-product convention) would avoid occasional index shifts later.
  3. Theorem 6 constructs the approximating random variables Y_n and heta_{n,j} but does not explicitly record that they remain independent of the past filtration F_n; a one-line remark would make the subsequent martingale arguments completely transparent.
  4. A few typographical slips: “Mutatis Mutandis”, “thestep size”, “istight”, and the arXiv date “11 Jul 2026” should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: self-contained almost-sure argument from contraction (2), deterministic class L, tightness decomposition and martingale convergence.

full rationale

The paper proves Theorem 1 by an explicit reduction: the recursive scheme (5) is rewritten as a weighted average (40) whose noise term belongs almost surely to the deterministic class L({β•}) under any of (J1)–(J3). Membership in L is obtained from the tightness decomposition of Theorem 6 (which produces a finite-rank mean-zero part plus a small remainder) together with classical L^{2}-bounded martingale convergence for the scalar coefficients; the deterministic comparison Theorem 5 then extracts ||X_n - x*|| o 0 from the structural contraction (2). All steps are proved in full inside the paper; the self-citations [13,16] are used only for historical comparison and are not load-bearing inputs. There are no fitted parameters, no uniqueness theorems imported from the authors, and no renaming of known results. The derivation is therefore independent of its own conclusions.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The paper is a pure existence/convergence theorem. It rests on standard Banach-space and martingale facts plus one structural hypothesis on G and the usual step-size and moment conditions. No free parameters are fitted; the only invented auxiliary object is the linear space L({γ_•}) of sequences that can be approximated by convergent weighted averages, introduced solely as a proof device.

assumptions (5)
  • standard math Standard real-valued martingale convergence: an L^2-bounded martingale converges almost surely.
    Invoked repeatedly in Steps 2–4 of the proof of Theorem 1 to obtain almost-sure convergence of the scalar series involving the noise components.
  • standard math Hoffmann-Jørgensen SLLN: i.i.d. Banach-valued random variables with finite first moment and mean zero satisfy the strong law on every Banach space.
    Cited in the introduction as the classical special case that already holds without geometry; used only for motivation, not as a black-box lemma inside the proof.
  • domain assumption Condition (2): ∃τ>0, ρ<1 such that ||x-x*-τG(x)||≤ρ||x-x*|| for all x.
    Standing structural hypothesis on the map G that guarantees uniqueness of the root and supplies the contraction needed for the deterministic comparison argument (Theorem 5).
  • domain assumption Step-size conditions (3): β_n→0 and ∑β_n=∞, together with the moment/summability requirements (8),(10) or (14).
    Classical Robbins–Monro-type hypotheses; without them the weighted averages need not converge.
  • domain assumption Tightness of the noise sequence when the variables are not identically distributed (condition (12)).
    The key probabilistic substitute for geometric assumptions on the Banach space; used in Theorem 6 to produce a finite-dimensional approximation of the noise.
invented entities (1)
  • The linear space L({γ_•}) of Banach-valued sequences that admit an ε-approximation by a convergent weighted average plus a small residual.
    purpose: Technical device that isolates the noise sequences for which the deterministic recursion of Theorem 5 converges; the random noise is then shown to lie in this space almost surely.
    Defined ad hoc in Section 3 solely for the proof; no independent existence claim is made outside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Stochastic Approximation in Banach Spaces Without Geometric Constraints." pith.science (2026). https://pith.science/paper/S64VIJLE

@misc{pith2026260710356,
  author       = {Pith},
  title        = {Pith review of: Stochastic Approximation in Banach Spaces Without Geometric Constraints},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S64VIJLE}},
  note         = {Machine review of arXiv:2607.10356}
}
read the original abstract

The thrust of this article is to show that on all Banach spaces, stochastic approximation holds when the noise sequence is an i.i.d. sequence with mean 0, without imposing any condition on the geometry of the space. Also, the same is true when the noise is a sequence of independent random variables under appropriate conditions on the moment. In this case, we need to require that the noise sequence is tight.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

21 extracted references · 2 linked inside Pith

  1. [1]

    and Priouret, P.Adaptive Algorithms and Stochastic Approxima- tion

    Benveniste, A, Metivier, M. and Priouret, P.Adaptive Algorithms and Stochastic Approxima- tion. Springer-Verlag, 1990

  2. [2]

    Bertsekas, D. P. Reinforcement Learning and Optimal Control.Athena Scientific, 2019

  3. [3]

    Blum, J. R. Approximation methods which converge with probability one.Ann. Math. Statist., 25: 382–386, 1954

  4. [4]

    Blum, J. R. Multidimensional stochastic approximation .Annals of Mathematical Statistics, 25, 737–744, 1954

  5. [5]

    Borkar, V. S. Asynchronous stochastic approximations.SIAM Journal on Control and Opti- mization, 36(3), 840–851, 1998

  6. [6]

    Borkar, V. S. Stochastic Approximation: A Dynamical Systems Viewpoint.Hindustan Book Agency, New Delhi, India and Cambridge University Press, Cambridge, UK, 2008. 18

  7. [7]

    Borkar, V. S. and Meyn, S.P. The O.D.E. method for convergence of stochastic approximation and reinforcement learning.SIAM Journal on Control and Optimization, 38, 447–469, 2000

  8. [8]

    On stochastic approximation.Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, 39–55, 1956

    Dvoretzky A. On stochastic approximation.Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, 39–55, 1956

Show all 21 references
  1. [9]

    Gladyshev, E. G. On stochastic approximation.Theory Probab Appl10: 275–278, 1965

  2. [10]

    Probability in Banach SpaceEcole d’Ete de Probabilites de Saint-Flour VI, 1976Lecture Notes in Mathematics, 598, 1977

    Hoffmann-Jørgensen, J. Probability in Banach SpaceEcole d’Ete de Probabilites de Saint-Flour VI, 1976Lecture Notes in Mathematics, 598, 1977

  3. [11]

    Jaakkola, T, Jordan, M. I. and Singh, S. P. Convergence of stochastic iterative dynamic programming algorithms.Neural Computation, 6, 1185–1201, 1994

  4. [12]

    and Wolfowitz, J

    Kiefer, J. and Wolfowitz, J. Stochastic estimation of the maximum of a regression function. Ann. Math. Statist., 23 462–466, 1952

  5. [13]

    and Rao, B

    Karandikar, R.L. and Rao, B. V. Stochastic approximation in infinite dimensionsInfinite Dimensional Analysis, Quantum Probability and Related Topics27, 2024

  6. [14]

    and Vidyasagar, M

    Karandikar, R.L. and Vidyasagar, M. Convergence of batch asynchronous stochastic approximation with applications to reinforcement learning. (2021) https://arxiv.org/pdf/2109.03445.pdf

  7. [15]

    Karandikar, R.L. and Vidyasagar, M.: Convergence Rates for Stochastic Approximation: Biased Noise with Unbounded Variance, and Applications.Journal of Optimization Theory and Applications, 203, 2412—2450, 2024.https://arxiv.org/pdf/2312.02828v2.pdf

  8. [16]

    Karandikar, R.L., Rao, B. V. and Vidyasagar, M.: Revisiting Stochastic Approximation and Stochastic Gradient Descent To appear in Pure and Applied Functional Analysis: Memory of Professor Allen Tannenbaum, 2026.https://arxiv.org/abs/2505.11343

  9. [17]

    Kushner, H. J. and Shwartz, A. Stochastic Approximation in Hilbert Space: Identification and Optimization of Linear Continuous Parameter Systems.SIAM Journal on Control and Optimization, 23(5), 774–793, 1985

  10. [18]

    Lai, T. L. Stochastic approximation (invited paper).The Annals of Statistics, 31, 391– 406, 2003. 19

  11. [19]

    and Monro, S

    Robbins, H. and Monro, S. A stochastic approximation method.Annals of Mathematical Statistics, 22, 400–407, 1951

  12. [20]

    and Siegmund, D

    Robbins, H. and Siegmund, D. A convergence theorem for nonnegative almost supermartin- gales and some applications. InOptimizing Methods in Statistics(J. S. Rustagi, ed.) 233–257, Academic Press, New York,1971

  13. [21]

    Stochastic approximation.Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability.587–609, 1960

    Schmetterer, L. Stochastic approximation.Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability.587–609, 1960. 20

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.