Pith. sign in

REVIEW 4 major objections 4 minor 15 references

Stochastic Krasnosel skii-Mann Iterations in Banach Spaces with Bregman Distances

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A Bregman-distance generalization of the stochastic Krasnosel'skii-Mann iteration converges almost surely to a fixed point in reflexive Banach spaces, with residual bounds governed by the uniform convexity modulus.

desk verdict The core Bregman-SKM update is undefined in general Banach spaces and the rate proof inverts a key inequality; the intended extension is natural but the technical execution doesn't hold up. read the letter →

arxiv 2506.08031 v1 pith:FVM7QQ5K submitted 2025-06-02 math.OC

classification math.OC MSC 47H0547J2549M2765K1090C25
keywords stochasticfixed-pointiterationBregmandistanceBanachspaceKrasnosel'skii-Mannalmost-sureconvergenceuniformconvexitymodulusmartingale-differencenoisemirrordescent
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes a stochastic version of the Krasnosel'skii-Mann iteration that works in reflexive Banach spaces instead of Hilbert spaces, using Bregman distances to adapt the geometry. It claims that, under martingale-difference noise and mild conditions on a Legendre distance-generating function, the iterates converge almost surely to a fixed point of a nonexpansive operator, and the Bregman residual goes to zero. It further derives non-asymptotic bounds on the averaged residual that depend on the uniform convexity modulus of the generating function. A sympathetic reader would care because many optimization and reinforcement-learning algorithms naturally live in non-Euclidean spaces, and this result would give them the same stochastic convergence guarantees that are already available in Hilbert spaces.

What carries the argument

The central object is the Bregman-SKM update, a two-step map that pulls the current point into the dual space through the gradient of a Legendre function, averages it with the nonexpansive operator's output, and returns via the conjugate gradient. The argument is carried by the three-point identity for Bregman distances and the uniform convexity modulus \(\delta\) of \(\vartheta\): together they give a one-step residual decrease (Lemma 3.2) with a shrinking term proportional to \(\alpha_n \delta(\|\zeta_n-\hbar(\zeta_n)\|)\). Summability of step-squares and a standard almost-supermartingale convergence lemma then force the residual to zero, while the rate exponent \(p\) in Theorem 4.3 appears from the polynomial lower bound on \(\delta\).

What would settle it

Look at Definition 3.1 in a concrete reflexive Banach space where \(X\neq X^*\) as sets, such as \(\ell^p\) for \(p\in(1,2)\): the expression \(\hbar(\zeta_n)+\mho_n\) sums an element of \(X\) with an element of \(X^*\), which is undefined, so the central algorithm has no meaning without an additional identification that the paper does not state.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that the Bregman-SKM iteration, defined by \(\varsigma_n = \nabla\vartheta^*\left((1-\alpha_n)\nabla\vartheta(\zeta_n)+\alpha_n\nabla\vartheta(\hbar(\zeta_n)+\mho_n)\right)\) and \(\zeta_{n+1}=\varsigma_n\), is almost-surely convergent: under assumptions (A1)-(A4), \(\zeta_n \to \zeta^*\) for some \(\zeta^* \in \mathrm{Fix}(\hbar)\) and \(D_\vartheta(\zeta_n,\hbar(\zeta_n)) \to 0\) almost surely. When the modulus of uniform convexity satisfies \(\delta(r) \ge c r^q\), the window-averaged residual satisfies \(\bar{R}_N = O($A_N^{{-p}}$)\) with \(p=(q-1)/q\), recovering the classical \(O(1/\sqrt{n})\) Hilbert-space rate for \(q=2\).

Load-bearing premise

The core premise is that the algorithmic update makes sense, but the update adds a dual-space noise term directly to a primal-space operator output, so the algorithm is well defined only when the space and its dual are identified—true in Hilbert spaces but not in the general reflexive Banach spaces the paper claims to cover.

Editorial extensions

If this is right

  • Almost-sure convergence now holds for stochastic fixed-point iterations in reflexive Banach spaces, so entropy-regularized reinforcement learning and mirror-descent variants can be analyzed in their native geometry.
  • When \(\vartheta(\zeta)=\tfrac12\|\zeta\|^2\) in a Hilbert space, the new bounds reduce to the known \(O(1/\sqrt{n})\) averaged residual, giving a unified framework rather than a separate theory.
  • With polynomial step-sizes \(\alpha_n=n^{-\gamma}\) for \(\gamma\in(1/2,1)\), the averaged residual scales as \(O(N^{-p(1-\gamma)})\), so the rate improves as the uniform-convexity exponent \(q\) grows.
  • Adaptive Bregman geometries (time-varying \(\vartheta_n\)) and trimmed heavy-tailed noise both preserve almost-sure convergence under the stated conditions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The update rule in Definition 3.1 adds a dual-space noise vector \(\mho_n \in X^*\) to the primal-space vector \(\hbar(\zeta_n) \in X\) inside the same argument; in a general reflexive Banach space this sum is not defined unless the paper silently identifies \(X\) with \(X^*\), which holds for Hilbert spaces but not for, say, \(\ell^p\) with \(p\neq 2\).
  • If the well-posedness gap is repaired, the assumption that the noise lives in the dual space suggests a natural interpretation: the stochastic perturbation affects the gradient of the Bregman function, not the operator evaluation itself, so a cleaner formulation might apply noise after \(\nabla\vartheta\) rather than inside it.
  • The rate exponent \(p=(q-1)/q\) implies that strengthening uniform convexity (larger \(q\)) drives the exponent to 1, so one could design distance-generating functions with high-order convexity to approach linear convergence; the paper leaves such a construction open.
  • The trimming results under heavy-tailed noise depend on an order-statistic bound for the removed coordinates; a direct numerical test in \(\ell^p\) with Student-t noise would reveal whether the logarithmic trimming schedule behaves as predicted when the spaces are not identified.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a stochastic Krasnosel'skii-Mann (SKM) iteration in reflexive Banach spaces using Bregman distances. The authors define a Bregman-SKM update, prove almost-sure convergence to a fixed point (Theorem 3.3), and derive non-asymptotic residual bounds (Theorem 4.3) under a uniform-convexity modulus condition. They also discuss adaptive Bregman geometries and heavy-tailed noise with trimming, and report numerical experiments on entropy-regularized policy iteration.

Significance. If correct, the results would generalize stochastic KM methods beyond Hilbert spaces and provide rates governed by the modulus of uniform convexity. However, the central algorithm is not well-defined as stated because it adds a dual-space noise element to a primal-space point, and the main proofs contain load-bearing gaps. The paper does provide a clear structure and the Hilbert-space special case reduces to known SKM, but the claimed Banach-space extension is not established.

major comments (4)
  1. [Definition 3.1] The Bregman-SKM update is not well-defined in general reflexive Banach spaces. The update reads ζ_{n+1} = ∇ϑ*((1−α_n)∇ϑ(ζ_n) + α_n∇ϑ(ℏ(ζ_n)+℧_n)), with ℏ(ζ_n) ∈ X and (by Definition 2.6 and (A4)) ℧_n ∈ X*. The sum ℏ(ζ_n)+℧_n therefore adds a primal vector and a dual vector, which is meaningless unless X and X* are identified. No such identification is assumed in (A1)–(A4), and reflexive Banach spaces generally do not admit a canonical isometric identification of X with X*. This defect propagates to Algorithms 1 and 2, Theorem 5.2, and Proposition 5.4, all of which use the same primal-plus-dual addition. A repair would require redefining the algorithm, e.g., placing the noise in X or adding it after applying ∇ϑ, and then re-deriving all subsequent results.
  2. [Lemma 3.2] The proof asserts without support that uniform convexity of ϑ implies Lipschitz continuity of ∇ϑ and ∇ϑ* on bounded sets. Uniform convexity alone does not imply differentiability beyond Gateaux differentiability, nor does it yield a Lipschitz gradient; standard results require additional smoothness assumptions such as uniform smoothness or a modulus of smoothness. Because the one-step decrease estimate (Lemma 3.2) is the foundation for Theorem 3.3 and Theorem 4.3, this missing hypothesis undermines the entire analysis. The paper needs to either add explicit smoothness assumptions or replace these steps with arguments that do not rely on unproved Lipschitz bounds.
  3. [Theorem 4.3] The proof contains a key inequality in the wrong direction. From D_n = Dϑ(ζ_n, ℏ(ζ_n)) ≥ δ(∥ζ_n−ℏ(ζ_n)∥) and δ(r) ≥ c r^q, one obtains D_n ≥ c∥ζ_n−ℏ(ζ_n)∥^q, i.e., c∥ζ_n−ℏ(ζ_n)∥^q ≤ D_n. The proof, however, substitutes δ(∥ζ_n−ℏ(ζ_n)∥) ≥ c∥ζ_n−ℏ(ζ_n)∥^q ≥ c(D_n/c) = D_n, which reverses the inequality. Consequently the drift term (1/2)D_n α_n in the displayed inequality is not justified. This invalidates the derivation of the averaged residual bound O(A_N^{-p}).
  4. [Theorem 3.3] The proof concludes that δ(∥ζ_n−ℏ(ζ_n)∥) → 0 a.s. from ∑ α_n δ(∥ζ_n−ℏ(ζ_n)∥) < ∞ a.s. and ∑ α_n = ∞. This implication is false in general: with α_n = 1/n, taking δ(x_n)=1 on a sparse subsequence and 0 elsewhere yields a finite sum ∑ α_n δ(x_n) while δ(x_n) does not tend to 0. Without an additional argument forcing δ(x_n) → 0, the conclusion that ∥ζ_n−ℏ(ζ_n)∥ → 0 and hence D_n → 0 does not follow from the Robbins–Siegmund lemma as applied. This is a load-bearing gap in the almost-sure convergence claim.
minor comments (4)
  1. [Throughout] The text contains numerous typographical and encoding issues, such as 'Krasnosel ski ¨A', 'Fej ˜A©r', 'Fix(⟨⌊⊣∇)', and inconsistent spacing in the title. These should be corrected in a revision.
  2. [Definition 2.6 and Definition 3.1] The noise sequence is defined as a martingale difference with E[℧_{n+1} | F_n] = 0, but the update uses ℧_n. Please clarify the indexing and the measurability of ℧_n with respect to F_n; the current notation makes the conditional expectation arguments in Lemma 3.2 ambiguous.
  3. [Definition 5.3] The trimming operator Trim_k is defined by 'zero out the k largest-magnitude coordinates ... in a chosen basis'. This depends on a basis choice, which is not natural in a general Banach space and is not invariant under basis changes; the paper does not discuss how this affects the analysis.
  4. [Section 6] The numerical experiments are performed on the probability simplex in R^d, which is a finite-dimensional Euclidean setting. They therefore do not exercise the claimed reflexive Banach-space framework and cannot validate the ill-posed primal-dual addition in Definition 3.1.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the claimed convergence and rate results are derived from stated assumptions via martingale arguments, with no fitted target, self-citation chain, or imported uniqueness result.

full rationale

The paper's main results (Theorems 3.3, 4.3, 5.2, Proposition 5.4) are presented as consequences of assumptions (A1)-(A4) / (B1)-(B2) through the Robbins-Siegmund lemma and Bregman identities, not as predictions fitted from the values they claim to bound. There are no load-bearing self-citations: none of the references [1]-[15] is authored by Hashemi Sababe or Lotfali Ghasab, and no uniqueness theorem or ansatz is imported from the authors' prior work. The Hilbert-space specialization in Definition 3.1 is explicitly identified as the classical SKM scheme, which is a sanity check rather than a renamed result. Accordingly, the derivation chain does not reduce to its own inputs by construction. Two non-circular mathematical defects should nevertheless be flagged for correctness: (i) Definition 3.1 and Algorithms 1-2 add a dual-space martingale difference ℇₙ ∈ X* to the primal-space point ℏ(ζₙ) ∈ X inside ∇ϑ, without any stated identification of X with X*; the iteration is therefore undefined in a general reflexive Banach space, and the same defect propagates to Theorems 3.3, 5.2 and Proposition 5.4. (ii) In the proof of Theorem 4.3, after correctly obtaining Dₙ ≥ δ(...) and c‖ζₙ − ℏ(ζₙ)‖^q ≤ Dₙ, the text substitutes δ(...) ≥ c‖...‖^q ≥ c(Dₙ/c) = Dₙ, reversing the inequality; the O(A_N^{-p}) bound is not justified as written. These are serious proof and well-posedness gaps, but they are not equivalences-by-construction or fitted-input predictions, so they do not raise the circularity score.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The main theoretical results rest on a Hilbertian identification of primal and dual spaces (unstated), on a false Lipschitz-gradient consequence, and on standard tools. These explain why the convergence proof fails in the claimed Banach generality.

assumptions (5)
  • domain assumption X is a reflexive Banach space and ϑ is a Legendre, uniformly convex function.
    Assumptions (A1)-(A2) in Section 3; this is the intended setting.
  • standard math Robbins-Siegmund theorem applies to the supermartingale inequality with α_n^2 summable.
    Cited as [15]; used in Theorem 3.3.
  • ad hoc to paper The primal vector ℏ(ζ_n)+℧_n is well-defined for ℧_n ∈ X*, i.e., X and X* are identified or ℧_n is embedded in X.
    Definition 3.1; this assumption is not stated and fails for general reflexive Banach spaces.
  • ad hoc to paper Uniform convexity of ϑ implies the gradients ∇ϑ and ∇ϑ* are Lipschitz on bounded sets with constant L.
    Used in Lemma 3.2 to bound ∥ζ_{n+1}-ζ_n∥; false in general, e.g., ϑ(x)=|x|^p with 1<p<2.
  • domain assumption Bregman-Fejér monotonicity Dϑ(ζ_{n+1}, ζ*) ≤ Dϑ(ζ_n, ζ*) holds for fixed points of ℏ.
    Invoked in Theorem 3.3 to show all weak cluster points coincide; not proved, but a known property in some settings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Stochastic Krasnosel skii-Mann Iterations in Banach Spaces with Bregman Distances." pith.science (2026). https://pith.science/paper/FVM7QQ5K

@misc{pith2026250608031,
  author       = {Pith},
  title        = {Pith review of: Stochastic Krasnosel skii-Mann Iterations in Banach Spaces with Bregman Distances},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FVM7QQ5K}},
  note         = {Machine review of arXiv:2506.08031}
}
abstract

We propose a generalization of the stochastic Krasnoselskil-Mann $(SKM)$ algorithm to reflexive Banach spaces endowed with Bregman distances. Under standard martingale-difference noise assumptions in the dual space and mild conditions on the distance-generating function, we establish almost-sure convergence to a fixed point and derive non-asymptotic residual bounds that depend on the uniform convexity modulus of the generating function. Extensions to adaptive Bregman geometries and robust noise models are also discussed. Numerical experiments on entropy-regularized reinforcement learning and mirror-descent illustrate the theoretical findings.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 15 canonical work pages

  1. [1]

    W. R. Mann, Mean value methods in iteration , Proc. Amer. Math. Soc. 4 (1953), 506-510

  2. [2]

    M. A. Krasnoselski ˘ ı,Two remarks on the method of successive approximations , Uspekhi Mat. Nauk 10 (1955), no. 1(63), 123-127 (in Russian)

  3. [3]

    H. H. Bauschke and P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces , Springer, New York, 2011

  4. [4]

    L. M. Bregman, The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming , USSR Comput. Math. Math. Phys. 7 (1967), 200-217

  5. [5]

    Csisz ˜A¡r, Information-type measures of difference of probability distributions and indirect observation, Studia Sci

    I. Csisz ˜A¡r, Information-type measures of difference of probability distributions and indirect observation, Studia Sci. Math. Hungar. 2 (1967), 299-318

  6. [6]

    A. S. Nemirovski and D. B. Yudin, Problem Complexity and Method Efficiency in Optimization , Wiley- Interscience, New York, 1983

  7. [7]

    Beck and M

    A. Beck and M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Oper. Res. Lett. 31 (2003), no. 3, 167-175

  8. [8]

    Censor and S

    Y. Censor and S. Reich, The Dykstra algorithm with Bregman projections , Commun. Appl. Anal. 5 (2001), no. 2, 113-121

Show all 15 references
  1. [9]

    Cegielski, Iterative Methods for Fixed Point Problems in Hilbert Spaces , Springer Monographs in Mathe- matics, Springer, Cham, 2012

    A. Cegielski, Iterative Methods for Fixed Point Problems in Hilbert Spaces , Springer Monographs in Mathe- matics, Springer, Cham, 2012

  2. [10]

    Juditsky, A

    A. Juditsky, A. Nemirovski, and C. Tauvel, Solving variational inequalities with stochastic mirror-prox algo- rithm, Math. Program. 127 (2011), no. 1, 205-226

  3. [11]

    R. T. Rockafellar, Convex Analysis, Princeton Univ. Press, Princeton, NJ, 1970

  4. [12]

    Nemirovski, Robust stochastic approximation approach to stochastic programming , SIAM J

    A. Nemirovski, Robust stochastic approximation approach to stochastic programming , SIAM J. Optim. 19 (2009), no. 4, 1574-1609

  5. [13]

    Cioranescu, Geometry of Banach Spaces, Duality Mapping and Nonlinear Problems , Kluwer Academic Pub- lishers, Dordrecht, 1990

    I. Cioranescu, Geometry of Banach Spaces, Duality Mapping and Nonlinear Problems , Kluwer Academic Pub- lishers, Dordrecht, 1990

  6. [14]

    Neveu, Discrete-Parameter Martingales, North-Holland, Amsterdam, 1975

    J. Neveu, Discrete-Parameter Martingales, North-Holland, Amsterdam, 1975

  7. [15]

    Robbins and D

    H. Robbins and D. Siegmund, A convergence theorem for non-negative almost supermartingales and some applications, in Proc. Sympos. Math. Statist. Probab. , Vol. 4, Academic Press, New York, 1971, 233-257. R&D Section, Data Premier Analytics, Edmonton, Canada. Email address : H...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.