{"id":"09d5168f-c40b-4aac-ac33-faf87d05bb9e","arxiv_id":"2506.08031","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper claims a Bregman-SKM algorithm in Banach spaces with a.s. convergence and O(A_N^{-p}) rates, but the update is ill-posed and the proof contains a reversed inequality.","lead":"This paper proposes a Bregman-distance version of the stochastic Krasnoselskii-Mann fixed-point algorithm for reflexive Banach spaces, claiming almost-sure convergence and non-asymptotic residual rates. It would matter for non-Euclidean stochastic optimization and reinforcement learning, but the core update is not well-defined in general Banach spaces and the rate proof reverses a key inequality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Definition 3.1 adds dual-space noise ℧_n ∈ X* to primal point ℏ(ζ_n) ∈ X inside ∇ϑ; since no identification of X with X* is assumed, the Bregman-SKM iteration is undefined in general reflexive Banach spaces and Theorems 3.3 and 4.3 do not hold as stated.","rationale":"The reader's verdict is REJECT with high correctness risk, and the load-bearing defect is exactly the one identified as the weakest assumption: the update in Definition 3.1 adds a noise vector declared to lie in X* to a primal-space vector ℏ(ζ_n) ∈ X, with no identification of X and X* assumed. This makes the algorithm undefined in the claimed general reflexive Banach-space setting, which in turn invalidates the statements of Theorems 3.3 and 4.3 as written. The Hilbert-space case in §3 and the finite-dimensional experiments mask the problem because the Riesz identification is available there; this is precisely the setting the paper claims to generalize. The proof of Lemma 3.2 and the rate proof in Theorem 4.3 also contain problematic steps, including the asserted local Lipschitz property of ∇ϑ following from uniform convexity and the reversed inequality in the residual bound, but these are secondary to the fact that the iteration itself is not defined in the claimed domain. My proposed ℓ^p check would settle the primary concern directly. Since the same issue was already identified by the reader and no repair appears in the manuscript, the REJECT verdict stands unchanged.","tokens_in":10933,"tokens_out":4395,"duration_ms":46076,"concrete_test":"Instantiate the algorithm in X = ℓ^p for p ∈ (1,∞), p ≠ 2, with ϑ(z) = (1/p)‖z‖_p^p and a nonexpansive ℏ that has a fixed point. Take a nonzero ℧_n ∈ X* (for example, a single coordinate functional in the standard dual basis) and write out the expression ℏ(ζ_n) + ℧_n required by Definition 3.1. If no injection X* → X is supplied and the expression is undefined, the central claim fails; if the authors instead intend the repaired update with noise added after ∇ϑ(ℏ(ζ_n)), verify that the stated assumptions include this correction and that the proofs in §3–§4 are re-derived for the corrected iteration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Definition 3.1 defines ζ_{n+1} = ∇ϑ*((1−α_n)∇ϑ(ζ_n) + α_n ∇ϑ(ℏ(ζ_n)+℧_n)), while Definition 2.6 and assumption (A4) require ℧_n ∈ X*. However, ℏ(ζ_n) ∈ X, so ℏ(ζ_n)+℧_n is a sum of vectors from two different spaces. The paper states no identification of X with X* in (A1)–(A4), and reflexive Banach spaces in general admit no canonical isometric identification of this kind; Hilbert spaces are a special case because Riesz identification makes X and X* interchangeable. The Hilbert-space reduction in §3 works only because of that identification, and the numerical experiments also take place in finite-dimensional Euclidean settings. Without such an identification, the central iteration is not well-defined, so Theorem 3.3's almost-sure convergence and Theorem 4.3's averaged residual bound are claims about an undefined process. A charitable repair would be to write the noisy update as ∇ϑ*((1−α_n)∇ϑ(ζ_n) + α_n(∇ϑ(ℏ(ζ_n)) + ℧_n)), or else to take the noise in X; either change would require a new algorithm definition and re-derivation of the results. The same defect propagates to Algorithm 1, Algorithm 2, Theorem 5.2, and Proposition 5.4, since all inherit the same primal-plus-dual addition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a stochastic Krasnosel'skii-Mann (SKM) iteration in reflexive Banach spaces using Bregman distances. The authors define a Bregman-SKM update, prove almost-sure convergence to a fixed point (Theorem 3.3), and derive non-asymptotic residual bounds (Theorem 4.3) under a uniform-convexity modulus condition. They also discuss adaptive Bregman geometries and heavy-tailed noise with trimming, and report numerical experiments on entropy-regularized policy iteration.","tokens_in":11332,"tokens_out":4144,"duration_ms":44511,"significance":"If correct, the results would generalize stochastic KM methods beyond Hilbert spaces and provide rates governed by the modulus of uniform convexity. However, the central algorithm is not well-defined as stated because it adds a dual-space noise element to a primal-space point, and the main proofs contain load-bearing gaps. The paper does provide a clear structure and the Hilbert-space special case reduces to known SKM, but the claimed Banach-space extension is not established.","major_comments":[{"comment":"The Bregman-SKM update is not well-defined in general reflexive Banach spaces. The update reads ζ_{n+1} = ∇ϑ*((1−α_n)∇ϑ(ζ_n) + α_n∇ϑ(ℏ(ζ_n)+℧_n)), with ℏ(ζ_n) ∈ X and (by Definition 2.6 and (A4)) ℧_n ∈ X*. The sum ℏ(ζ_n)+℧_n therefore adds a primal vector and a dual vector, which is meaningless unless X and X* are identified. No such identification is assumed in (A1)–(A4), and reflexive Banach spaces generally do not admit a canonical isometric identification of X with X*. This defect propagates to Algorithms 1 and 2, Theorem 5.2, and Proposition 5.4, all of which use the same primal-plus-dual addition. A repair would require redefining the algorithm, e.g., placing the noise in X or adding it after applying ∇ϑ, and then re-deriving all subsequent results.","section":"Definition 3.1"},{"comment":"The proof asserts without support that uniform convexity of ϑ implies Lipschitz continuity of ∇ϑ and ∇ϑ* on bounded sets. Uniform convexity alone does not imply differentiability beyond Gateaux differentiability, nor does it yield a Lipschitz gradient; standard results require additional smoothness assumptions such as uniform smoothness or a modulus of smoothness. Because the one-step decrease estimate (Lemma 3.2) is the foundation for Theorem 3.3 and Theorem 4.3, this missing hypothesis undermines the entire analysis. The paper needs to either add explicit smoothness assumptions or replace these steps with arguments that do not rely on unproved Lipschitz bounds.","section":"Lemma 3.2"},{"comment":"The proof contains a key inequality in the wrong direction. From D_n = Dϑ(ζ_n, ℏ(ζ_n)) ≥ δ(∥ζ_n−ℏ(ζ_n)∥) and δ(r) ≥ c r^q, one obtains D_n ≥ c∥ζ_n−ℏ(ζ_n)∥^q, i.e., c∥ζ_n−ℏ(ζ_n)∥^q ≤ D_n. The proof, however, substitutes δ(∥ζ_n−ℏ(ζ_n)∥) ≥ c∥ζ_n−ℏ(ζ_n)∥^q ≥ c(D_n/c) = D_n, which reverses the inequality. Consequently the drift term (1/2)D_n α_n in the displayed inequality is not justified. This invalidates the derivation of the averaged residual bound O(A_N^{-p}).","section":"Theorem 4.3"},{"comment":"The proof concludes that δ(∥ζ_n−ℏ(ζ_n)∥) → 0 a.s. from ∑ α_n δ(∥ζ_n−ℏ(ζ_n)∥) < ∞ a.s. and ∑ α_n = ∞. This implication is false in general: with α_n = 1/n, taking δ(x_n)=1 on a sparse subsequence and 0 elsewhere yields a finite sum ∑ α_n δ(x_n) while δ(x_n) does not tend to 0. Without an additional argument forcing δ(x_n) → 0, the conclusion that ∥ζ_n−ℏ(ζ_n)∥ → 0 and hence D_n → 0 does not follow from the Robbins–Siegmund lemma as applied. This is a load-bearing gap in the almost-sure convergence claim.","section":"Theorem 3.3"}],"minor_comments":[{"comment":"The text contains numerous typographical and encoding issues, such as 'Krasnosel ski ¨A', 'Fej ˜A©r', 'Fix(⟨⌊⊣∇)', and inconsistent spacing in the title. These should be corrected in a revision.","section":"Throughout"},{"comment":"The noise sequence is defined as a martingale difference with E[℧_{n+1} | F_n] = 0, but the update uses ℧_n. Please clarify the indexing and the measurability of ℧_n with respect to F_n; the current notation makes the conditional expectation arguments in Lemma 3.2 ambiguous.","section":"Definition 2.6 and Definition 3.1"},{"comment":"The trimming operator Trim_k is defined by 'zero out the k largest-magnitude coordinates ... in a chosen basis'. This depends on a basis choice, which is not natural in a general Banach space and is not invariant under basis changes; the paper does not discuss how this affects the analysis.","section":"Definition 5.3"},{"comment":"The numerical experiments are performed on the probability simplex in R^d, which is a finite-dimensional Euclidean setting. They therefore do not exercise the claimed reflexive Banach-space framework and cannot validate the ill-posed primal-dual addition in Definition 3.1.","section":"Section 6"}],"recommendation":"reject","confidential_remarks":"The central algorithm is not well-posed and the main technical results rely on incorrect or unsupported inequalities. These are not local presentation issues but fundamental flaws in the paper's contribution. A substantial rewrite would be required, including a new algorithm definition, additional smoothness assumptions, and corrected proofs of the convergence and rate theorems. Given the scope of the necessary changes, rejection is appropriate. The authors may wish to consult standard references on Bregman projections and duality mappings in Banach spaces before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's main iteration adds a dual-space noise term ℧_n ∈ X* to a primal-space point ℏ(ζ_n) ∈ X inside ∇ϑ. No identification of X with X* is stated, so in a general reflexive Banach space that sum is not well-defined. This is a load-bearing flaw, not a cosmetic gap: every theorem that follows describes an undefined process. A charitable fix exists (add the noise after differentiation, as ∇ϑ(ℏ(ζ_n)) + ℧_n), but that changes the algorithm and the analysis would need to be redone.\n\nThere is a reasonable motivation behind the paper. Extending Cegielski's stochastic KM to Bregman geometries is a natural step, and making rates depend on the uniform convexity modulus of ϑ is a sensible goal. The adaptive and heavy-tailed extensions point to real directions, and the entropy-regularized RL example is a nice showcase.\n\nThe problems are serious. Lemma 3.2 asserts that uniform convexity of ϑ implies ∇ϑ and ∇ϑ* are Lipschitz on bounded sets; that is false — for ϑ(x)=|x|^p/p with p<2, the derivative has singular behavior near zero. In the proof of Theorem 4.3, the substitution δ(∥ζ_n−ℏ(ζ_n)∥) ≥ D_n is backwards; the assumptions give D_n ≥ c∥·∥^q, hence ∥·∥^q ≤ D_n/c, not the opposite. So the O(A_N^{-p}) rate bound does not follow. The almost-sure convergence proof also hand-waves the weak-cluster-point argument, and the numerical experiments are entirely Euclidean (probability simplex with Gaussian or Student-t noise), so they do not exercise the claimed Banach-space setting.\n\nThis paper is best read as a rough working paper by someone already familiar with Bregman fixed-point methods, not as a reliable reference. The core defect is easy to state, and the paper needs major revision before it deserves referee time. As it stands, I would return it rather than send it to peer review.\n\nBottom line: reject as is; the right response is to ask the authors to fix the definition of the update and rework the proofs.","headline":"The core Bregman-SKM update is undefined in general Banach spaces and the rate proof inverts a key inequality; the intended extension is natural but the technical execution doesn't hold up.","tokens_in":11822,"tokens_out":4918,"would_cite":false,"duration_ms":53991,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["47H05","47J25","49M27","65K10","90C25"],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bregman-distance generalization of the stochastic Krasnosel'skii-Mann iteration converges almost surely to a fixed point in reflexive Banach spaces, with residual bounds governed by the uniform convexity modulus.","keywords":["stochastic fixed-point iteration","Bregman distance","Banach space","Krasnosel'skii-Mann","almost-sure convergence","uniform convexity modulus","martingale-difference noise","mirror descent"],"falsifier":"Look at Definition 3.1 in a concrete reflexive Banach space where \\(X\\neq X^*\\) as sets, such as \\(\\ell^p\\) for \\(p\\in(1,2)\\): the expression \\(\\hbar(\\zeta_n)+\\mho_n\\) sums an element of \\(X\\) with an element of \\(X^*\\), which is undefined, so the central algorithm has no meaning without an additional identification that the paper does not state.","tokens_in":10709,"feed_emoji":"📉","tokens_out":6859,"duration_ms":67371,"temperature":0.7,"pith_summary":"The paper proposes a stochastic version of the Krasnosel'skii-Mann iteration that works in reflexive Banach spaces instead of Hilbert spaces, using Bregman distances to adapt the geometry. It claims that, under martingale-difference noise and mild conditions on a Legendre distance-generating function, the iterates converge almost surely to a fixed point of a nonexpansive operator, and the Bregman residual goes to zero. It further derives non-asymptotic bounds on the averaged residual that depend on the uniform convexity modulus of the generating function. A sympathetic reader would care because many optimization and reinforcement-learning algorithms naturally live in non-Euclidean spaces, and this result would give them the same stochastic convergence guarantees that are already available in Hilbert spaces.","feed_headline":"Bregman-KM iterates converge almost surely in Banach spaces","feed_subtitle":"A distance-aware stochastic fixed-point algorithm now covers non-Euclidean settings like mirror descent and entropy-regularized RL.","key_machinery":"The central object is the Bregman-SKM update, a two-step map that pulls the current point into the dual space through the gradient of a Legendre function, averages it with the nonexpansive operator's output, and returns via the conjugate gradient. The argument is carried by the three-point identity for Bregman distances and the uniform convexity modulus \\(\\delta\\) of \\(\\vartheta\\): together they give a one-step residual decrease (Lemma 3.2) with a shrinking term proportional to \\(\\alpha_n \\delta(\\|\\zeta_n-\\hbar(\\zeta_n)\\|)\\). Summability of step-squares and a standard almost-supermartingale convergence lemma then force the residual to zero, while the rate exponent \\(p\\) in Theorem 4.3 appears from the polynomial lower bound on \\(\\delta\\).","core_discovery":"On its own terms, the paper establishes that the Bregman-SKM iteration, defined by \\(\\varsigma_n = \\nabla\\vartheta^*\\left((1-\\alpha_n)\\nabla\\vartheta(\\zeta_n)+\\alpha_n\\nabla\\vartheta(\\hbar(\\zeta_n)+\\mho_n)\\right)\\) and \\(\\zeta_{n+1}=\\varsigma_n\\), is almost-surely convergent: under assumptions (A1)-(A4), \\(\\zeta_n \\to \\zeta^*\\) for some \\(\\zeta^* \\in \\mathrm{Fix}(\\hbar)\\) and \\(D_\\vartheta(\\zeta_n,\\hbar(\\zeta_n)) \\to 0\\) almost surely. When the modulus of uniform convexity satisfies \\(\\delta(r) \\ge c r^q\\), the window-averaged residual satisfies \\(\\bar{R}_N = O($A_N^{{-p}}$)\\) with \\(p=(q-1)/q\\), recovering the classical \\(O(1/\\sqrt{n})\\) Hilbert-space rate for \\(q=2\\).","pith_inferences":["The update rule in Definition 3.1 adds a dual-space noise vector \\(\\mho_n \\in X^*\\) to the primal-space vector \\(\\hbar(\\zeta_n) \\in X\\) inside the same argument; in a general reflexive Banach space this sum is not defined unless the paper silently identifies \\(X\\) with \\(X^*\\), which holds for Hilbert spaces but not for, say, \\(\\ell^p\\) with \\(p\\neq 2\\).","If the well-posedness gap is repaired, the assumption that the noise lives in the dual space suggests a natural interpretation: the stochastic perturbation affects the gradient of the Bregman function, not the operator evaluation itself, so a cleaner formulation might apply noise after \\(\\nabla\\vartheta\\) rather than inside it.","The rate exponent \\(p=(q-1)/q\\) implies that strengthening uniform convexity (larger \\(q\\)) drives the exponent to 1, so one could design distance-generating functions with high-order convexity to approach linear convergence; the paper leaves such a construction open.","The trimming results under heavy-tailed noise depend on an order-statistic bound for the removed coordinates; a direct numerical test in \\(\\ell^p\\) with Student-t noise would reveal whether the logarithmic trimming schedule behaves as predicted when the spaces are not identified."],"forward_implications":["Almost-sure convergence now holds for stochastic fixed-point iterations in reflexive Banach spaces, so entropy-regularized reinforcement learning and mirror-descent variants can be analyzed in their native geometry.","When \\(\\vartheta(\\zeta)=\\tfrac12\\|\\zeta\\|^2\\) in a Hilbert space, the new bounds reduce to the known \\(O(1/\\sqrt{n})\\) averaged residual, giving a unified framework rather than a separate theory.","With polynomial step-sizes \\(\\alpha_n=n^{-\\gamma}\\) for \\(\\gamma\\in(1/2,1)\\), the averaged residual scales as \\(O(N^{-p(1-\\gamma)})\\), so the rate improves as the uniform-convexity exponent \\(q\\) grows.","Adaptive Bregman geometries (time-varying \\(\\vartheta_n\\)) and trimmed heavy-tailed noise both preserve almost-sure convergence under the stated conditions."],"supporting_citations":[{"why":"introduces the stochastic KM iteration in Hilbert spaces that this paper generalizes, including step-size conditions.","marker":"[9]"},{"why":"provides the almost-supermartingale convergence lemma used to prove almost-sure convergence in Theorem 3.3.","marker":"[15]"},{"why":"supplies the uniform-convexity estimate linking norm error to Bregman distance, used for the residual decrease.","marker":"[7]"},{"why":"gives the three-point identity for Bregman distances used in the one-step decrease lemma.","marker":"[8]"},{"why":"provides fixed-point structure of nonexpansive maps in reflexive spaces, used to identify cluster points.","marker":"[3]"}],"fun_headline_variants":["Bregman-KM converges almost surely in reflexive Banach spaces","Stochastic KM with Bregman distances: a.s. convergence and rates","Banach-space Bregman-KM: almost-sure fixed-point convergence","Bregman-KM in Banach spaces: almost-sure convergence and rates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The core premise is that the algorithmic update makes sense, but the update adds a dual-space noise term directly to a primal-space operator output, so the algorithm is well defined only when the space and its dual are identified—true in Hilbert spaces but not in the general reflexive Banach spaces the paper claims to cover.","fun_headline_variants_meta":{"raw":{"variants":["Bregman-KM converges almost surely in reflexive Banach spaces","Stochastic KM with Bregman distances: a.s. convergence and rates","Banach-space Bregman-KM: almost-sure fixed-point convergence","Bregman-KM in Banach spaces: almost-sure convergence and rates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000604,"raw_usage":{"total_tokens":2770,"prompt_tokens":848,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1840}},"tokens_in":464,"tokens_out":1922,"duration_ms":15399,"temperature":1.0,"reasoning_tokens":1840,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:44:32.905565+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look at Definition 3.1 in a concrete reflexive Banach space where \\(X\\neq X^*\\) as sets, such as \\(\\ell^p\\) for \\(p\\in(1,2)\\): the expression \\(\\hbar(\\zeta_n)+\\mho_n\\) sums an element of \\(X\\) with an element of \\(X^*\\), which is undefined, so the central algorithm has no meaning without an additional identification that the paper does not state.","supporting_citations":[{"cited_title":"Cegielski, Iterative Methods for Fixed Point Problems in Hilbert Spaces , Springer Monographs in Mathe- matics, Springer, Cham, 2012","cited_arxiv_id":null,"evidence_quote":"introduces the stochastic KM iteration in Hilbert spaces that this paper generalizes, including step-size conditions."},{"cited_title":"Robbins and D","cited_arxiv_id":null,"evidence_quote":"provides the almost-supermartingale convergence lemma used to prove almost-sure convergence in Theorem 3.3."},{"cited_title":"Beck and M","cited_arxiv_id":null,"evidence_quote":"supplies the uniform-convexity estimate linking norm error to Bregman distance, used for the residual decrease."},{"cited_title":"Censor and S","cited_arxiv_id":null,"evidence_quote":"gives the three-point identity for Bregman distances used in the one-step decrease lemma."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides fixed-point structure of nonexpansive maps in reflexive spaces, used to identify cluster points."}],"review_version":1}