{"id":"97d38a95-bd6c-4681-8ad9-9cbccae6793e","arxiv_id":"2505.23517","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Inexact JKO and proximal-gradient schemes in Wasserstein space converge weakly to minimizers with O(1/σ_n) energy rates, provided the error sequences are summable and the stepsize sums diverge.","lead":"This paper proves that JKO and proximal-gradient algorithms in the space of probability distributions still converge to a minimizer when each optimization step is solved only approximately, as long as the errors are controlled and summable. The results give rigorous guarantees for the approximate JKO steps that numerical solvers actually produce, which matters for sampling, mean-field optimization, and generative modeling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central assumption—generalized-geodesic convexity with arbitrary base in P2—fails for the paper's flagship free-energy example, so the main JKO convergence results may not apply to the motivating case.","rationale":"The reader's weakest_assumption already identifies the load-bearing issue: the paper needs generalized-geodesic convexity with base points in all of P2, not merely in D(G) or in the regular measures, because inexact iterates can leave the domain. My review confirms this is genuinely load-bearing and shows it is not just a stronger-than-usual condition but one violated by the paper's own flagship example. The proof skeleton of Theorem 3.3 is otherwise plausible: the quasi-Fejér monotonicity, summability arguments, and Opial-based convergence follow the standard inexact proximal-point template, and the presentation errors flagged by the reader (e.g., equation (3.8) using G(µ_{n+1}) instead of G(Jτn(µn)), the W2/W2² mixing in the Opial step, and the factor mismatch in the variational error definition) appear fixable without changing the mathematical architecture. The most serious limitation is that the main theorems' assumptions are not satisfied by the free-energy example when approximate iterates are atomic/discrete, which is exactly the scenario the inexact framework is designed to address. This warrants keeping the verdict CONDITIONAL: the paper should either prove the claimed extension of generalized-geodesic convexity from regular to arbitrary base points under lower semicontinuity (which the text asserts but does not prove), or clearly state that the convergence results apply only to functionals satisfying the stronger Definition 2.9 and rephrase the free-energy motivation accordingly. Since the reader already conditioned acceptance on these points, I do not move the verdict.","tokens_in":26501,"tokens_out":16964,"duration_ms":164188,"concrete_test":"Let G(µ) = ∫(1/2)||x||² dµ(x) + Ent(µ), take µ0 = N(−1, I_d), µ1 = N(1, I_d), and base ν = δ0. Compute the unique generalized geodesic µ_t = (1−t)µ0 + tµ1 (any gluing gives the same linear mixture) and evaluate the convexity inequality at t = 1/2: check whether Ent(µ_{1/2}) ≤ (Ent(µ0) + Ent(µ1))/2. For two well-separated unit Gaussians, the mixture has variance 1 + ||1−(−1)||²/4 = 1 + 1 = 2 per coordinate for d=1, while each component has variance 1, so Ent(µ_{1/2}) − (Ent(µ0)+Ent(µ1))/2 = (1/2)log(2) > 0. This explicit positive gap confirms failure of Definition 2.9 for the paper's flagship example; rerun the check with higher separation to make the gap arbitrarily large.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The discrete EVI is extended to all base points in Theorem 2.13(ii) only under Definition 2.9, where G must be convex along generalized geodesics with base points in all of P2(Rd), not just D(G). This extension is indispensable because inexact iterates µn can leave D(G), and Lemma 3.1 applies the EVI with µ = µn. The paper's motivating functional is G(µ) = ∫F dµ + Ent(µ), with H = Ent. But Ent is not convex along generalized geodesics with an atomic base. For base ν = δ0, any generalized geodesic between two regular measures µ0, µ1 is their linear mixture (1−t)µ0 + tµ1, because the only optimal plans from δ0 to µ0 and δ0 to µ1 are product plans, and the gluing must have δ0 as first marginal. Along mixtures, entropy is concave, so Ent((1−t)µ0 + tµ1) > (1−t)Ent(µ0) + tEnt(µ1) for distinct, sufficiently separated measures. Thus the required inequality (2.3) fails for the prototype functional whenever the inexact iterate is atomic. The text's assertion that the Definition 2.9 notion 'coincides' with the regular-base notion under lower semicontinuity is not substantiated; if it were true, it would force Ent to be convex along atomic-base generalized geodesics, contradicting the explicit computation above. Consequently, Theorem 3.3 and Theorem 3.8 do not cover the free-energy example in the practical regime where approximate JKO outputs are discrete or atomic, despite the abstract and introduction presenting this as the central application. This is not a minor technicality: the whole motivation of the inexact framework is that solvers return approximate, often discrete, measures.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proves convergence results for inexact JKO schemes and inexact proximal-gradient algorithms in the 2-Wasserstein space. Two inexactness models are considered: a distance-type error W2(μ_{n+1}, J_{τ_n}(μ_n)) ≤ ε_n and a variational-type error in the value of the proximal objective. The main theorems (Theorems 3.3, 3.8, 4.6, 4.10) assert O(1/σ_n) rates for the objective at the best or last iterates, and convergence of the generated sequence in the τ_{w,2} topology (hence in W_p for p < 2) to a minimizer of G, under summability assumptions on the error and weighted-error sequences. The analysis for the proximal-gradient scheme relies on a refined discrete EVI (Lemma 4.2) and on the composite structure G = E_F + H with dom(H) contained in the regular measures.","tokens_in":26736,"tokens_out":31758,"duration_ms":347466,"significance":"If the results hold, they fill a genuine gap: numerical solvers for the JKO step are inexact, yet convergence analyses in the Wasserstein literature mostly assume exact proximal steps. Extending the classical inexact-proximal-point template to a non-Hilbertian setting via the Opial property is a useful contribution, and the proximal-gradient part goes beyond the Bures–Wasserstein results of Diao et al. The assumptions are explicit and the error conditions are natural analogues of those used in Hilbert-space analyses. The paper would be strengthened by making the proof of the general-base convexity extension fully explicit, but the main body of the analysis appears sound; the stress-test concern about the free-energy example does not, on close reading, invalidate the main theorems.","major_comments":[{"comment":"The paper assumes convexity along generalized geodesics with base points ν in all of P2(R^d), whereas the standard notion in [4, Def. 9.2.4] and in [63] only uses base points in D(G) or in P^r_2(R^d). This extension is necessary because inexact iterates can leave D(G), but the assertion in Definition 2.9 that the two notions coincide under lower semicontinuity is not proved. A proof can be supplied by approximating an arbitrary base ν by regular measures ν_k, extracting a limit of the associated triples of couplings, and using lower semicontinuity of G to pass to the limit in (2.3). The potential counterexample with Ent and base δ0 does not land: for that base one may take the optimal coupling between the endpoints, so the generalized geodesic reduces to a displacement geodesic, and Ent is convex along displacement geodesics. Please add the missing argument, since Theorem 2.13(ii) and its use in Lemma 3.1 depend on this point.","section":"Definition 2.9 and Theorem 2.13"},{"comment":"The displayed inequality in (3.8) uses G(μ_{n+1}), but β_N is defined as the best iterate among {G(J_{τ_i}(μ_i))}. The correct inequality is G(β_N) ≤ (1/σ_N)∑_{n=0}^{N-1} τ_n G(J_{τ_n}(μ_n)); with G(μ_{n+1}) the bound (3.9) does not follow from (3.7). Please correct this line and the surrounding text. In the same proof, Lemma 3.1 writes W2(J_{τ_n}(μ_{n+1}),ν), which should read W2(J_{τ_{n+1}}(μ_{n+1}),ν); otherwise the claimed convergence of {W2(J_{τ_n}(μ_n),ν)} is not established.","section":"Theorem 3.3, Eq. (3.8)"},{"comment":"The reduction from the variational error (3.15) to W2(μ_{n+1}, J_{τ_n}(μ_n)) ≤ ε_n contains inconsistent constants: the two displayed inequalities have (1−t)2ε_n^2/τ_n and (1−t)ε_n^2/(2τ_n), respectively, and as written the limiting argument yields 4ε_n^2 rather than ε_n^2. The intended conclusion is correct with δ = ε_n^2/(2τ_n), since the standard derivation gives W2^2 ≤ 2τ_nδ/t and hence W2^2 ≤ ε_n^2 as t → 1, but the displayed computation should be fixed.","section":"Theorem 3.8, proof after Eq. (3.16)"}],"minor_comments":[{"comment":"In the proof of Theorem 2.13(iii), the line defining ν as J_{τ}(μ_n) should read J_{τ}(μ); otherwise the subsequent continuity argument is not the intended one.","section":"Theorem 2.13(iii)"},{"comment":"In the proof of Lemma 2.5, the statement that τ_{w,2}-convergence of μ̄_n to μ* implies W1(μ̄_n, μ*) → 0 should cite the relevant fact that τ_{w,2} convergence gives W_p convergence for p < 2; as written the sentence looks circular.","section":"Lemma 2.5"},{"comment":"The function ℓ(ν) is defined as the limit of W2(J_{τ_n}(μ_n), ν), but in the Opial argument it is used as the limit of squared distances; please define ℓ with W2^2 consistently.","section":"Theorem 3.3, Opial step"},{"comment":"There are several typos and minor notational slips: 'Wassertein', 'contant', 'demostrate', 'propriety'; Theorem 4.6 writes ε_{n−1} for n = 0; Corollary 4.9 has S_{w,2} instead of τ_{w,2}; and the display (3.8) mislabels the index set for j_n. These should be corrected in a revision.","section":"Throughout"},{"comment":"The ergodic remarks cite [1, Proposition 7.6] for convexity along barycenters; please state explicitly that the JKO images are assumed regular there, since the cited result is stated for regular measures and the general-energy-density case is left as an expectation.","section":"Remarks 3.6 and 4.8"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about Definition 2.9 excluding the free-energy example does not survive scrutiny: for base δ0 one can choose an optimal coupling between the endpoints and obtain a displacement geodesic, along which Ent is convex, and the regular-base case can be extended by approximation under lower semicontinuity. The manuscript is therefore likely correct in its central claims, but the missing proof of that extension and the algebra/indexing errors in the proofs of Theorems 3.3 and 3.8 are load-bearing and should be fixed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper proves what it claims, the results are new, and the stress-test note you passed me is wrong. The worry was that entropy fails the paper's Definition 2.9, which would sink the main free-energy application. That specific concern misreads the definition. For base ν = δ0, the definition only requires existence of a generalized geodesic along which convexity holds, and you can choose the coupling between μ0 and μ1 to be optimal. The resulting generalized geodesic is the standard displacement interpolation, and entropy is convex along that. The stress-test incorrectly assumes the coupling must be the product plan, so the motivating example is not a counterexample.\n\nWhat is actually new: the paper extends inexact proximal-point analysis with summable errors to the Wasserstein setting, where nonexpansivity of Jτ is unavailable, using the Opial property. The refined discrete EVI in Lemma 4.2, which drops the regularity condition on the input for the proximal-gradient map, is a genuine improvement over Salim–Korba–Luise. The theorems give O(1/σ_n) rates on the best iterates and weak convergence under weighted summable errors. That is a real step forward for anyone doing approximate JKO with ADMM, entropic, or neural approximations.\n\nSoft spots: the paper is not polished enough as written. Equation (3.8) uses G(μ_{n+1}) where it should be G(J_{τ_n}(μ_n)); the Opial step mixes W2 and W2^2; the variational error in the introduction differs from Theorem 3.8 by a factor 1/(2τ_n); and (4.16) mixes τ and τ_n. All are fixable, but a referee should ask for a careful revision. The stronger convexity notion (Definition 2.9) is genuinely restrictive, and the paper's claim that it coincides with the regular-base notion under lower semicontinuity needs a proof or a reference. For the proximal-gradient part, assumption (A3) excludes atomic measures, which limits direct application to discrete solvers; the authors acknowledge this and conjecture it can be removed.\n\nCitation pattern is fine; reliance on Naldi–Savaré for the Opial property is legitimate. The proofs are genuine derivations, no fitted constants. This paper deserves a serious referee and should not be desk rejected. Send it to review, with expectations of a minor-to-moderate revision. I'd bring it to reading group and cite it if I worked on Wasserstein optimization.","headline":"A genuinely useful paper; the stress-test's entropy counterexample fails, and the remaining issues are fixable typos and a real but nonfatal convexity restriction.","tokens_in":27406,"tokens_out":7109,"would_cite":true,"duration_ms":67868,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","90C25","65K10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that approximate JKO steps, with errors measured in Wasserstein distance or in energy, still converge weakly to minimizers at the usual convex rate $O(1/\\sigma_n)$, provided the per-step errors are summable and satisfy a…","keywords":["JKO scheme","inexact proximal point method","proximal-gradient algorithm","Wasserstein space","weak convergence","discrete EVI","geodesic convexity","optimal transport"],"falsifier":"Check whether the discrete EVI (2.5), $W_2^2(J_\\tau(\\mu),\\nu) \\le W_2^2(\\mu,\\nu) - 2\\tau(G(J_\\tau(\\mu))-G(\\nu)) - W_2^2(J_\\tau(\\mu),\\mu)$, holds for a proper lower semicontinuous geodesically convex $G$ when the input $\\mu$ is a two-atomic measure outside $D(G)$; a single failure would break the proof of Theorem 3.3 for functionals whose domain excludes such measures.","tokens_in":26191,"feed_emoji":"🎯","tokens_out":9420,"duration_ms":91426,"temperature":0.7,"pith_summary":"The paper asks what happens when the central minimization step of the Jordan–Kinderlehrer–Otto (JKO) scheme, and of the related proximal-gradient algorithm in the Wasserstein space, is solved only approximately—as is unavoidable in practice. Its main theorems say that if each approximate step is within a controlled error of the exact one, measured either in 2-Wasserstein distance or in the value of the energy being minimized, then the generated sequence still converges weakly to a minimizer of the functional at the convex rate $O(1/\\sigma_n)$, where $\\sigma_n$ is the cumulative sum of stepsizes. The price is that the errors must be summable and must satisfy a weighted summability condition involving $\\sigma_n/\\tau_n$, conditions that parallel the ones used in Hilbert-space inexact proximal-point theory. A sympathetic reader should care because exact JKO steps are known in closed form only in very few cases, so the practical validity of Wasserstein gradient-flow algorithms rests on exactly the error control this paper formalizes.","feed_headline":"Inexact JKO still converges to minimizers","feed_subtitle":"Summable per-step errors preserve weak convergence and the usual convex convergence rates.","key_machinery":"The load-bearing object is the discrete evolution variational inequality (EVI) for the proximal map, extended to all input measures rather than only to inputs inside the domain of $G$: $W_2^2(J_\\tau(\\mu),\\nu) \\le W_2^2(\\mu,\\nu) - 2\\tau(G(J_\\tau(\\mu))-G(\\nu)) - W_2^2(J_\\tau(\\mu),\\mu)$. This inequality converts each inexact step into a quasi-contraction of squared Wasserstein distances, and the extra negative term $W_2^2(J_\\tau(\\mu),\\mu)$ is what makes the increments $\\sum_n W_2^2(J_{\\tau_n}(\\mu_n),\\mu_n)$ summable. Full sequence convergence is then obtained from the Opial property of the weak topology $\\tau_{w,2}$, which upgrades the fact that every cluster point is a minimizer to convergence of the whole sequence. The proximal-gradient part uses a finer EVI, $W_2^2(S_\\tau(\\mu),\\nu) \\le (1-\\tau\\lambda)W_2^2(\\mu,\\nu) - 2\\tau(G(S_\\tau(\\mu))-G(\\nu)) - (1-\\tau L)W_2^2(\\mu,S_\\tau(\\mu))$, whose extra coefficient $(1-\\tau L)$ plays the same summability role.","core_discovery":"The central claim is that approximate computation of JKO steps does not destroy convergence. For a proper, lower semicontinuous functional $G$ that is convex along generalized geodesics and admits a minimizer, any sequence satisfying the distance-type error bound $W_2(\\mu_{n+1}, J_{\\tau_n}(\\mu_n)) \\le \\epsilon_n$ with $\\sum_n \\epsilon_n < \\infty$, $\\sum_n \\tau_n = \\infty$, and $\\sum_n (\\sigma_n/\\tau_n)\\epsilon_{n-1}^2 < \\infty$ has the property that the energies of the exact JKO outputs converge at rate $G(J_{\\tau_n}(\\mu_n))-\\inf G = O(1/\\sigma_n)$, and the sequence itself converges in the $\\tau_{w,2}$ topology—hence in $W_p$ for every $p<2$—to a minimizer of $G$. The same conclusion holds under the variational-type error bound (1.3), with the added benefit that $G(\\mu_n)-\\inf G$ itself is $O(1/\\sigma_n)$. The proximal-gradient analogue, for $G=E_F+H$ under assumptions (A1)–(A3), inherits both convergence statements, with the corresponding weighted error conditions; Theorem 4.6 thereby extends the known weak-convergence result for exact proximal-gradient methods beyond the Gaussian/Bures–Wasserstein setting.","pith_inferences":["The weighted condition $\\sum_n(\\sigma_n/\\tau_n)\\epsilon_{n-1}^2<\\infty$ gives a concrete budget for inner solvers: early errors may be large, but errors must decay fast enough relative to the accumulated stepsize; this is a ready-made stopping criterion for algorithms such as ALG2 once a $W_2$ error estimate is available.","The authors' conjecture that the domain restriction $\\mathrm{dom}(H)\\subseteq P_2^r$ can be removed via a different subdifferential calculus, if true, would let the proximal-gradient results cover entropy-type functionals evaluated at atomic or singular measures.","The same proof template could apply to other approximate JKO constructions, such as entropic-regularized JKO, provided one can quantify the distance between the entropic proximal output and the exact JKO output—a step the paper explicitly leaves open."],"forward_implications":["If a numerical solver returns each JKO step with a $W_2$ error that is summable and satisfies the weighted condition $\\sum_n(\\sigma_n/\\tau_n)\\epsilon_{n-1}^2<\\infty$, the computed iterates are guaranteed to converge weakly to a minimizer, so approximate solvers can replace exact JKO steps without destroying convergence.","Under variational-type errors, the actually accessible values $G(\\mu_n)$ converge at the rate $O(1/\\sigma_n)$, not only the unobservable values $G(J_{\\tau_n}(\\mu_n))$.","The proximal-gradient algorithm for composite functionals $E_F+H$ keeps the same guarantees under the stated error conditions, generalizing existing weak-convergence results from the Gaussian/Bures–Wasserstein case to the full Wasserstein space.","Variable and vanishing stepsizes are allowed without changing the rates, which is relevant for correcting the bias of unadjusted Langevin-type sampling schemes."],"supporting_citations":[{"why":"Supplies the classical JKO well-posedness and the base-point-in-domain EVI that Theorem 2.13 extends to all inputs.","marker":"[4]"},{"why":"Defines the exact JKO scheme whose inexact version is the subject of the paper.","marker":"[47]"},{"why":"Provides the weak topology $\\tau_{w,2}$, sequential compactness, and Opial property used to upgrade cluster points to full convergence.","marker":"[58]"},{"why":"Establishes the exact Wasserstein proximal-gradient algorithm and the EVI whose refined form Lemma 4.1 derives.","marker":"[63]"},{"why":"Gives the weak-convergence result in the Bures–Wasserstein space that Theorem 4.6 generalizes.","marker":"[34]"},{"why":"Provides the ALG2/ADMM numerical method that motivates the distance-type and variational-type inexactness analysis.","marker":"[9]"},{"why":"Interprets Langevin sampling as a composite optimization problem, motivating the inexact proximal-gradient analysis with variable stepsizes.","marker":"[69]"}],"fun_headline_variants":["Inexact JKO steps still converge to minimizers","Approximate JKO keeps convergence alive","Weak convergence survives inexact JKO","JKO tolerates errors without losing limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The key premise is that the energy is convex along generalized geodesics with base points allowed anywhere in the space of probability measures, not just inside the energy's domain, because approximate JKO iterates can leave the domain of $G$; for many natural energies only convexity along ordinary geodesics is known, and the discrete EVI used throughout the proof requires the stronger property.","fun_headline_variants_meta":{"raw":{"variants":["Inexact JKO steps still converge to minimizers","Approximate JKO keeps convergence alive","Weak convergence survives inexact JKO","JKO tolerates errors without losing limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1352,"prompt_tokens":961,"completion_tokens":391,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":335}},"tokens_in":577,"tokens_out":391,"duration_ms":4064,"temperature":1.0,"reasoning_tokens":335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:45:02.898118+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether the discrete EVI (2.5), $W_2^2(J_\\tau(\\mu),\\nu) \\le W_2^2(\\mu,\\nu) - 2\\tau(G(J_\\tau(\\mu))-G(\\nu)) - W_2^2(J_\\tau(\\mu),\\mu)$, holds for a proper lower semicontinuous geodesically convex $G$ when the input $\\mu$ is a two-atomic measure outside $D(G)$; a single failure would break the proof of Theorem 3.3 for functionals whose domain excludes such measures.","supporting_citations":[{"cited_title":"Lectures in Mathematics ETH Z¨ urich","cited_arxiv_id":null,"evidence_quote":"Supplies the classical JKO well-posedness and the base-point-in-domain EVI that Theorem 2.13 extends to all inputs."},{"cited_title":"Weak topology and opial property in Wasserstein spaces, with applications to gradient flows and proximal point algorithms of geodesically convex functionals","cited_arxiv_id":null,"evidence_quote":"Provides the weak topology $\\tau_{w,2}$, sequential compactness, and Opial property used to upgrade cluster points to full convergence."},{"cited_title":"The Wasserstein proximal gradient algorithm","cited_arxiv_id":null,"evidence_quote":"Establishes the exact Wasserstein proximal-gradient algorithm and the EVI whose refined form Lemma 4.1 derives."},{"cited_title":"Forward- backward Gaussian variational inference via JKO in the Bures–Wasserstein space, 2023","cited_arxiv_id":null,"evidence_quote":"Gives the weak-convergence result in the Bures–Wasserstein space that Theorem 4.6 generalizes."},{"cited_title":"An augmented la- grangian approach to Wasserstein gradient flows and applications","cited_arxiv_id":null,"evidence_quote":"Provides the ALG2/ADMM numerical method that motivates the distance-type and variational-type inexactness analysis."},{"cited_title":"Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem","cited_arxiv_id":null,"evidence_quote":"Interprets Langevin sampling as a composite optimization problem, motivating the inexact proximal-gradient analysis with variable stepsizes."}],"review_version":1}