{"id":"b09c2b9c-be5a-4cc4-a5f1-4a7bff5ca8a6","arxiv_id":"2411.15067","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Linear convergence of JKO-based proximal point, prox-linear, and proximal gradient schemes is proved for entropy-regularized flat-convex functionals, with iterates shown to have finite relative Fisher information.","lead":"The paper proves linear convergence for three step-by-step (proximal) optimization schemes on the space of probability distributions, for entropy-regularized objectives that are flat-convex but not geodesically convex. This gives new theoretical guarantees for mean-field machine learning problems such as training two-layer neural networks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform flat-derivative boundedness (Assumption 2.2(2.3)) is load-bearing: without it the Holley–Stroock LSI constant can degenerate and the linear-rate proof has no positive κ.","rationale":"The reader's weakest-assumption identification matches my own: Assumption 2.2(2.3) is the condition on which the uniform LSI (2.7) and hence the positive rate κ in all three schemes depends. The concern is not that the theorem is false under its stated hypotheses; rather, it delineates the actual scope of the result, which is narrower than 'under flat convexity' alone. The separate issue flagged by the reader — Lemma 6.6 being applied to μ0 which is not shown to lie in C — is real but minor and does not threaten the qualitative linear-convergence claim. Consequently, the reader's CONDITIONAL verdict remains appropriate: the main proof is convincing under the stated assumptions, but a revision should either relax or prominently justify (2.3), and should fix the Lemma 6.6 initialization gap. I do not see grounds for REJECT or for ACCEPT without revision, so UNCHANGED is the honest recommendation.","tokens_in":31033,"tokens_out":16055,"duration_ms":131943,"concrete_test":"Check analytically whether [8, Thm 2.7(2)] can be applied to Φ[μ] under only the uniform Lipschitz condition (2.2), without (2.3), yielding a constant θ>0 independent of μ. Concretely, take d=1, π∝e^{-x²/2}, σ=1, and F(μ)=∫x² dμ (so δF/δμ is unbounded in x); compute or upper-bound the log-Sobolev constant of Φ[μ] ∝ e^{-x²/2 - (x² - ∫y²dμ)/σ}. If the constant can be chosen uniformly over μ with 0<θ<∞, then (2.3) is not necessary for the proof mechanism and Remark 2.6 should be promoted to a theorem; if no uniform θ exists, the stated Theorem 3.1 requires (2.3) and does not cover this F.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 3.1 hinges on the uniform LSI (2.7) for the proximal Gibbs measures Φ[μ], obtained in (2.7) via the Holley–Stroock criterion from the pointwise bound |δF/δμ(μ,x)| ≤ C_F. This bound is also used in Theorem 5.1 (Step 3) to pass to the limit by dominated convergence with integrand bounded by 2C_F. If (2.3) is dropped, the constant e^{-4C_F/σ} in the rates of Theorem 3.1 becomes 0, so the proof yields no positive κ. Natural flat-convex objectives such as F(μ)=∫V dμ with unbounded V (e.g., V(x)=|x|²) violate (2.3), and the paper's main example (Example 3.3) relies on a bounded activation and a clipping function to satisfy it. Remark 2.6 suggests replacing (2.3) by uniform Lipschitzness of x↦δF/δμ(μ,x), citing [8, Thm 2.7(2)], but Theorem 3.1 is not stated under that weaker condition, and no uniform-in-μ lower bound on the resulting LSI constant is verified there. Thus the advertised relaxation from geodesic convexity to flat convexity is accompanied by a strong uniform-boundedness condition that excludes common quadratic and other unbounded potentials; this is a scope limitation of the central claim rather than an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the entropy-regularized mean-field optimization problem min_μ F_σ(μ)=F(μ)+σ KL(μ|π) on the Wasserstein space and analyzes three discrete-time proximal/JKO-type schemes: the proximal point scheme (1.3), the prox-linear scheme (1.4), and the proximal gradient scheme (1.5). Under flat convexity of F (Assumption 2.1), Lipschitzness and uniform boundedness of the flat derivative (Assumption 2.2), strong convexity and growth conditions on the potential U of π (Assumption 2.3), and suitable Wasserstein smoothness assumptions (Assumptions 2.8–2.9), Theorem 3.1 establishes Q-linear convergence of the objective gap with explicit contraction rates; Corollary 3.2 transfers these rates to KL divergence and squared Wasserstein distance. The proof strategy combines a uniform logarithmic Sobolev inequality for the proximal Gibbs measures Φ[μ], obtained via Holley–Stroock, with the entropy sandwich lemma (Lemma 2.7), first-order optimality conditions, and a technical regularity result showing that the relative entropy admits a unique Wasserstein subgradient along the iterates.","tokens_in":31333,"tokens_out":21671,"duration_ms":190628,"significance":"If the result is correct, it is a useful discrete-time counterpart to the continuous-time LSI-based analysis of mean-field Langevin dynamics and extends JKO-type convergence theory beyond geodesically convex functionals. The paper's strengths are its rigor and completeness: existence and uniqueness of minimizers for each scheme are proved rather than assumed, the Wasserstein subdifferentiability of the relative entropy along the iterates is handled via a Sobolev-regularity class C, and all rates are given with explicit constants. The main caveat is that Assumption 2.2(2.3), the uniform bound on δF/δμ, is essential for the uniform LSI constant in (2.7) and also enters the proof of existence in Theorem 5.1; this condition excludes natural flat-convex objectives such as F(μ)=∫|x|^2 dμ(dx), so the advertised relaxation from geodesic convexity is narrower than the abstract suggests. This is a scope limitation rather than an internal inconsistency of the proof.","major_comments":[],"minor_comments":[{"comment":"Lemma 6.6 is stated only for μ′, μ ∈ C, but in the proof of Theorem 3.1(ii) it is applied with μ = μ0, and μ0 is assumed only to lie in P^λ_2. The proof of the lemma itself only needs μ ∈ P^λ_2 and μ′ ∈ C, so the statement should be relaxed accordingly; the sentence just before (4.4) claiming that all iterates (μn)n∈N lie in C is also inaccurate at n = 0.","section":"Lemma 6.6 and §4.2"},{"comment":"The remark that condition (2.3) can be replaced by uniform Lipschitzness of x ↦ δF/δμ(μ,x) is not sufficient as written, because (2.3) is also used to obtain the dominated convergence bound in Step 3 of the proof of Theorem 5.1. Either prove the full statement under the weaker assumption or restrict the remark to the LSI constant.","section":"Remark 2.6 and Theorem 5.1 Step 3"},{"comment":"The displayed formula for δF/δμ appears to omit the factor 2 from the derivative of the squared L2 loss and the normalization constant required by Definition B.1; please verify the computation and report the normalized derivative.","section":"Example 3.3"},{"comment":"The phrase 'relaxing the common reliance on geodesic convexity' should be qualified: Assumption 2.2(2.3), the uniform boundedness of the flat derivative, is essential for the uniform LSI constant in (2.7) and excludes basic flat-convex functionals such as F(μ)=∫|x|^2 dμ(dx). A sentence stating this scope would improve the paper.","section":"Abstract and Introduction"},{"comment":"There are minor typographical issues ('auxilliary' in Section 7, 'if identical' in Proposition 6.3, and the header of Theorem 3.1 reading '1.5'); these do not affect the mathematics.","section":"Various"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a mathematically solid contribution with detailed proofs and explicit rates. The main revisions I request are local: correct the statement and application of Lemma 6.6, and temper the claims about the relaxation of geodesic convexity to reflect the uniform boundedness assumption. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you care about mean-field optimization or JKO schemes. The genuinely new thing is Theorem 3.1: three proximal schemes—proximal point, prox-linear, proximal gradient—converge linearly in objective gap for entropy-regularized flat-convex functionals, with explicit rates and no geodesic convexity and no discrete EVI. They also handle the technical core that usually gets swept under the rug: showing each iterate lies in a Sobolev-regularity class so the relative entropy has a unique Wasserstein subgradient and the Fisher information is finite. That part is real work, and the existence/uniqueness proofs for each scheme fill gaps where earlier papers just postulated minimizers.\n\nThe proof strategy is the continuous-time LSI plus entropy sandwich from [27,12], adapted to discrete time. The adaptation is the contribution, not the machinery, and the authors are upfront about that. The rates are explicit and the constants trace through; I did not find a load-bearing error.\n\nThe soft spot is Assumption 2.2(2.3): the flat derivative δF/δμ(μ,x) is uniformly bounded in both arguments. That bound enters the Holley–Stroock step and gives the uniform LSI constant e^{-4C_F/σ}; without it the proof has no positive κ. Natural objectives like F(μ)=∫V dμ with V(x)=|x|² violate it, and the paper's own NN example needs a bounded activation and clipping to fit. Remark 2.6 points to a relaxation via uniform Lipschitzness, but Theorem 3.1 is not stated under that condition and no uniform LSI constant is verified there. So the advertised scope—flat convex rather than geodesic convex—still carries a uniform boundedness condition that excludes common unbounded potentials. That is a real limitation, but it is a scope limitation, not a broken proof.\n\nOne small technical gap: Lemma 6.6 is stated for μ',μ∈C but is applied in the proof of Theorem 3.1(ii) with μ=μ0, which is only assumed to be in P^λ_2. The proof seems to support the wider statement by approximation or density, but as written the stated hypotheses do not cover the application. That should be fixed in revision.\n\nVerdict: the paper earns a serious referee. It is a solid, technically demanding contribution with honest attribution. I would send it out; the revision should address the Lemma 6.6 hypothesis gap and ideally state the weaker uniform-Lipschitz version if the authors can verify it.","headline":"The first discrete-time linear convergence rates for JKO-type schemes under flat convexity; the uniform boundedness assumption on the flat derivative is the real price of admission.","tokens_in":31907,"tokens_out":1811,"would_cite":true,"duration_ms":18075,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["46N10","49Q22","49K30","58E30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that three proximal descent schemes on the Wasserstein space converge linearly to the minimizer of an entropy-regularized objective, under flat convexity rather than geodesic convexity.","keywords":["entropy regularization","proximal JKO schemes","Wasserstein space","mean-field optimization","logarithmic Sobolev inequality","flat convexity","optimal transport","linear convergence"],"falsifier":"In the exactly solvable Gaussian case with U quadratic and F = 0, the JKO iterates are explicit; compute the objective gap and compare the contraction ratio with the predicted kappa = 1 + tau sigma alpha_U. A single asymptotic contraction factor smaller than the predicted one would refute the claimed Q-linear rate of Theorem 3.1(i).","tokens_in":30846,"feed_emoji":"📉","tokens_out":5515,"duration_ms":51270,"temperature":0.7,"pith_summary":"The paper establishes linear convergence for three discrete-time proximal methods on the space of probability measures: a proximal point scheme, a prox-linear scheme, and a proximal gradient scheme. Each method targets the minimizer of an entropy-regularized objective F_sigma(mu) = F(mu) + sigma KL(mu|pi), and the authors show the objective gap contracts by a fixed factor every step. This relaxes the common requirement that F be geodesically convex in the Wasserstein space, replacing it with flat convexity, a bounded flat derivative, and a strongly convex reference potential. The proof avoids discrete-time Evolution Variational Inequalities and instead uses a uniform logarithmic Sobolev inequality for the proximal Gibbs measures together with an entropy sandwich lemma. The main technical challenge is proving that the relative entropy is Wasserstein subdifferentiable along the iterates, which is handled by showing the iterates stay in a Sobolev regularity class.","feed_headline":"Linear convergence proven for Wasserstein proximal descent","feed_subtitle":"Three JKO-based schemes reach the entropy-regularized minimizer at an exponential rate under flat convexity alone.","key_machinery":"The load-bearing object is the proximal Gibbs measure Phi[mu] proportional to exp(-$sigma^{{-1}}$ delta F/delta mu (mu, .) - U), which for every mu satisfies a logarithmic Sobolev inequality with uniform constant alpha_U $e^{{-4 C_F / sigma}}$, obtained by the Holley–Stroock criterion applied to the strongly log-concave reference pi proportional to $e^{{-U}}$. The entropy sandwich lemma (Lemma 2.7) bounds the objective gap F_sigma(mu) - F_sigma(mu*_sigma) between $\\sigma$ KL(mu|mu*_sigma) and $\\sigma$ KL(mu|Phi[mu]). The proof then shows each scheme's first-order optimality condition identifies the relative Fisher information I(mu_{n+1}|Phi[mu]) with a Wasserstein displacement, and the log-Sobolev inequality converts that Fisher information into KL, producing a contraction in the objective gap. A Sobolev regularity class C of densities is introduced so that the relative entropy KL(·|pi) has a unique Wasserstein subgradient, given by the log-density gradient, at every iterate, which makes the Fisher information finite and the optimality conditions valid.","core_discovery":"The central claim is Theorem 3.1: for each of the schemes (1.3), (1.4), and (1.5), under Assumptions 2.1–2.3 and 2.8 (plus Assumption 2.9 for the latter two, with small step-size bounds), there exists kappa > 1 such that 0 <= F_sigma(mu_n) - F_sigma(mu*_sigma) <= $kappa^{{-n}}$(F_sigma(mu_0) - F_sigma(mu*_sigma)) for all n. Corollary 3.2 transfers this exponential contraction to the relative entropy KL(mu_n|mu*_sigma) and the squared Wasserstein distance $W_2^{2}$(mu_n, mu*_sigma). The purpose is to show that geodesic convexity is not needed for linear rates in mean-field optimization: flat convexity plus a uniformly bounded flat derivative and a strongly log-concave reference measure suffice. The proof works by identifying, through first-order optimality conditions, the relative Fisher information with the squared Wasserstein displacement at each step, then using the uniform log-Sobolev inequality and the sandwich lemma to convert Fisher information into a one-step objective decrease.","pith_inferences":["The explicit contraction factors suggest the practical step-size constraints are real and quantitative: the rate degrades as the Wasserstein gradient becomes less smooth, and particle implementations would need to respect the same tau bounds.","Remark 2.6 points to a relaxation: if uniform Lipschitzness of x |-> delta F/delta mu (mu, x) replaces the boundedness of delta F/delta mu, the two-layer network example can drop its clipping function; this relaxation is left as a remark but is a testable route to broader applicability.","The linear rates are in the mean-field limit; the paper does not quantify how the contraction interacts with particle discretization error, which seems to be the natural next question for practical algorithms."],"forward_implications":["The objective gap F_sigma(mu_n) - F_sigma(mu*_sigma) contracts by a fixed factor at every step for all three schemes, so the methods are Q-linearly convergent in objective value.","The same exponential contraction transfers to KL(mu_n|mu*_sigma) and W_2^2(mu_n, mu*_sigma), with explicit constants in Corollary 3.2.","Geodesic convexity of F is not needed; flat convexity, bounded flat derivative, and a strongly convex reference potential are enough for linear rates.","The small step-size restrictions are explicit: tau < 2/L'_F for the prox-linear scheme and tau < 1/L'_F for the proximal gradient scheme.","The two-layer mean-field neural network with a clipped activation function satisfies the assumptions, so the rates apply to that training objective."],"supporting_citations":[{"why":"Introduces the Jordan–Kinderlehrer–Otto scheme whose existence proof is modified in Theorem 5.1 and whose JKO step underlies all three schemes.","marker":"[20]"},{"why":"Supplies the convex-analysis view of mean-field Langevin dynamics and the LSI-plus-sandwich strategy that the paper adapts to discrete time.","marker":"[27]"},{"why":"Develops the entropy sandwich lemma and the uniform log-Sobolev approach for continuous-time mean-field Langevin dynamics, which this paper extends.","marker":"[12]"},{"why":"The Holley–Stroock criterion converts bounded perturbations of a strongly log-concave reference into the uniform LSI constant for the proximal Gibbs measures.","marker":"[16]"},{"why":"The Bakry–Émery criterion yields the log-Sobolev inequality for pi from the strong convexity of U.","marker":"[3]"},{"why":"The Otto–Villani Talagrand inequality transfers KL contraction to squared Wasserstein distance in Corollary 3.2.","marker":"[28]"},{"why":"Provides the Wasserstein subdifferential calculus, optimal transport map existence theorems, and geodesic convexity of KL needed for the first-order optimality conditions.","marker":"[1]"},{"why":"Gives the Wasserstein differentiability of F from the flat derivative, used in Assumption 2.8 and the identification of the Wasserstein gradient.","marker":"[7]"}],"fun_headline_variants":["Flat convexity suffices for linear rates in Wasserstein descent","Proximal descent on Wasserstein space: linear convergence without geodesic convexity","Linear convergence for JKO schemes under flat convexity alone","Wasserstein proximal descent: linear rates with flat convexity","Relaxing geodesic convexity: linear convergence in mean-field optimization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The uniform bound |delta F/delta mu (mu, x)| <= C_F for all mu and x is what keeps the log-Sobolev constant of the proximal Gibbs measure from degenerating; if that bound fails, the proof has no positive constant to work with.","fun_headline_variants_meta":{"raw":{"variants":["Flat convexity suffices for linear rates in Wasserstein descent","Proximal descent on Wasserstein space: linear convergence without geodesic convexity","Linear convergence for JKO schemes under flat convexity alone","Wasserstein proximal descent: linear rates with flat convexity","Relaxing geodesic convexity: linear convergence in mean-field optimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1559,"prompt_tokens":967,"completion_tokens":592,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":501}},"tokens_in":583,"tokens_out":592,"duration_ms":5969,"temperature":1.0,"reasoning_tokens":501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:32:57.795884+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the exactly solvable Gaussian case with U quadratic and F = 0, the JKO iterates are explicit; compute the objective gap and compare the contraction ratio with the predicted kappa = 1 + tau sigma alpha_U. A single asymptotic contraction factor smaller than the predicted one would refute the claimed Q-linear rate of Theorem 3.1(i).","supporting_citations":[{"cited_title":"Jordan, D","cited_arxiv_id":null,"evidence_quote":"Introduces the Jordan–Kinderlehrer–Otto scheme whose existence proof is modified in Theorem 5.1 and whose JKO step underlies all three schemes."},{"cited_title":"Nitanda, D","cited_arxiv_id":null,"evidence_quote":"Supplies the convex-analysis view of mean-field Langevin dynamics and the LSI-plus-sandwich strategy that the paper adapts to discrete time."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Develops the entropy sandwich lemma and the uniform log-Sobolev approach for continuous-time mean-field Langevin dynamics, which this paper extends."},{"cited_title":"Holley and D","cited_arxiv_id":null,"evidence_quote":"The Holley–Stroock criterion converts bounded perturbations of a strongly log-concave reference into the uniform LSI constant for the proximal Gibbs measures."},{"cited_title":"Bakry and M","cited_arxiv_id":null,"evidence_quote":"The Bakry–Émery criterion yields the log-Sobolev inequality for pi from the strong convexity of U."},{"cited_title":"Otto and C","cited_arxiv_id":null,"evidence_quote":"The Otto–Villani Talagrand inequality transfers KL contraction to squared Wasserstein distance in Corollary 3.2."},{"cited_title":"Ambrosio, N","cited_arxiv_id":null,"evidence_quote":"Provides the Wasserstein subdifferential calculus, optimal transport map existence theorems, and geodesic convexity of KL needed for the first-order optimality conditions."}],"review_version":1}