{"id":"89efe549-50f5-40d4-950d-99964c906396","arxiv_id":"2607.23008","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Nesterov and heavy-ball acceleration are extended to probability-measure optimization through dual lifting, with claimed non-asymptotic rates matching the Euclidean case.","lead":"This paper develops momentum-accelerated optimization algorithms for probability distributions, using a two-step 'lifting' so classical Nesterov convergence arguments apply. It claims Euclidean-matching rates for the measure-space algorithms and quantifies how particle approximations perturb those rates.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prop 2.3's m>0 equivalence is internally inconsistent: m-strong convexity of the L² lift plus invariance under rearrangements forces a Dirac minimizer, so Theorem 3.2(i)'s proof does not cover non-Dirac strongly geodesically convex targets.","rationale":"The paper has a clear architecture and the numerical experiments are honest about heavy-ball's poor performance. However, the central convergence claims rest on the L²-lifting equivalence in Proposition 2.3. The reader identified this as unproved; we find an additional internal consistency problem in the m>0 case, which is more decisive. For m>0, the lifted functional eF is automatically invariant under all measure-preserving transformations of the reference space. If eF is m-strongly convex, it has a unique minimizer; invariance then forces that minimizer to be fixed by all such transformations, so it must be constant a.e. Consequently, ρ⋆ must be a Dirac measure. This means Proposition 2.3 cannot hold for any m-strongly geodesically convex functional with a non-Dirac minimizer. Unless the paper explicitly restricts its strong-convexity results to Dirac minimizers—which would substantially narrow their scope—the proof of Theorem 3.2(i) is invalid as written. For m=0, the direction from geodesic convexity to lifted convexity also lacks a proof: geodesic convexity only controls optimal couplings, whereas arbitrary lifts X,Y need not be optimally coupled, so the linear interpolation (1−t)X+tY need not be a Wasserstein geodesic. The manuscript does not supply the missing argument. The particle-discretization bound in Theorem 3.3 is also not combined with the iteration error, but that is a separate limitation; the failure of Proposition 2.3 is more fundamental because it is the first step of the entire proof. We are not questioning the authors' integrity, only the logical completeness of the argument. The concrete test above would settle whether the m>0 equivalence is tenable or must be restricted. Until that is resolved, the reader's REJECT verdict stands.","tokens_in":19756,"tokens_out":37708,"duration_ms":364518,"concrete_test":"Analytical verification: Let F satisfy A1 and be m-strongly geodesically convex, and let X⋆ be a minimizer of eF. For any measure-preserving φ of (Ω,ν), eF(X⋆∘φ)=eF(X⋆); m-strong convexity then forces X⋆=X⋆∘φ a.e. Prove this implies X⋆ is constant a.e., hence ρ⋆ is a Dirac. Next, instantiate this with the paper's regularized KL functional (4.1) for a strongly log-concave target: if F_ε is m-strongly geodesically convex with non-Dirac minimizer (the Bayesian sampling regime), Proposition 2.3 is false; if it is not m-strongly convex, state that restriction and amend Theorem 3.2(i) so it no longer claims the equivalence for arbitrary non-Dirac cases.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central rates in Theorems 3.1–3.2 are derived by lifting to L²(ν) and invoking Proposition 2.3: F is m-geodesically convex iff eF(X):=F[X#ν] is m-convex on L²(ν). This proposition is not proved here; it is cited to Gangbo–Mészáros [2022]. More than a missing proof, the m>0 direction appears internally inconsistent. eF is invariant under every measure-preserving transformation φ of (Ω,ν): eF(X∘φ)=eF(X). If eF is m-strongly convex, its minimizer is unique. Hence any minimizer X⋆ must satisfy X⋆∘φ=X⋆ a.e. for all φ, which forces X⋆ to be constant a.e. and ρ⋆=X⋆#ν to be a Dirac mass. Thus Proposition 2.3 with m>0 can hold only for functionals whose minimizer is a Dirac. The paper applies Theorem 3.2(i) to strongly convex potentials (Example 1), whose minimizers are Dirac, so the proof may survive there; but the theorem as stated for arbitrary m-strongly geodesically convex F is either vacuous or false for non-Dirac minimizers. This is a concrete obstruction, not merely an absence of a citation. In addition, the 'geodesic convex ⇒ lifted convex' direction is non-obvious because geodesic convexity only controls optimal couplings, while arbitrary lifts need not be optimal couplings; the manuscript provides no argument covering that gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes heavy-ball and Nesterov-type accelerated algorithms for minimizing functionals over the Wasserstein space P2(R^d). Two lifts are used: a phase-space lift adding velocity marginals, and an L2(ν) lift representing measures as pushforwards of a reference measure. The main theoretical results are Theorems 3.1–3.2, which claim that the continuous measure-level iterates converge at the Euclidean rates (1 + m/(16L))^{-k} (heavy-ball), (1 − √(m/L))^k (Nesterov, strongly convex), and O(1/k²) (Nesterov, convex). Theorem 3.3 bounds the particle approximation error by C e^{(3+L)t} N^{-1/d} for d > 4. The proofs reduce the measure-space update to a Euclidean Nesterov sequence in L2(ν) by invoking structural equivalences between geodesic convexity/smoothness of F and Hilbert-space convexity/smoothness of the lift eF (Propositions 2.3–2.4), cited to prior work. Numerical experiments on potential energies, regularized KL sampling, mean-field neural networks, and porous-media flow compare the proposed methods with Wasserstein gradient descent.","tokens_in":20222,"tokens_out":19093,"duration_ms":192198,"significance":"Should the results be correct, the paper would provide a conceptually clean template for momentum methods on P2 and the first end-to-end non-asymptotic guarantee tracking both iteration and particle number. The phase-space/L2 dual lift is attractive, and the finite-particle tracking argument in Theorem 3.3 is a simple and elegant Lipschitz-stability argument. The paper also presents numerical comparisons across four problems. However, the central structural Proposition 2.3 is not proved and is false for m > 0 as stated; since the convergence theorems rely on it, the advertised generality of the rates is not established. The contribution is therefore presently a heuristic algorithm with encouraging numerics, not a proven acceleration theory.","major_comments":[{"comment":"Proposition 2.3 (m > 0 direction) is internally inconsistent. For every measure-preserving map φ of (Ω,ν), eF(X∘φ)=eF(X). If eF were m-strongly convex on L2(ν), its minimizer X⋆ would be unique; invariance would force X⋆∘φ=X⋆ a.e. for all φ, hence X⋆ constant a.e. and ρ⋆=X⋆#ν a Dirac. Thus Prop. 2.3 can hold for m > 0 only when the minimizer of F is a Dirac. This contradicts the stated scope of Theorem 3.2(i) and excludes the regularized KL functional (4.1) used in Example 2, which has a non-Dirac minimizer. The proof at Eq. (3.24) therefore does not establish the claimed rate for general m-strongly geodesically convex F.","section":"§2.3 (Prop. 2.3); §3.2.1, proof of Thm. 3.2(i)"},{"comment":"The m = 0 direction of Prop. 2.3 is also unproved. Definition 2.4 gives the convexity inequality only along optimal couplings, while (2.27) must hold for arbitrary X,Y ∈ L2(ν); the linear interpolation (1−t)X + tY is not an optimal coupling in general. No argument or precise quoted theorem is provided for this implication. This gap is independent of the m > 0 obstruction.","section":"§2.3 (Prop. 2.3, m = 0 direction)"}],"minor_comments":[{"comment":"The constant C is said to depend on dimension only, but the proof via Fournier–Guillin also depends on M3(µ0) and on the growth of the initial measure; these dependencies should be stated explicitly.","section":"Thm. 3.3, Eq. (3.26)"},{"comment":"The numerical experiments do not validate the N^{-1/d} particle rate. In Example 1, N=100 in d=500 gives N^{-1/d}≈0.99, so the bound is vacuous there; a scaling plot in N would be needed.","section":"§4"},{"comment":"The notation M3(µ0) is used without definition. Also, the bound (L')^{t/√s} ≤ e^{(3+L)t} in Eq. (3.34) absorbs an e^{2√s(L+1)t} factor into the constant; this should be acknowledged or s should be fixed.","section":"Thm. 3.3"},{"comment":"All three structural propositions are cited to prior papers without proofs or precise theorem numbers. For a self-contained journal submission, the exact statements and proofs (or precise references) should be included, especially because the main theorems rest on them.","section":"§2.3, Props. 2.2–2.4"},{"comment":"The estimates m≈10^{-5} and L≈1 are asserted without derivation. Please specify how the strong-convexity and smoothness parameters are estimated in the experiments.","section":"§4, Example 1"}],"recommendation":"reject","confidential_remarks":"The paper has a promising algorithmic idea, but the theoretical core is not sound as written. The false Proposition 2.3 is not a local gap; it invalidates the main convergence theorems. A restricted version for potential-type functionals with Dirac minimizers might be salvageable, but that would require substantial reframing of the paper's claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-written and honest attempt to build discrete-time Nesterov acceleration over P_2 by lifting to L^2(ν). The algorithmic translation is clean, the numerical section is sober about heavy-ball's performance, and the m=0 case may well be correct. But the strong-convexity theorem (3.2(i)) rests on Proposition 2.3, which is not proved and, worse, looks internally inconsistent for m>0. That is a load-bearing flaw, not a missing citation.\n\nWhat's new: the explicit phase-space discretization plus Hilbert-space bridge, with non-asymptotic iteration rates and a particle-discretization bound. The reduction of the measure update to the Euclidean Nesterov sequence is elegant and dimension-free. If the lifting equivalences hold, the rates follow immediately. The authors also deserve credit for flagging that KL divergence needs regularization and that heavy-ball underperforms.\n\nThe soft spots: the main one is Proposition 2.3. As stated, it says F is m-geodesically convex iff eF is m-convex on L^2(ν) for m>0. But eF is invariant under measure-preserving rearrangements, so m-strong convexity forces a unique minimizer that must be fixed by every rearrangement; hence the minimizer is constant and the target measure is Dirac. For any non-Dirac strongly geodesically convex functional (e.g., relative entropy with a smooth target), the equivalence is false. So Theorem 3.2(i) as stated is either vacuous or restricted to Dirac minimizers. The proof also needs the 'geodesic convex ⇒ lifted convex' direction, which is non-obvious since geodesic convexity only controls optimal couplings; no argument is given. The particle bound (Theorem 3.3) is exponential in time, requires d>4, and is not combined with the optimization error, so the advertised end-to-end rate is not actually delivered. And there is no code or error bars in the experiments.\n\nBottom line: the paper is worth a serious referee, but the strong-convexity claim needs to be either proved under the right structural assumptions (or restricted to Dirac-target functionals) or corrected. The m=0 and potential-energy examples may survive. If I were the editor, I'd send it to review with the clear request that the authors prove or precisely restate the lifting lemmas and de-emphasize the overbroad theorem.","headline":"Elegant dual-lifting framework for accelerated Wasserstein optimization, but the strong-convexity theorem leans on a lifting equivalence that appears false as stated for non-Dirac minimizers.","tokens_in":20639,"tokens_out":3704,"would_cite":false,"duration_ms":39445,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","65K10","90C25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that Nesterov's accelerated gradient method can be transplanted from Euclidean space to optimization over probability measures—preserving its classic rates—via two complementary lifts: a phase-space momentum lift and a Hil","keywords":["Nesterov acceleration","probability measures","Wasserstein space","phase-space lifting","Hilbert-space lifting","particle methods","convergence rates","geodesic convexity"],"falsifier":"Run the particle Nesterov scheme (3.13) on a smooth, strongly geodesically convex functional in dimension d=5, with N ranging from 10^3 to 10^5. If the expected W2 error at a fixed time t does not decay roughly like N^{−1/5}, Theorem 3.3 is wrong. Also, construct a functional where the cited geodesic-convexity-to-L²-convexity equivalence can be tested directly; a counterexample would collapse the proof of Theorem 3.2.","tokens_in":19623,"feed_emoji":"🚀","tokens_out":4610,"duration_ms":46222,"temperature":0.7,"pith_summary":"The paper tries to establish that Nesterov-style momentum acceleration works on the nonlinear space of probability measures, at the same speed as in Euclidean space. It does so by lifting the problem twice: first into a phase space of positions and velocities, then into a flat Hilbert space of random variables. If correct, the paper's rates mean that accelerated optimization over distributions—important for Bayesian sampling, mean-field neural networks, and related tasks—does not lose its iteration complexity to the curvature of the Wasserstein geometry. The finite-particle version is also shown to track the ideal trajectory with an explicit, dimension-cursed error bound.","feed_headline":"Measure-space Nesterov matches Euclidean convergence rates","feed_subtitle":"A dual lifting trick carries accelerated gradient methods onto Wasserstein space, preserving rates and bounding particle error.","key_machinery":"The central device is a double lift. First, a phase-space lift inserts velocity variables so momentum is encoded in a joint position–velocity measure. Second, a Hilbert-space lift represents any probability measure ρ as X#ν for a random variable X ∈ L²(ν), pulling the functional back to eF(X)=F[X#ν]. This restores a single flat geometry in which the Wasserstein gradient becomes an ordinary Fréchet gradient, geodesic convexity becomes m-convexity, and gradient smoothness becomes L-Lipschitz continuity. The lifted Nesterov updates are then two linear operators S and R acting on (X,V), whose pushforwards reproduce the measure iteration; particle implementations apply the same operators to empir","core_discovery":"The paper claims that the phase-space Nesterov iteration over probability measures (3.12), when analyzed through the L²(ν) lift, converges at exactly the Euclidean Nesterov rates: linear convergence F[ρ_k]−m_F ≤ O((1−√(m/L))^k) for m-strongly geodesically convex functionals and O(1/k²) for geodesically convex ones, under a global Wasserstein-gradient regularity condition. The finite-particle version (3.13), initialized by N independent samples, is shown to track the continuous trajectory with expected Wasserstein error bounded by C e^{(3+L)t} N^{−1/d} for d>4, so iteration error and Monte Carlo error add without destroying the acceleration. The whole argument reduces the nonlinear measure-sp","pith_inferences":["Editorial inference: the N^{−1/d} particle error means that for high-dimensional problems the sampling error will quickly dominate the iteration error; structured or quasi-Monte Carlo sampling could be tested as a way to soften the curse of dimension.","Editorial inference: the same two-step lifting recipe might transplant other Euclidean accelerated schemes—for example accelerated proximal methods—to measure spaces, provided convexity and smoothness are preserved by the lift.","Editorial inference: the numerical experiments include settings that only satisfy the paper's assumptions locally or approximately, suggesting the rates may hold beyond the stated global geodesic-convexity hypothesis; a dedicated local-convexity analysis would be a natural extension."],"forward_implications":["In the strongly geodesically convex case, the continuous phase-space Nesterov iterates satisfy F[ρ_k]−m_F = O((1−√(m/L))^k), the Euclidean Nesterov rate, not the slower gradient-descent rate.","In the merely geodesically convex case, the same scheme achieves O(1/k²), matching Nesterov's convex rate.","The interacting-particle implementation tracks the ideal measure trajectory: expected W2 error after time t is at most C e^{(3+L)t} N^{−1/d} for d>4, so iteration and sampling errors decouple.","Heavy-ball momentum, by contrast, only reaches the slower O((1+m/(16L))^{-k}) rate, reproducing the Euclidean gap.","Because the lifted sequence is exactly the classical Nesterov sequence in Hilbert space, all dimension-free Euclidean proof techniques transfer automatically to the measure setting."],"fun_headline_variants":["Nesterov on probability measures closes Euclidean gap","Dual lift gives Wasserstein Nesterov matching convergence rates","Phase-space lift brings O(1/k²) Nesterov to measure space","Particle-tracking Nesterov: measure-space rates match Euclidean","Lifting to Hilbert space unlocks Nesterov for Wasserstein"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is a cited equivalence—geodesic convexity of the functional on probability space is equivalent to ordinary convexity of its lifted version—and it is imported, not proven here; if that equivalence fails or needs extra conditions, the main rates do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Nesterov on probability measures closes Euclidean gap","Dual lift gives Wasserstein Nesterov matching convergence rates","Phase-space lift brings O(1/k²) Nesterov to measure space","Particle-tracking Nesterov: measure-space rates match Euclidean","Lifting to Hilbert space unlocks Nesterov for Wasserstein"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000968,"raw_usage":{"total_tokens":3951,"prompt_tokens":737,"completion_tokens":3214,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":3126}},"tokens_in":481,"tokens_out":3214,"duration_ms":24169,"temperature":1.0,"reasoning_tokens":3126,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:54:09.305761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the particle Nesterov scheme (3.13) on a smooth, strongly geodesically convex functional in dimension d=5, with N ranging from 10^3 to 10^5. If the expected W2 error at a fixed time t does not decay roughly like N^{−1/5}, Theorem 3.3 is wrong. Also, construct a functional where the cited geodesic-convexity-to-L²-convexity equivalence can be tested directly; a counterexample would collapse the proof of Theorem 3.2.","supporting_citations":[],"review_version":1}