{"id":"22a7977f-1ef1-4ebf-9bd1-eb98dbc7da56","arxiv_id":"2506.23429","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A simple two-term loss whose minimizer is the Monge map, with a stability bound showing the learned map converges to the optimal transport map as the loss gap shrinks.","lead":"A new training objective combines a transport cost with the Wasserstein distance to the target, and its global minimizer is the exact optimal transport map. The method trains with a stable min-min scheme and comes with a quantitative error bound, making it a simple architecture-agnostic alternative to adversarial or convex-network OT solvers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.6's O(√ε) map-error bound rests on an unproved uniform strong-convexity/C² regularity estimate for the Kantorovich potentials ψϵ; without it, Proposition A.5 cannot be invoked and the quantitative guarantee is not established.","rationale":"The reader's weakest assumption correctly identifies the uniform strong convexity of the Kantorovich potentials as the key unsupported ingredient. My stress-test sharpens this: even if every individual ψϵ is smooth and strongly convex, the proof requires the convexity modulus and the C² bounds to be uniform in ϵ. Proposition A.5 cannot be applied without fixed constants K independent of ϵ, and the paper's two references to this uniformity — the sentence after Lemma 3.3 and Remark 3.7 — are assertions, not derivations. The regulariy theory cited in Remark 3.7 gives smoothness and bounds only for fixed ϵ; it does not automatically give constants uniform in ϵ unless additional hypotheses on the sequence Tϵ (or νϵ) control derivatives. Since (3.19) is the only bridge from the dual-gap estimate to the L2 map error, the title's 'convergence guarantee' is conditional as written. That said, the consistency proof is elementary and correct, and the weak convergence argument uses standard stability of OT plans, so the central idea is plausible and the missing regularity argument is likely repairable under strengthened assumptions. I would therefore keep the reader's conditional verdict rather than moving to acceptance or rejection.","tokens_in":25826,"tokens_out":8713,"duration_ms":93827,"concrete_test":"Independently prove the claim in Lemma 3.3 and Remark 3.7: under Assumption 3.2 plus supp(ν) ⊆ supp(νϵ), show that there exist m > 0 and M < ∞ independent of ϵ such that mI ⪯ D²ψϵ ⪯ MI on supp(µ), where ψϵ is the Kantorovich potential for (µ, νϵ). Work through the argument in [17] (Gigli 2011) step by step and track the dependence of the convexity modulus on the constants in Assumption 3.2. If a uniform upper bound on D²ψϵ (or on ∇ψϵ) is needed and cannot be obtained from the stated assumptions, then the proof of Theorem 3.6 has a real gap. Alternatively, exhibit a family νϵ satisfying Assumption 3.2 for which the potentials ψϵ have D²ψϵ unbounded in ϵ while det(D²ψϵ) remains bounded, which would disprove the asserted uniform strong convexity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The decisive step is the passage from the dual-gap estimate (3.18) to the map error (3.19). This invokes Proposition A.5, which requires ψϵ to lie in a class X(2K) with K⁻¹I ⪯ D²ψϵ ⪯ KI, |∇ψϵ| ≤ K, and |ψϵ| ≤ 2K² on a fixed convex domain, with K independent of ϵ. Assumption 3.2 gives C² uniformly convex supports and uniformly bounded densities, which yield interior C^{2,α} regularity and uniform bounds on det(D²ψϵ) through the Monge–Ampère equation. But in dimension n > 1, uniform lower and upper bounds on det(D²ψϵ) do not by themselves give a uniform lower bound on the smallest eigenvalue unless the largest eigenvalue is also uniformly bounded; that upper bound is not derived anywhere. The sentence after Lemma 3.3 ('uniformity of the supremum convex modulus of ψϵ ... derived from the upper bound of |det(D²ψϵ)| as the same argument in [17]') and Remark 3.7's appeal to Theorem 12.50 in [40] assert, rather than prove, ϵ-independent constants. Without a uniform m > 0 such that ψϵ − (m/2)|x|² is convex for all ϵ, the factor 1/M in Lemma 3.3 and the constant 8K in Proposition A.5 can both diverge as ϵ → 0, so the O(√ε) bound in (3.9) is not established. The consistency theorem (3.1) and the weak convergence theorem (3.5) are not affected by this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes DPOT, a neural-network method for computing Monge maps between continuous distributions by minimizing P(T) = λ(I_µ(T))^{1/2} + W2(T♯µ, ν) over maps T, with 0 < λ < 1. Section 3 proves that the unique minimizer of P is the exact optimal transport map (Theorem 3.1), establishes weak convergence of approximately optimal maps (Theorem 3.5), and claims a quantitative L2 error bound of order √ε in terms of the optimality gap ε (Theorem 3.6). Section 3.3 gives a discrete min-min training objective, including a conditional version, and Section 4 reports experiments on synthetic OT problems, inverse maps, a compartmental epidemic model, and image color transfer.","tokens_in":26218,"tokens_out":19275,"duration_ms":203704,"significance":"Theorem 3.1 is a clean and correct elementary result: the term λ(I_µ(T))^{1/2} removes the trivial identity-map degeneracy of the pure Wasserstein objective while preserving the Monge solution. If Theorem 3.6 were fully established, the paper would provide a useful primal-loss-gap to map-error guarantee for neural OT without architectural constraints, which is a valuable complement to dual-potential analyses. The numerical study is broad, spanning multiple architectures and real-world tasks, and the experiments are generally consistent with the qualitative message of the theory. However, the advertised quantitative convergence guarantee is not yet fully supported: the proof of Theorem 3.6 relies on an unproved uniform strong-convexity regularity estimate, and the discretized loss in Eq. (3.21) does not match the continuous objective as written. These issues need to be resolved before the central claim can be accepted.","major_comments":[{"comment":"The quantitative bound (3.9) is not established as stated. The passage from (3.18) to (3.19) invokes Proposition A.5, which requires the Kantorovich potentials ψ_ε to belong to the class X(2K) with a constant K independent of ε. Assumption 3.2 and the Monge–Ampère equation give uniform bounds on det(D²ψ_ε) and interior regularity, but in dimension n > 1 a uniform lower and upper bound on det(D²ψ_ε) does not by itself imply a uniform lower bound on the smallest eigenvalue of D²ψ_ε unless a uniform upper bound on D²ψ_ε (or on its trace) is also available; that upper bound is not proved anywhere. The sentence following Lemma 3.3 and Remark 3.7 assert this uniformity by citing [17] and [40, Theorem 12.50], but they do not supply the required ε-independent constants. Consequently Lemma 3.3, the use of Lemma 3.3 in Theorem 3.5, and Theorem 3.6 all rest on an unproved regularity hypothesis. The authors should either prove the uniform strong convexity of ψ_ε under Assumption 3.2, or add it explicitly as an assumption, and should also verify that the pair (µ, ν) and the potentials satisfy the hypotheses of Proposition A.5 with a common K.","section":"§3.3, Eq. (3.21)"},{"comment":"The discretized first term in (3.21) does not approximate I_µ(T) as defined in (3.1). The continuous functional uses I_µ(T) = ½∫|T(x) − x|²dµ(x), whose empirical counterpart is (2N)^{-1}∑_{i=1}^N |T_θ(x_i) − x_i|². As written, the first term is (2N)^{-1}∑_{i,j=1}^N |T_θ(x_i) − x_j|², which converges to ½∫∫|T(x) − x'|² dµ(x)dµ(x'), not to I_µ(T). Thus the objective actually minimized in the experiments is different from the loss analyzed in Theorems 3.1–3.6. If the double sum is a typo, it should be corrected; if not, the numerical validation does not test the theoretical objective, and the claimed empirical support for the convergence bound needs to be re-examined.","section":"Theorem 3.6, Eq. (3.15), Remark 3.7"},{"comment":"The regularity of ψ*_ε on supp(ν) used in Eq. (3.15) is not justified by the stated assumptions. The proof applies the first-order convexity inequality for ψ*_ε at points y ∈ supp(ν) and bounds ∇ψ*_ε(y) by a constant R(µ) using only the fact that ∇ψ*_ε pushes ν_ε forward to µ. This requires ψ*_ε to be differentiable and its gradient to be uniformly bounded on supp(ν) (not merely ν_ε-a.e.), which is not derived from Assumption 3.2. The assumption supp(ν) ⊆ supp(ν_ε) is explicit but is an extra condition on T_ε that is not implied by a small gap ε, and Remark 3.7's proposed relaxation is only a sketch: the cited standard extension and [40, Theorem 12.50] are not shown to produce ε-independent constants. This affects the bound on the term I and hence the estimate (3.18).","section":null}],"minor_comments":[{"comment":"The statement of Proposition A.5 is not self-contained: it defines the class M(K) but never specifies the hypotheses on the target measure Q, even though the conclusion refers to an optimal transport map from P to Q. Please either quote the full statement from [23] or state the needed assumptions explicitly.","section":"Appendix A.1, Proposition A.5"},{"comment":"The displayed inequality bounds M(ν) − M(ν_ε) from above, but the term II in the decomposition contains M(ν_ε) − M(ν). The needed upper bound follows by the same argument with the roles of ν and ν_ε swapped, but this should be stated to make the proof complete.","section":"Eq. (3.16)"},{"comment":"The text says that the L2 relative error 'correlates well' with the optimality gap ε, but no quantitative fit is reported. Since Theorem 3.6 predicts a √ε rate, a log-log plot with a fitted slope or a table of exponents would make the empirical validation more convincing.","section":"Section 4, Figures 3a and 3b"},{"comment":"The sentence 'This optimization scheme maintains convergence while significantly lowering computational overhead' is not supported by any theoretical or experimental analysis of the alternating update of γ and θ; please soften or provide evidence for this claim.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The consistency theorem and weak convergence arguments are sound in spirit, and the proposed objective is attractive. The main risk is the uniform-regularity gap in Theorem 3.6, which may require either a substantial proof or a significantly strengthened assumption. In addition, the discrepancy between the continuous loss (3.1) and the discrete loss (3.21) must be resolved; if the double sum is a typo, correcting it will make the numerical section align with the theory. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: the paper introduces a two-term loss for Monge map estimation, P(T)=λ(I_µ(T))^{1/2}+W2(T♯µ,ν), and proves that for any λ∈(0,1) the unique minimizer is the true Monge map. The consistency proof is a clean triangle-inequality argument and is genuinely new. DeepParticle, the authors' own baseline, is the λ=0 special case; the added cost term is what forces a Monge map. That construction alone is worth a look from anyone doing neural OT.\n\nThe weak convergence theorem (Thm 3.5) is standard but competently assembled. The quantitative bound (Thm 3.6) is the right kind of stability result: small gap in P implies closeness to the OT map in L2(µ). That theorem has a real gap. The proof needs the Kantorovich potentials ψ_ϵ to be uniformly strongly convex, with modulus independent of ϵ; the paper asserts this via citation to [17] and Villani rather than proving it. The stress-test note is accurate: uniform bounds on det(D²ψ_ϵ) don't control the smallest eigenvalue in n>1 without also bounding the largest one. Remark 3.7 gestures at a fix, but the ϵ-independence of constants is not established. Theorem 3.6 is not proven as written. The consistency and weak convergence results do not depend on this, so it's a fixable gap, but a referee should ask for the full regularity argument.\n\nTwo lesser issues. The title claims a \"convergence guarantee,\" but the result is a deterministic stability bound conditional on a small loss gap; there is no sample-complexity or optimization guarantee for the min-min procedure. That's an overstatement but not a fatal one. The experiments are reasonable — they confirm the λ dependence, show gap vs. error correlation, and include a real posterior sampling task — but they lack comparisons to other neural OT methods on the synthetic problems where ground truth is known, so the practical claims are less convincing than the theory.\n\nThe citation pattern is clean; self-citation is limited to the λ=0 baseline and is not load-bearing. Code is released for a representative example, with the rest on request.\n\nThis paper deserves a serious referee. The construction is simple and likely to become a baseline, and the consistency proof should be in the literature. The referee's main job is to determine whether the uniform strong convexity can be proved under Assumption 3.2 or the assumption needs strengthening. I'd engage with it.","headline":"A simple, genuinely new loss for Monge map estimation with a clean consistency proof; the quantitative stability bound has a fixable but real regularity gap.","tokens_in":26736,"tokens_out":4007,"would_cite":true,"duration_ms":42165,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that a map-valued loss has a unique minimizer, and that minimizer is exactly the optimal transport map.","keywords":["optimal transport","Monge map","Wasserstein distance","min-min optimization","neural transport map","convergence guarantee","conditional transport","generative modeling"],"falsifier":"Construct a family of smooth, uniformly convex densities satisfying Assumption 3.2 with $\\operatorname{supp}(\\nu)\\subseteq\\operatorname{supp}(\\nu_\\epsilon)$ and bounded $|\\det(D^2\\psi_\\epsilon)|$, and compute the minimal eigenvalue of $D^2\\psi_\\epsilon$ as $\\nu_\\epsilon\\to\\nu$; if that eigenvalue tends to zero while the determinant bound stays finite, the asserted uniformity in Lemma 3.3 fails and the $\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2}$ bound is not established by the given proof.","tokens_in":1726,"feed_emoji":"🎯","tokens_out":9280,"duration_ms":193849,"temperature":0.7,"pith_summary":"This paper proposes a way to learn the optimal transport (Monge) map between two continuous distributions from unpaired samples, using a loss that combines the map's own transport cost with the Wasserstein-2 distance between the pushed-forward distribution and the target. The central theoretical claim is that this loss has a unique minimizer and that minimizer is exactly the Monge map, so training never needs adversarial or input-convex architectures. It further proves that a small gap in the loss implies a small $L^2$ error to the true map, giving a quantitative certificate of accuracy. The authors demonstrate the method on synthetic maps, conditional families, inverse maps, posterior sampling, and image color transfer. A sympathetic reader would care because this turns optimal-transport-map learning into a stable min-min optimization with a direct convergence guarantee.","feed_headline":"One loss provably recovers the exact optimal transport map","feed_subtitle":"A small gap in the training loss certifies a small error to the Monge map, with no adversarial training.","key_machinery":"The load-bearing object is the functional $P(T)=\\lambda\\,(I_\\mu(T))^{1/2}+W_2(T_\\sharp\\mu,\\nu)$ with $0<\\lambda<1$, where $I_\\mu(T)$ is the quadratic transport cost of $T$ and $W_2$ is the Wasserstein-2 distance. The consistency proof runs on one inequality chain: $P(T)\\ge\\lambda W_2(T_\\sharp\\mu,\\mu)+W_2(T_\\sharp\\mu,\\nu)\\ge\\lambda W_2(\\mu,\\nu)$, with equality only if $T$ is the Monge map. The quantitative proof decomposes the optimality gap $\\epsilon$ into $\\epsilon_1+\\epsilon_2+\\epsilon_3$ and bounds $\\|T_\\epsilon-\\bar T\\|_{L^2(\\mu)}$ by the sum of a perturbative estimate and a potential-comparison estimate, yielding $\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2}$.","core_discovery":"The paper claims that a single map-valued functional, $\\lambda$ times the transport cost of $T$ plus the Wasserstein-2 distance from $T_\\sharp\\mu$ to $\\nu$, has a unique minimizer and that minimizer is exactly the Monge map between $\\mu$ and $\\nu$. Moreover, if a learned map $T_\\epsilon$ has gap $\\epsilon$ in this functional, its $L^2(\\mu)$ distance to the true map is $O(\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2})$, so a small training loss is a certificate of near-optimality. The training objective is min-min rather than adversarial, and it imposes no convexity or Lipschitz constraints on the neural network. The paper also extends the loss to conditional families of maps and to inverse maps, and validates the claims numerically.","pith_inferences":["The paper does not propose using the loss gap as a training diagnostic, but Theorem 3.6 makes that a direct corollary: plotting $\\epsilon$ during training gives an upper-bound certificate for the map error, so one could use it for early stopping or model selection.","The quantitative proof's dependence on uniform strong convexity suggests the $\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2}$ rate is tied to smooth, uniformly convex geometry; testing the method on a target with non-convex support, such as an annulus, would show whether the rate survives outside those assumptions even though weak convergence likely does.","The consistency argument is specific to the quadratic cost and the Wasserstein-2 triangle inequality; extending the same loss to other costs would require a new lower bound relating the map's cost to the distance between $T_\\sharp\\mu$ and $\\mu$, so the min-min scheme should not be expected to transfer unchanged."],"forward_implications":["Any standard neural architecture can be used for optimal-transport-map estimation, because the loss itself enforces push-forward correctness without Lipschitz or convexity constraints.","The loss gap $\\epsilon$ functions as a training certificate, so reaching a small gap controls the $L^2(\\mu)$ error by $\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2}$.","Conditional maps $T(\\cdot|\\kappa)$ can be trained once and evaluated at unseen parameters $\\kappa$, allowing the method to scale to families of related transport problems without retraining.","With added cycle-consistency terms, the same loss learns forward and inverse maps, enabling bidirectional transport in applications such as color transfer.","Because inference is a single forward pass with no entropic smoothing or iterative sampling, the trained map can generate large numbers of samples quickly, as demonstrated on posterior sampling."],"supporting_citations":[{"why":"It supplies the existence and uniqueness of the optimal map and its representation as the gradient of a convex function, which Theorem 3.1 takes as its target.","marker":"[40]"},{"why":"It supplies the weak-convergence result for transported measures and the dual formulation that the quantitative proof relies on.","marker":"[41]"},{"why":"It supplies the perturbative $L^2$ estimate used in Lemma 3.3 and the cited argument for uniform strong convexity of the convex potentials.","marker":"[17]"},{"why":"It supplies the comparison bound between the duality gap and the $L^2$ difference of transport maps, used in Theorem 3.6.","marker":"[23]"},{"why":"It defines the DeepParticle objective that the new loss extends by adding the $\\lambda$-weighted transport-cost term.","marker":"[42]"}],"fun_headline_variants":["Training loss gap bounds OT map error","Min-min loss provably yields optimal transport","One loss certifies near-optimal Monge map","No adversarial training, provable OT map","Convergent DPOT: loss gap controls map error"],"cache_read_input_tokens":28672,"weakest_assumption_plain":"The quantitative error bound rests on assuming the source and target supports are smooth, uniformly convex, with densities bounded above and below, and on the asserted but not fully derived claim that the convex potentials stay uniformly strongly convex as the approximation improves; if these regularity conditions fail, only the weak convergence guarantee remains.","fun_headline_variants_meta":{"raw":{"variants":["Training loss gap bounds OT map error","Min-min loss provably yields optimal transport","One loss certifies near-optimal Monge map","No adversarial training, provable OT map","Convergent DPOT: loss gap controls map error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1363,"prompt_tokens":790,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":406,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":406,"tokens_out":573,"duration_ms":6155,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:44:33.918277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a family of smooth, uniformly convex densities satisfying Assumption 3.2 with $\\operatorname{supp}(\\nu)\\subseteq\\operatorname{supp}(\\nu_\\epsilon)$ and bounded $|\\det(D^2\\psi_\\epsilon)|$, and compute the minimal eigenvalue of $D^2\\psi_\\epsilon$ as $\\nu_\\epsilon\\to\\nu$; if that eigenvalue tends to zero while the determinant bound stays finite, the asserted uniformity in Lemma 3.3 fails and the $\\sqrt{\\epsilon_1}+\\sqrt{\\epsilon_2}$ bound is not established by the given proof.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the existence and uniqueness of the optimal map and its representation as the gradient of a convex function, which Theorem 3.1 takes as its target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the weak-convergence result for transported measures and the dual formulation that the quantitative proof relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the perturbative $L^2$ estimate used in Lemma 3.3 and the cited argument for uniform strong convexity of the convex potentials."},{"cited_title":"The Annals of Statistics 49(2) (2021)","cited_arxiv_id":null,"evidence_quote":"It supplies the comparison bound between the duality gap and the $L^2$ difference of transport maps, used in Theorem 3.6."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the DeepParticle objective that the new loss extends by adding the $\\lambda$-weighted transport-cost term."}],"review_version":1}