{"id":"10f05015-676c-4983-861c-69dc64dc3faf","arxiv_id":"2412.15861","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Three robust variants of Gromov-Wasserstein (Tukey and Huber GW, locally robust GW, and a robust reversible Gromov-Monge distance) are introduced, with partial theoretical guarantees and empirical gains on contaminated shape and image alignment.","lead":"The paper proposes three ways to make Gromov-Wasserstein alignment between different data domains resistant to outliers: penalizing large geometry distortions, trimming the metric distances before comparing them, and regularizing the transport plan with denoising maps. It proves some metric and robustness properties for the new distances and reports improved shape matching and image translation under contamination.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 3's same-space restriction leaves the cross-domain robustness claim unproved; a triangle-inequality corollary would close the gap.","rationale":"The reader identified Proposition 3's same-space premise as the fragile load-bearing assumption, and I agree that the paper as written does not provide a population-level cross-domain robustness theorem. However, the gap is not deep: the triangle inequality for TGW (Proposition 2(i)) plus a same-space bound on dTGW(µ', µ) yields exactly the cross-domain statement needed, dTGW(µ', ν) − dGW(µ, ν) ≤ τ ε^{1/p}. Thus the central idea survives, but the manuscript should state and prove this corollary or explicitly qualify the abstract. Other issues, such as the overstatement in the contributions bullet for Proposition 11 (an upper bound on F2 being called 'calculating an OT') and the lack of error bars, are secondary and do not change the verdict. The reader's conditional verdict already requires these revisions, so I leave it unchanged.","tokens_in":38882,"tokens_out":17054,"duration_ms":146657,"concrete_test":"Analytical check: write the missing corollary in full, replacing Proposition 3's hypothesis by arbitrary Polish mm spaces (X, dX, µ) and (Y, dY, ν) with µ' = (1−ε)µ + εµ_c. Step 1: use Proposition 2(i) with three spaces (X, µ'), (X, µ), (Y, ν) to get dTGW(µ', ν) ≤ dTGW(µ', µ) + dTGW(µ, ν). Step 2: repeat the proof of Proposition 3 but only on X to show dTGW(µ', µ) ≤ W_Tp(µ', µ) ≤ τ ε^{1/p}. Step 3: combine with dTGW(µ, ν) ≤ dGW(µ, ν). If the derivation is valid, the same-space restriction is non-essential and the abstract's cross-domain claim is supported; if any step fails (e.g., the O∪I empirical version in Remark 7 requires a separate argument), the paper must add this theorem or qualify the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central population-level robustness guarantee, Proposition 3, is stated only for distributions µ and ν on the same mm space (X, dX): the bound dTGW(µ', ν) ≤ τ ε^{1/p} + W_Tp(µ, ν) uses W_Tp, an OT distance on X, and the proof (Appendix A, eqs. 33–35) applies the triangle inequality of dX to points drawn from both distributions. In every headline application—cat/heart shape matching, MNIST↔USPS, Apples↔Oranges—the two domains are different metric spaces, so no common dX exists and Proposition 3 does not apply. The paper nevertheless asserts (abstract and §4.1) that Tukey-penalized GW 'becomes robust to Huber contamination' in the cross-domain setting. Remark 7's empirical bound also cites Proposition 3, inheriting the same-space premise. The gap is not fatal to the idea: applying Proposition 2(i) with intermediate space (X, dX, µ) gives dTGW(µ', ν) ≤ dTGW(µ', µ) + dTGW(µ, ν); the same-space bound dTGW(µ', µ) ≤ W_Tp(µ', µ) ≤ τ ε^{1/p} follows from the existing proof restricted to X, yielding dTGW(µ', ν) ≤ τ ε^{1/p} + dTGW(µ, ν) ≤ τ ε^{1/p} + dGW(µ, ν). This corollary is absent from the manuscript, so the central cross-domain robustness claim currently rests on an unstated proof.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes three robustifications of the Gromov-Wasserstein (GW) distance for cross-domain alignment: Tukey/Huber penalized GW (TGW/HGW), locally robust GW (LRGW) via truncated metrics, and plan-robustification through RRGM. For each variant, the authors study metric properties, robustness guarantees, and relations to GW, and they report experiments on shape matching and image-to-image translation under contamination. The central claims are that TGW/HGW are robust to Huber contamination, that LRGW reduces to an OT problem between truncated observations, and that RRGM yields robust transport maps; each claim is supported by at least one theorem, algorithm, or experimental comparison.","tokens_in":39237,"tokens_out":5651,"duration_ms":51347,"significance":"The paper addresses an important and timely problem: robustifying GW-type distances, which are widely used for cross-domain alignment but are known to be brittle under outliers. The concrete contributions include the triangle inequality for TGW (Proposition 2), a same-space contamination bound (Proposition 3), an upper-bound duality for the cross-term of LRIGW (Proposition 11), and non-asymptotic concentration inequalities for the RRGM loss (Propositions 21–22). The experimental sections are detailed and the code is made available, which strengthens reproducibility. However, as detailed below, the most prominent claims—especially cross-domain robustness and the 'boils down to OT' statement—are not fully supported by the theorems as written, so the significance of the paper depends on whether these gaps can be closed by revision.","major_comments":[{"comment":"The robustness guarantee is stated only for distributions µ and ν on the same metric measure space, yet the abstract and Section 4.1 present Proposition 3 as establishing robustness for cross-domain GW. The proof (Appendix A, eqs. 33–35) uses the triangle inequality of a single metric dX for points drawn from both distributions, which is unavailable when X ≠ Y. Remark 7 then applies Proposition 3 to empirical distributions on two different spaces by calling it a corollary, but no proof of that extension is given. Adding the triangle-inequality corollary via the intermediate space (X, dX, µ), namely dTGW(µ′,ν) ≤ τ ε^{1/p} + dGW(µ,ν), would make the cross-domain claim true; as it stands the claim is unsupported.","section":"Section 4.1, Proposition 3"},{"comment":"The contribution bullet states that 'solving the same boils down to calculating an OT between truncated observations from µ and ν (Proposition 11).' However, Proposition 11 provides an upper bound on the F2 component of the squared LRIGW cost, not an equality for the full distance d²_LRIGW = F1 + F2. The F1 term is a sum of marginal moments and is not expressed as an OT problem. The text later acknowledges an 'upper bound' (eq. 16), but the contribution language is not qualified accordingly. The authors should either prove an equality for the full minimal cost or rewrite the claim to say that the OT reduction applies to the cross-term upper bound only.","section":"Section 4.2, Proposition 11"},{"comment":"HGW is introduced as the second penalized variant, and the text claims that 'HGW poses as a robust estimate of the corresponding GW value as it follows a property similar to Proposition 3.' No theorem or proof is given for this robustness property, and the Huber loss in Definition 8 is not monotone, so the TGW proof does not carry over directly. Since HGW is one of the three headline methods and is used in the shape-matching experiments, this unproved assertion is load-bearing. The authors should either provide a rigorous statement with proof or explicitly downgrade the claim.","section":"Section 4.1, Definition 8 and following text"},{"comment":"The RRGM construction relies on a denoising map Id̃_ε satisfying Id̃_ε#α ≤ (1−ε)α, and the non-emptiness of Π_ε(µ,ν) is justified by an appeal to partial mass transport. However, no construction or learnable parametrization of Id̃_ε is given, and the loss (30) treats this map as an input. Since the claimed robust translation capability depends on the existence and identifiability of such a map, the authors should either specify how Id̃_ε is obtained in practice or present the existence argument as a formal lemma with the required assumptions.","section":"Section 4.3, eq. (29)"}],"minor_comments":[{"comment":"The y-axis labels 'Cost Metrics' are unhelpful; please specify the plotted quantity, e.g., 'average loss value' or 'distance value'.","section":"Figure 3 and Figure 7"},{"comment":"The step WTp((1−ε)µ + εµc, µ) ≤ WTp(εµ, εµc) is used without justification; a one-sentence explanation of why deleting the common part cannot increase the OT cost would improve readability.","section":"Appendix A, proof of Proposition 3, inequality (34)"},{"comment":"The text defines τ = m̃ + 3σ̃, where m̃ and σ̃ are the median and mean absolute deviation about the median, but the subsequent discussion in Appendix B.1 refers to 'mean deviation about median' without spelling out the estimator; please align the terminology.","section":"Section 4.1, parameter selection"}],"recommendation":"major_revision","confidential_remarks":"The paper's headline claims are stronger than the theorems currently support. The missing triangle-inequality corollary for Proposition 3 is a small but necessary addition, and the 'boils down to OT' phrasing for Proposition 11 should be corrected or substantiated. The experimental work is extensive and the code release is a strength, but the revision should prioritize aligning the claims with the proved statements."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Three new ways to robustify Gromov-Wasserstein, and they are not just unbalanced GW in disguise. The Tukey/Huber penalized distances (TGW/HGW), the locally truncated LRGW, and the plan-regularized RRGM all have real content, and the paper gives them structural properties that largely check out. Proposition 2's metric properties are plausible, and Proposition 3 is a legitimate robustness bound within its stated scope. The algorithms for HGW and the shape-matching experiments are reasonable, and the connection between LRGW and trimmed observations is a nice idea.\n\nThe soft spot is the one that matters. Proposition 3 is proved only for two distributions on the same metric space, yet the abstract and Section 4.1 use it to claim cross-domain robustness. The stress-test note is right about the gap, and also right that it is fixable: applying the triangle inequality through an intermediate space gives dTGW(µ', ν) ≤ dTGW(µ', µ) + dTGW(µ, ν), and the same-space bound on the first term yields τ ε^{1/p} + dGW(µ, ν). That corollary is missing from the manuscript, so the central claim currently rests on an unstated proof. The authors should either add it or qualify the claim.\n\nSecond soft spot: the 'solving boils down to OT' bullet for LRGW is overstated. Proposition 11 is an upper bound on the F2 term only, not an equality for the full distance. That is a meaningful difference. A careful rewrite of that contribution would help.\n\nThe experiments are suggestive but not decisive. FID differences of a few points, single seeds, no error bars, and a parameter ablation that is qualitative rather than statistical. That is acceptable for a first version but needs tightening.\n\nOverall, this is a paper that deserves a serious referee. The ideas are new, the proofs are mostly careful, and the gap is addressable rather than fatal. The right outcome is likely a revise-and-resubmit with a requirement to fix the same-space claim and tone down the OT-reduction statement.","headline":"Genuinely new robust GW formulations with a real proof gap: the headline cross-domain guarantee is only proved for same-space distributions, though a one-line triangle-argument corollary would likely close it.","tokens_in":39756,"tokens_out":1816,"would_cite":true,"duration_ms":19192,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","62F35","60B05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims three outlier-resistant variants of the Gromov-Wasserstein distance—Tukey/Huber penalization, local metric truncation, and robust Monge-map regularization—preserve metric structure and keep transport plans as probability…","keywords":["Gromov-Wasserstein distance","robust optimal transport","Huber contamination","Tukey loss","domain alignment","image-to-image translation","Gromov-Monge distance","metric measure spaces"],"falsifier":"For two distributions on genuinely different metric spaces, contaminate one side with a single outlier moved arbitrarily far and record $d_{\\mathrm{TGW}}(\\mu',\\nu)-d_{\\mathrm{GW}}(\\mu,\\nu)$ at fixed $\\tau$; if this gap grows without bound while the same-space bound stays finite, the claimed cross-domain robustness is not a population-level property.","tokens_in":38675,"feed_emoji":"🛡️","tokens_out":11219,"duration_ms":93003,"temperature":0.7,"pith_summary":"The paper's project is to make the Gromov-Wasserstein (GW) distance—the standard measure of geometric alignment between distributions on different spaces—resistant to outliers without sacrificing the structure that makes it useful. It argues that robustifications imported from optimal transport, which rely on partial mass transport or unbalanced couplings, give up metric properties and handle one-sided contamination poorly. In their place it proposes three families: Tukey and Huber penalized GW distances, locally robust GW built on truncated metrics, and a robust reversible Gromov-Monge formulation for learning maps. For each it proves metric and robustness properties and shows empirically that they beat comparison baselines in shape matching and image-to-image translation under contamination. If correct, the payoff is that outlier-resistant alignment can keep the transport plan a true probability distribution while staying close to the unperturbed GW value.","feed_headline":"Three distances make Gromov-Wasserstein alignment outlier-safe","feed_subtitle":"Tukey-Huber penalization, metric truncation, and Monge-plan regularization each preserve metric structure under contamination.","key_machinery":"The load-bearing objects are three distinct attack points on the GW objective. The Tukey loss $T_p(x)=\\min\\{|x|^p,\\tau^p\\}$ and the Huber loss $H(x)=x^2/(2\\tau)$ for $|x|\\leq\\tau$ and $|x|-\\tau/2$ otherwise cap the contribution of huge pairwise-distance discrepancies; the triangle inequality for the induced $T_p$ 'norm', which follows from subadditivity of $T_p^{1/p}$, is what lets TGW inherit GW's metric structure. The truncated metric $l_\\lambda(d)=\\min\\{d,\\lambda\\}$ modifies the distances themselves before they enter the cost; Proposition 11 shows the locally robust inner-product GW cost is bounded above by an optimal-transport problem on truncated observations, which ties trimming to OT and enables sample-complexity arguments. The coupling family $\\Pi_\\epsilon(\\mu,\\nu)=\\{\\pi:\\pi=(\\tilde{\\mathrm{Id}}_\\epsilon,F)_\\#\\mu=(G,\\tilde{\\mathrm{Id}}_\\epsilon)_\\#\\nu\\}$, with $\\tilde{\\mathrm{Id}}_\\epsilon{}_\\#\\alpha\\leq(1-\\epsilon)\\alpha$, implements partial alignment while keeping plans as probability distributions and leads to the RRGM Lagrangian with $W_1$ measure-preservation terms. The chain $\\mathrm{LRGW}\\leq\\mathrm{TGW}\\leq\\mathrm{GW}$ and the reduction of RRGM to robust OT tie the three proposals together.","core_discovery":"The central claim is that GW is best robustified at the level of the distortion, the metric, or the coupling, not by relaxing marginal constraints. First, replacing the raw distortion in the GW objective with the Tukey loss $T_p(x)=\\min\\{|x|^p,\\tau^p\\}$ gives the Tukey-GW distance $d_{\\mathrm{TGW}}$, a pseudometric that lower-bounds GW, converges to it as $\\tau\\to\\infty$, and under Huber's $\\epsilon$-contamination satisfies $d_{\\mathrm{TGW}}(\\mu',\\nu)\\leq \\tau\\epsilon^{1/p}+W_{T_p}(\\mu,\\nu)$ when both measures live on the same metric space; the smooth Huber variant gives a computable HGW. Second, truncating the base metrics $d_X,d_Y$ with $l_\\lambda(d)=\\min\\{d,\\lambda\\}$ yields locally robust GW, a lower bound to Tukey-GW whose cost is shown to become an optimal-transport problem on trimmed observations, and the construction extends to probabilistic metric-measure spaces via the robust Wasserstein distance $W_p^\\epsilon$. Third, regularizing the admissible couplings with clean-proxy marginals or partial maps $\\tilde{\\mathrm{Id}}_\\epsilon$ whose pushforward is dominated by $(1-\\epsilon)$ of the original measure produces the RRGM loss for robust image translation. The three formulations are positioned as complementary: TGW protects against extreme distortions, LRGW protects at the nascency of pairwise distances, and RRGM protects the learned measure-preserving map.","pith_inferences":["The same truncation and penalization ideas should transfer to other GW-type objectives, such as Fused GW or Z-GW, wherever the distortion or base metric is the vulnerable part; the paper's Lemma 9 suggests this transfer is direct for any metrics satisfying the same pointwise inequality.","Because Proposition 3 is restricted to one shared metric space, a genuine cross-domain guarantee would likely need a bound in terms of the dissimilarity between the two spaces, such as their Gromov-Hausdorff distance or an optimal embedding distortion; the experiments do not substitute for that.","The proposed threshold choice $\\tau=\\tilde{m}+3\\tilde{\\sigma}$ is a heuristic; a data-dependent, contamination-level-adaptive rule that provably preserves the metric and robustness bounds would be a natural next test.","Since RRGM and LRGW both reduce to robust OT on trimmed or proxy distributions, their sample complexity can probably be analyzed with existing robust-OT bounds, and Proposition 21 is a first concentration result in that direction."],"forward_implications":["Tukey-GW is a pseudometric lower bound to GW that recovers it as $\\tau\\to\\infty$, so robust alignment can stay inside the balanced-coupling framework.","Under Huber contamination and same-space support, TGW is provably a robust estimator of GW with an explicit bound, and the resilience of distributions under $W_{T_p}$ follows from Corollary 5.","HGW inherits the entropic GW algorithm with a Huber cost at the same $O(m^2n^2)$ complexity class, yielding robust shape matching with full marginal distributions.","LRGW reduces to an OT problem on truncated observations, so its computation and sample complexity can be approached with OT tools, and it extends to Gaussian-mixture and probabilistic mm spaces.","RRGM improves noisy MNIST-to-USPS translation over CycleGAN and reversible Gromov-Monge baselines in reported FID scores while keeping plans as probability distributions."],"supporting_citations":[{"why":"It defines the Gromov-Wasserstein distance and its metric properties, which are the objects the paper robustifies and the baselines for the proposed distances.","marker":"Mémoli (2011)"},{"why":"It introduces the $\\epsilon$-contamination model and robust estimation principles that motivate Tukey-GW and HGW.","marker":"Huber (1964)"},{"why":"It supplies the dual formulation of the robust Wasserstein distance $W_p^\\epsilon$ that LRGW uses and the resilience concept behind Corollary 5.","marker":"Nietert et al. (2022)"},{"why":"It provides the robust-optimal-transport approach with proxy distributions that RRGM's plan regularization extends and compares against.","marker":"Balaji et al. (2020)"},{"why":"It defines partial optimal transport, the PGW baseline used in experiments and the connection point for RRGM's partial alignment.","marker":"Chapel et al. (2020)"},{"why":"It defines unbalanced Gromov-Wasserstein, the main approach the paper argues against and the UGW baseline in experiments.","marker":"Séjourné et al. (2021)"},{"why":"It gives the entropic Gromov-Wasserstein algorithm and cost decomposition that Algorithm 1 adapts to the Huber loss.","marker":"Peyré et al. (2016)"},{"why":"It introduces the reversible Gromov-Monge sampler whose loss RRGM robustifies and extends to contaminated data.","marker":"Hur et al. (2024)"}],"fun_headline_variants":["Three ways to protect Gromov-Wasserstein from outliers","Robust GW: Tukey loss, metric truncation, and Monge-plan ties","GW made outlier-safe: distortion, metric, or coupling fixes","Tukey, truncation, and Monge plans: a robust GW trio","Three complementary robustifications for cross-domain alignment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative robustness guarantee for Tukey-GW is proved only when both distributions are supported on the same metric space, so the paper's strongest cross-domain guarantee does not actually apply to cross-domain problems and instead rests on heuristic thresholds and experiments.","fun_headline_variants_meta":{"raw":{"variants":["Three ways to protect Gromov-Wasserstein from outliers","Robust GW: Tukey loss, metric truncation, and Monge-plan ties","GW made outlier-safe: distortion, metric, or coupling fixes","Tukey, truncation, and Monge plans: a robust GW trio","Three complementary robustifications for cross-domain alignment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1400,"prompt_tokens":1013,"completion_tokens":387,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":296}},"tokens_in":629,"tokens_out":387,"duration_ms":4130,"temperature":1.0,"reasoning_tokens":296,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:01:40.667203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For two distributions on genuinely different metric spaces, contaminate one side with a single outlier moved arbitrarily far and record $d_{\\mathrm{TGW}}(\\mu',\\nu)-d_{\\mathrm{GW}}(\\mu,\\nu)$ at fixed $\\tau$; if this gap grows without bound while the same-space bound stays finite, the claimed cross-domain robustness is not a population-level property.","supporting_citations":[],"review_version":1}