{"id":"1e6ba3ee-e924-4c12-9535-342562ad27be","arxiv_id":"2608.13374","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The optimal Hölder exponent for max-sliced 1-Wasserstein distance on the unit ball is 2/(d+2), and under a Voronoi matching condition W_p is bounded by a nearly optimal multiple of the sliced distance.","lead":"This math paper proves the sharpest possible comparison inequalities between the full Wasserstein distance and its cheaper sliced and max-sliced versions in Euclidean space. The results settle an open question about the optimal 'curse of dimensionality' exponent and quantify when a discrete approximation can give nearly lossless comparisons.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Main internal proofs appear sound; the least secure link is the unverified import of the Chen–Travaglini L1 discrepancy bound (4.2) supporting Theorem 1.4.","rationale":"The reader's weakest_assumption correctly identifies the external Chen–Travaglini discrepancy bound as the most load-bearing unverified input; I agree with that assessment. After checking the internal proofs, I found no mathematical gap in Theorem 1.1 or in the upper-bound Theorems 1.3 and 1.5. The only concern that could materially alter the stated near-optimality is the possibility that (4.2) is misquoted or not satisfied by the constructed lattice, which would weaken Theorem 1.4 from a polylogarithmic gap to a power-of-N gap. However, relying on a published theorem is standard mathematical practice, and the cited result appears well-matched to the application, so I do not regard this as grounds to change the reader's ACCEPT verdict. The concrete numerical check would settle whether the imported bound behaves as stated; if it does, the paper's claims stand. The finiteness of K noted by the reader is a minor presentational caveat: when K is infinite the inequality is vacuous unless SW_{p,k} = 0, in which case µ = ν and the bound is trivial, so it does not affect the central results.","tokens_in":19365,"tokens_out":22763,"duration_ms":216852,"concrete_test":"Verify (4.2) for the specific lattice Λ_M in dimensions d = 2 and 3 with M = 16, 32, 64, 128. For a Monte-Carlo sample of at least 10^4 directions θ uniform on S^{d−1}, compute sup_{t∈[−√d,√d]} |card(Λ_M ∩ P_{θ,t}) − N 2^{−d}|P_{θ,t}|| and average over θ; compare the growth of this empirical L1 discrepancy with (log N)^d. If the growth is not polylogarithmic in N, Lemma 4.1(iv) is unsupported. In parallel, consult the published Chen–Travaglini theorem statement to confirm that it indeed supplies an L1 bound over the sphere with exponent d for half-spaces with arbitrary orientation in [−1,1]^d.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result, Theorem 1.1, is supported by a self-contained construction (Lemmas 2.2–2.5) that we checked: the moment-matched perturbation, projective packing, probabilistic shift selection, and normalization all hold together, and the exponent calculation 1 + d/2 is consistent. Theorems 1.3 and 1.5 also check out: Lemma 3.3's anticoncentration bound is uniform in k, and the choice of τ in (3.6) makes the union-bound probability at most 1/2, yielding W_p ≤ C sqrt(d/k) K SW_{p,k}. The weakest load-bearing link is in Theorem 1.4: the lower bound in (4.3) inherits its entire quantitative content from the external discrepancy estimate (4.2), sup_t ∫ |card(Λ_M ∩ P_{θ,t}) − N 2^{−d}|P_{θ,t}|| dσ(θ) ≤ c_d (log N)^d, imported from Chen and Travaglini without proof. Lemma 4.1(iv) converts this into SW_{1,1} ≤ C_d (log N)^d / N. If (4.2) is misstated (for example, the exponent should be different, or the cited theorem controls L2 discrepancy or a different class of test sets) or if the bound fails for the specific grid Λ_M with arbitrary half-space orientations, then the ratio in (4.3) would carry an extra power of N, and the claimed near-optimal linear dependence on K would not follow. This concern does not affect Theorem 1.1 and does not indicate an internal inconsistency; it marks the point where the proof of Theorem 1.4 depends on an unverified imported result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies quantitative comparisons between the p-Wasserstein distance W_p and its sliced and max-sliced analogues SW_{p,k} and tilde-W_{p,1} on Euclidean space. Theorem 1.1 constructs, for every d≥2 and every sufficiently small rho>0, probability measures μ,ν supported on the unit ball such that W_1(μ,ν)≥rho and tilde-W_{1,1}(μ,ν)≤C_d rho^{1+d/2}; this proves that the Hölder exponent 2/(d+2) obtained by Bobkov and Götze for the max-sliced 1-Wasserstein distance is optimal for every dimension. Theorems 1.3 and 1.5 establish that under Condition 1.2 (an optimal coupling transporting μ-a.e. point to a nearest atom of a discrete measure ν), W_p(μ,ν)≤C√(d/k) K^{(k)}_{μ,ν} SW_{p,k}(μ,ν), where the complexity parameter K^{(k)} is at most the number of atoms and can be substantially smaller. Theorem 1.4, based on the Chen–Travaglini half-space discrepancy estimate, shows that the linear dependence on K in Theorem 1.3 cannot be improved by more than a polylogarithmic factor. The paper also contains a useful comparison with the recent bound of Park and Slepčev and a discussion of the optimal-exponent results of Carlier, Figalli, Mérigot, and Wang.","tokens_in":19626,"tokens_out":28547,"duration_ms":281570,"significance":"These results are significant. The sharpness of the exponent 2/(d+2) for max-sliced Wasserstein distances resolves a question left open in Bobkov and Götze's work, and the construction is a nontrivial adaptation of the Gaussian-pancake method to the averaging regime. The structural comparison under Condition 1.2 is clean and likely useful: it replaces a smallness assumption on W_infty in the Park–Slepčev theorem by a transport-geometric condition and gives explicit dependence on the complexity parameter. The proof of the upper bounds is self-contained, with a careful anticoncentration lemma on the Stiefel manifold, and the calculations in Section 2 are checkable; I verified the exponent bookkeeping and the packing/sparsity constraints. The lower bound in Theorem 1.4 is less self-contained because it imports the Chen–Travaglini discrepancy bound, but this is a legitimate external input and the way it is converted into the sliced-Wasserstein lower bound is correct. Overall the manuscript is a strong contribution to the optimal-transport literature.","major_comments":[],"minor_comments":[{"comment":"The entire quantitative content of Theorem 1.4 is inherited from the Chen–Travaglini discrepancy bound. As written, (4.2) involves sup_t outside the angular integral. If the cited result is the usual integrated L1 form with ∫∫ |...| dt dσ, please state that form explicitly and rewrite Lemma 4.1(iv) accordingly; this is a local fix but important for the reader to verify the constant's dependence.","section":"Section 4, Eq. (4.2)"},{"comment":"The comparison is non-vacuous only when K^{(k)}_{μ,ν}<∞; the statements should say this explicitly. The proof already says 'we may assume K<∞', but the theorem statements should not leave the reader to infer that the sum in (1.15) is finite.","section":"Theorems 1.3 and 1.5"},{"comment":"There is an index typo: in 1{d(x+sρu_j, P_i)≤|s|ρ} the vector should be u_i, since the point belongs to the i-th slab; as written it conflicts with the outer summation index j.","section":"Eq. (2.22)"},{"comment":"After reducing to z=0, the shorthand λ^{(ρ)}_{(u,r)} and λ_{(u,r)} is introduced without an explicit definition; please add a sentence identifying these as λ^{(ρ)}_{(u,0,r)} and λ_{(u,0,r)}.","section":"Lemma 2.2 proof"},{"comment":"There are formatting artifacts in the typeset text, e.g., 'itsufficestobound' in Section 2.3 and missing spaces in the display after (2.13); please correct. Also, in Section 4, 'By the symmetry, the bound (4.2) holds for all t∈R' would be clearer if it said 'by reflecting θ to -θ and t to -t'.","section":"Typesetting and wording"}],"recommendation":"minor_revision","confidential_remarks":"None beyond the report. If the editor can conveniently verify that the Chen–Travaglini inequality is quoted in the exact form needed, that would resolve the only external-input concern; I did not find an internal error in the derivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jon, the paper is worth your time. It gives the first example showing Bobkov–Götze's exponent 2/(d+2) for max-sliced 1-Wasserstein on the unit ball is optimal in every dimension d≥2. The construction is a genuine improvement over the single-direction one for sliced Wasserstein: it superimposes moment-matched perturbations along an ε-packing of projective space, and the dyadic summation over angles is clean. I checked the main steps—the smoothing lemma, the probabilistic shift selection, the normalization—and they work. Theorem 1.1 is new and complete.\n\nTheorems 1.3–1.5 also look solid. The Voronoi condition is a natural structural assumption, and the comparison bound with the complexity parameter K is useful because K can be much smaller than the number of atoms. The proof via anticoncentration on the Stiefel manifold is short and correct. The lower bound Theorem 1.4 is the right companion: it shows linear dependence on K is necessary up to polylog, using the Chen–Travaglini half-space discrepancy estimate.\n\nThe soft spots are minor. The lower bound in Theorem 1.4 leans entirely on the Chen–Travaglini bound (4.2), which is imported without proof. I didn't chase down the exact statement in the original paper, and neither should you trust it blindly; if that bound had a different exponent, the near-optimal dependence on K would weaken. But this is a citation-check, not a flaw in the internal reasoning, and it doesn't touch Theorem 1.1. The second caveat is that K needs to be finite; the paper says this implicitly, but it's worth flagging for the reader. The AI disclosure is honest and doesn't affect the math.\n\nOverall: this is a strong paper for optimal transport and discrepancy theory. It deserves a serious referee. The main things to ask the authors are to double-check the statement of (4.2) and maybe add a remark about the finite-K assumption. I'd take it to reading group and cite it if I worked on sliced Wasserstein.","headline":"A clean, citable resolution of the optimal exponent for max-sliced Wasserstein comparison, with a genuinely new construction; the only real soft spot is the unproved import of a discrepancy bound for the lower-bound half.","tokens_in":20250,"tokens_out":2479,"would_cite":true,"duration_ms":22738,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B10","49Q22","60D05","11K38"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves the Hölder exponent 2/(d+2) for max-sliced Wasserstein distance on the unit ball is optimal in every dimension, and that structural bounds need a nearly linear complexity factor.","keywords":["sliced Wasserstein distance","max-sliced Wasserstein distance","Hölder comparison exponents","optimal transport","half-space discrepancy","Voronoi matching","Kantorovich–Rubinstein norm","random projections anticoncentration"],"falsifier":"Evaluate numerically, for $d=3$ and increasing $M$, the quantity $\\sup_t \\int_{S^{d-1}} |\\mathrm{card}(\\Lambda_M \\cap P_{\\theta,t}) - N 2^{-d}|P_{\\theta,t}||\\, d\\sigma(\\theta)$ for the $N=(2M+1)^d$ lattice in $[-1,1]^d$; if this discrepancy grows faster than $(\\log N)^d$, then the bound in Lemma 4.1(iv) collapses and the proof of Theorem 1.4 fails. Alternatively, produce a pair $(\\mu,\\nu)$ satisfying Condition 1.2 with $K_{\\mu,\\nu}=K$ and $W_1(\\mu,\\nu)/\\mathrm{SW}_{1,1}(\\mu,\\nu) = o(K/(\\log K)^d)$, which would refute the claimed near-optimality directly.","tokens_in":19081,"feed_emoji":"📏","tokens_out":16794,"duration_ms":136268,"temperature":0.7,"pith_summary":"This paper pins down how much of the Wasserstein distance $W_1$ between two distributions survives when one records only one-dimensional projections, either averaged (sliced) or worst-case (max-sliced). It proves that the known upper bound $W_1 \\le C \\widetilde{W}_{1,1}^{2/(d+2)}$ on the unit ball cannot be improved: for every $d\\ge 2$ and small $\\rho$ there are measures on the ball with $W_1 \\ge \\rho$ but $\\widetilde{W}_{1,1} \\le C_d \\rho^{1+d/2}$, so no comparison with a larger Hölder exponent exists. Under a nearest-atom matching condition, it proves $W_p \\le C \\sqrt{d}\\, K \\,\\mathrm{SW}_{p,1}$, with an analogous bound for $k$-dimensional projections replacing $K$ by $K^{(k)} \\le N^{1/k}$. It also proves, via a half-space discrepancy construction, that the linear factor $K$ is necessary up to a polylogarithmic factor.","feed_headline":"Max-sliced Wasserstein comparisons: exponent 2/(d+2) is optimal","feed_subtitle":"For every dimension d≥2 the old bound cannot be improved, and discrete comparisons must pay a near-linear factor K.","key_machinery":"The central identity is Condition 1.2, under which $W_p^p(\\mu,\\nu)$ equals $\\int d(x,Y)^p\\, d\\mu(x)$ (Lemma 3.1), because the optimal coupling is exactly the nearest-atom Voronoi matching. The proof of the upper bound then reduces to an anticoncentration estimate for uniform random projections: for $U$ uniform on the Stiefel manifold $G_{d,k}$, $\\mathbb{P}(\\|Uv\\| < t\\sqrt{k/d}\\,\\|v\\|) \\le (C_0 t)^k$ (Lemma 3.3), which turns a small average projected distance into a pointwise lower bound on $\\mathbb{E}[d(Ux,UY)^p]$ in terms of $d(x,Y)/K^{(k)}$. For the sharpness of the Hölder exponent, the machinery is a 'star of Gaussian pancakes': signed one-dimensional perturbations whose moments up to order $d+1$ vanish, planted along an $\\epsilon$-packing of directions in $\\mathbb{RP}^{d-1}$, with Lemma 2.2 bounding the projected Kantorovich–Rubinstein norm by $C r \\min(\\epsilon, \\epsilon^{d+2}/\\beta^{d+1})$ and the local entropy bound (2.16) limiting how many pancakes any fixed direction sees. The lower bound on $K$ uses the cited half-space discrepancy estimate, which says the lattice of cube centers has $L^1$ average half-space discrepancy $O((\\log N)^d)$, giving $\\mathrm{SW}_{1,1} \\le C_d (\\log N)^d / N$.","core_discovery":"On the paper's own terms, the headline findings are two. First (Theorem 1.1), the max-sliced 1-Wasserstein distance cannot control $W_1$ on $\\mathcal{P}(B_1)$ with a Hölder exponent better than $2/(d+2)$: the authors construct explicit pairs $(\\mu_\\epsilon, \\nu_\\epsilon)$ with $W_1(\\mu_\\epsilon, \\nu_\\epsilon) \\asymp \\epsilon^{2-2/d}$ and $\\widetilde{W}_{1,1}(\\mu_\\epsilon, \\nu_\\epsilon) \\lesssim \\epsilon^{(d+1)-2/d}$, and the ratio of exponents is exactly $1+d/2$, forcing $\\beta \\le 2/(d+2)$ when combined with the known upper bound. Second (Theorems 1.3 and 1.4), under Condition 1.2 — that some optimal coupling matches every $\\mu$-mass point to a nearest atom of the discrete measure $\\nu$ — the comparison $W_p \\le C \\sqrt{d}\\, K_{\\mu,\\nu} \\,\\mathrm{SW}_{p,1}$ holds with $K_{\\mu,\\nu}$ the essential supremum of $d(x,Y) \\sum_{y\\in Y} 1/\\|x-y\\|$, and the linear dependence on $K$ is near-optimal: there are examples with $K_{\\mu,\\nu} \\asymp_d K$ and $W_1 \\ge c_d K/(\\log(K+1))^d \\,\\mathrm{SW}_{1,1}$. The $k$-dimensional projection version (Theorem 1.5) replaces $K$ by $K^{(k)} \\le N^{1/k}$ and the factor $\\sqrt{d}$ by $\\sqrt{d/k}$.","pith_inferences":["A practical reading is that algorithms which replace Wasserstein by sliced Wasserstein on discrete data should expect an effective cost factor comparable to $K$, i.e., to the number of atoms within a few multiples of the nearest-neighbour distance; this is testable on point clouds by computing $K$ empirically and comparing $W_1/\\mathrm{SW}_{1,1}$.","The discrepancy input of Theorem 1.4 is the only non-elementary step; substituting any point set with sub-polylogarithmic $L^1$ half-space discrepancy into the same cube construction would yield the same lower bound, suggesting the $K/(\\log K)^d$ phenomenon is a projection-averaging effect rather than a special feature of the lattice.","The moment-matching 'pancake' construction is tailored to $p=1$ through the Kantorovich–Rubinstein norm; a natural extension is to build analogous signed perturbations with more vanishing moments to test whether the optimal exponents for $p>1$ obey the same dimensional barrier.","For Gaussian mixtures with well-separated components, $K$ should be close to $1$, so Theorem 1.3 predicts near-Lipschitz $W_p$--$\\mathrm{SW}_{p,1}$ equivalence in that regime — a quantitative prediction that can be checked by simulation."],"forward_implications":["For $p=1$ on the unit ball, the comparison $W_1 \\le C \\widetilde{W}_1^\\beta$ cannot hold with $\\beta > 2/(d+2)$, so the max-sliced distance is exponentially coarser than $W_1$ in high dimension.","If $\\nu$ is discrete and the optimal transport is a nearest-atom matching, then $W_p$ and $\\mathrm{SW}_{p,1}$ become equivalent up to the factor $\\sqrt{d}\\,K$, so the information loss of slicing is governed by the local crowding $K$ rather than by the total number of atoms.","Using $k$-dimensional projections improves the constant: for $N$ atoms the multiplier is at most $C\\sqrt{d/k}\\, N^{1/k}$, so higher-dimensional projections are strictly less wasteful.","The near-linear dependence on $K$ is intrinsic: some pairs satisfying Condition 1.2 force $W_1 \\ge c_d K/(\\log(K+1))^d \\,\\mathrm{SW}_{1,1}$, so no method can remove the $K$ factor entirely below polylog.","Since $\\mathrm{SW}_{p,1} \\le \\widetilde{W}_{p,1}$, the same bounds and lower-bound construction hold verbatim with the max-sliced distance in place of the sliced distance."],"supporting_citations":[{"why":"Supplies the upper bound $W_1 \\le C \\widetilde{W}_{1,1}^{2/(d+2)}$ on the unit ball and poses the question of its optimality, which Theorem 1.1 settles.","marker":"[BG24]"},{"why":"Provides the moment-matched one-dimensional perturbation technique and the optimal sliced-distance exponent $\\alpha=1/d$ that Section 2 adapts to the max-sliced setting.","marker":"[Car+25]"},{"why":"States the $L^1$ half-space discrepancy estimate for the cube lattice, used in Lemma 4.1(iv) to bound $\\mathrm{SW}_{1,1}$ by $O((\\log N)^d/N)$.","marker":"[CT11]"},{"why":"Gives the local $W_2$--$\\mathrm{SW}_{2,1}$ comparison under a small $W_\\infty$ condition that Theorem 1.3 extends to full nearest-atom matching and all $p\\ge 1$.","marker":"[PS25]"},{"why":"Establishes the one-dimensional identity $W_1 = \\|F_\\mu-F_\\nu\\|_{L^1}$, used in the projected smoothing bound and in the discrepancy-based $\\mathrm{SW}_{1,1}$ computation.","marker":"[BL19]"},{"why":"Supplies the Beta distribution of the squared norm of a uniform random projection, underlying the anticoncentration estimate Lemma 3.3.","marker":"[FM90]"},{"why":"Bounds the metric entropy of the Grassmannian, giving the $\\epsilon$-packing size $m \\asymp \\epsilon^{-(d-1)}$ in the star construction.","marker":"[Paj98]"},{"why":"Gives the volume bounds for balls in $\\mathbb{RP}^{d-1}$ used to establish the local packing bound (2.16).","marker":"[PV14]"}],"fun_headline_variants":["Max-sliced Wasserstein: optimal exponent in all dimensions","Sliced Wasserstein comparisons are sharp in every dimension","Near-optimal linear factor for discrete sliced Wasserstein","Optimal 2/(d+2) exponent for max-sliced Wasserstein","Wasserstein vs sliced: tight bounds and near-optimal K"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing external input is the half-space discrepancy bound for the cubic lattice imported from the cited reference and stated as (4.2); if the true discrepancy of this lattice is larger than polylogarithmic, the lower bound $W_1/\\mathrm{SW}_{1,1} \\gtrsim K/(\\log K)^d$ no longer follows, and the near-optimality of the $K$ factor in Theorem 1.3 is not established.","fun_headline_variants_meta":{"raw":{"variants":["Max-sliced Wasserstein: optimal exponent in all dimensions","Sliced Wasserstein comparisons are sharp in every dimension","Near-optimal linear factor for discrete sliced Wasserstein","Optimal 2/(d+2) exponent for max-sliced Wasserstein","Wasserstein vs sliced: tight bounds and near-optimal K"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00046,"raw_usage":{"total_tokens":2418,"prompt_tokens":1172,"completion_tokens":1246,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":788,"completion_tokens_details":{"reasoning_tokens":1157}},"tokens_in":788,"tokens_out":1246,"duration_ms":10611,"temperature":1.0,"reasoning_tokens":1157,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:34:45.769232+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate numerically, for $d=3$ and increasing $M$, the quantity $\\sup_t \\int_{S^{d-1}} |\\mathrm{card}(\\Lambda_M \\cap P_{\\theta,t}) - N 2^{-d}|P_{\\theta,t}||\\, d\\sigma(\\theta)$ for the $N=(2M+1)^d$ lattice in $[-1,1]^d$; if this discrepancy grows faster than $(\\log N)^d$, then the bound in Lemma 4.1(iv) collapses and the proof of Theorem 1.4 fails. Alternatively, produce a pair $(\\mu,\\nu)$ satisfying Condition 1.2 with $K_{\\mu,\\nu}=K$ and $W_1(\\mu,\\nu)/\\mathrm{SW}_{1,1}(\\mu,\\nu) = o(K/(\\log K)^d)$, which would refute the claimed near-optimality directly.","supporting_citations":[],"review_version":1}