REVIEW 2 major objections 3 minor 1 cited by
Generalized Optimal Transport
T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A generalized optimal transport program claims to deliver sharp bounds at parametric √n speed for partial identification problems with overlapping marginals and independence restrictions.
desk verdict A serious new framework for partial identification via characteristic-kernel moment restrictions, with a real gap in the √n proof for independence and an abstract that overstates the conditional-independence scope. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The objects that carry the argument are characteristic kernels: integral operators P ↦ ∫ K(s,t)dP(t) that are injective on probability measures, so that equality of kernel means becomes equality of distributions. This converts identified marginals of subvectors, and unconditional or conditional independence, into a continuum of moment equalities E_P[a(T)(s)] = c(s). The resulting infinite-dimensional linear program is rewritten, through Banach-space duality and exact penalization with a fixed penalty level, as an unconstrained convex program over L2, and then projected onto spaces of polynomials on each cube. The projection is tractable because the orthogonal projection operator is self-adjoint, so the sieve bias factors through the projection applied to the known functions c and a(t), yielding a Jackson-type bound $Cκ^{{-r}}$ under Sobolev smoothness; a compactification term Δ(θ,γ) remains, and its rate controls the slower, upward approach to the bound.
What would settle it
Take a feasible problem with identified moment functions c and a(t) that lie outside the Sobolev class of Assumption 3, compute the sieve gap δ_p = ||(I-p_J)c|| + max_t ||(I-p_J)a(t)|| as the polynomial degree J grows, and check whether it decays at the predicted polynomial rate; then simulate the one-sided coverage of the inequality (3). If the sieve gap still decays polynomially, or if the one-sided uniform validity holds despite the nonsmooth moments, the smoothness assumption is not load-bearing; if both fail, the theorem's rate mechanism is confirmed to depend on it.
Extended reading notes
Core claim
The central claim is that the sharp upper bound β(θ) = sup_P E_P[b(T)] over all Borel probability measures on a compact set T satisfying a ν-almost-everywhere continuum of moment equalities E_P[a(T)(s)] = c(s) is captured by a penalized L2 program, and that projecting that program onto polynomials of degree κ yields a finite semi-infinite linear program whose approximation error is bounded by $Cκ^{{-r}}$ when c and a(t) have Sobolev smoothness of order r. Theorem 3.2 turns this into a deterministic perturbation bound in the estimation errors δ_a, δ_b, δ_c and the sieve gap δ_p. Theorem 3.3 then states that, when the restrictions are estimated by sample analogues, the resulting estimator satisfies $$O_P(\gamma_n/\sqrt{n}) \le \beta_{J_n}(\hat{\$\theta$};\gamma_n) - \$\beta$(\$\theta$) \le O_P(\gamma_n/\sqrt{n}) + o_P(1)$$ uniformly over the class, so the bound is approached from the conservative side at a parametric rate; the upward error is governed by a compactification bias Δ(θ,γ) that is related to the ill-posed inverse problem and vanishes as γ grows. In the optimal-transport special case, this asymmetry is exactly what permits one-sided √n validity despite the two-sided minimax rate $n^{{2/d}}$.
Load-bearing premise
The clean √n result rests on the identified moment functions being smooth enough that a polynomial sieve approximates them at the known rate $Cκ^{{-r}}$; if that smoothness is absent, the sieve bias need not vanish at the assumed rate and the one-sided uniform √n validity is lost.
Editorial extensions
If this is right
- Any problem fitting the format—identified marginals, possibly overlapping, plus independence restrictions, with continuous or discrete variables—gets a single computational procedure instead of bespoke linear-programming or optimal-transport solvers.
- The estimator is conservative by construction: for an upper bound it converges from below at rate √n/γ_n uniformly, so it does not understate the identified set's upper limit in large samples.
- For empirical optimal transport with a Lipschitz cost, one obtains a one-sided √n (up to logs) estimator despite the two-sided minimax rate n^{2/d}, and tuning a parameter η trades the two one-sided rates against each other.
- A growing number of discrete moment conditions can be added—for example, spline moments of identified bounded functions—at rate (γ_n + γ_n log l_n)/√n, as long as the added moments are redundant for identification.
- Conditional-independence restrictions can be embedded by the same construction, though the paper notes they require nonparametric estimation of conditional moment generating functions and therefore only nonparametric rates.
Reading between the lines
- An implicit consequence is that problems of combining experimental and observational data, such as the long-term treatment-effect example in the paper, become a routine computation once the variable vector, support, and characteristic kernels are specified.
- The one-sided nature of the estimator suggests a practical reporting rule: use the conservative bound as the policy number, and treat the vanishing upward bias as a diagnostic for how much regularization (the choice of γ) is doing.
- A natural extension the paper leaves open is fully data-driven choice of the kernel, sieve order, and tuning parameter; the theorems only require κ_n to diverge fast enough relative to n^{1/(2r*)}.
- The link to the ill-posed inverse problem hints that the framework could be dualized to deliver inference for nonparametric instrumental-variable-like objects, but that direction is not developed in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a method, Generalized Optimal Transport (GOT), for computing sharp bounds on a linear functional E_P[b(T)] when the econometric model identifies only certain joint subdistributions of T and/or imposes (conditional) independence restrictions. The identification restrictions are encoded as a continuum of moment equalities using characteristic kernels. The resulting infinite-dimensional linear program is dualized, exactly penalized with a fixed penalty, and solved over a finite-dimensional polynomial sieve. The main theoretical results are a deterministic perturbation bound (Theorem 3.2), a one-sided uniform convergence rate for the sieve estimator (Theorem 3.3), an extension to many discrete moments (Theorem 3.5), and an analysis of the compactification bias Delta(theta; gamma), including a sharp characterization in the optimal transport special case (Section 4).
Significance. If the results hold, this is a valuable contribution that unifies and extends LP-based partial identification and empirical optimal transport. The one-sided uniform validity result (Theorem 3.3, eq. (3)) is an interesting new phenomenon: the estimator does not underestimate an upper bound, and in the OT case it achieves sqrt(n)-type one-sided rates despite the n^{2/d} two-sided minimax rate. The paper also delivers a computable sieve formulation and a simulation illustration. The main chain from moment encoding to duality to exact penalization to sieve approximation is coherent and largely proved in the appendix.
major comments (2)
- The step establishing delta_a^n = O_P(1/sqrt(n)) uniformly in P for the independence restrictions of Section 2.4 is not proved. The class F is asserted to be bounded in H^{d2/2+nu}(R^{d2}) 'by continuity' and is then declared uniformly Donsker by citing Corollary 2 in Nickl and Potscher (2007). The proof does not verify two prerequisites: (i) a specific norm computation or a detailed continuity argument for the Sobolev norm of t2 -> xi(t2)K(s,(t1,t2)) as a function of (s,t1); for the exponential kernel one needs an explicit mollifier argument because exp(s2' t2) is not in L2(R^{d2}); and (ii) that the bracketing entropy bound is uniform in the law P2 of T2, since a bound with respect to Lebesgue measure does not automatically transfer to arbitrary probability measures. Because Theorem 3.3's rate depends on this Donsker step, the proof is incomplete. Please supply a self-contained lemma with a full verification of the uniform Donsker property.
- [Abstract and Remark 3.3 (scope)] The abstract lists '(conditional) independence' among the structural conditions and claims the approach yields a 'sqrt(n)-uniformly valid' estimator for the sharp bounds. Remark 3.3, however, states that conditional independence is not covered by Theorem 3.3 and that estimating the conditional MGFs in Lemma 2.3 would only give a nonparametric rate for delta_a^n. Thus the headline sqrt(n)/uniform-validity result applies to identified marginals and unconditional independence, not to conditional independence. The abstract and introduction should be revised to state this restriction, or the paper should extend Theorem 3.3 to conditional independence with the appropriate nonparametric rates.
minor comments (3)
- [Lemma 2.2] The text contains a duplicated word: 'for for s in H, lambda-a.e.' Also, the proof in Appendix A.2 uses the notation 'E_{P~}[integral K dP~_{I2}]', which is confusing; the argument is correct but should be written with explicit integrals over P_{I1} and P_{I2} to avoid ambiguity.
- [Theorem 3.3] The displayed relation 'OP(gamma_n/sqrt(n)) <= beta_{J_n}(theta_hat;gamma_n) - beta(theta) <= OP(gamma_n/sqrt(n)) + oP(1)' is notationally nonstandard. Combined with (3) the intended meaning appears to be that the difference is at most O_P(gamma_n/sqrt(n)) from above and bounded away from zero at the same rate from below; please clarify this statement, since the lower bound in Theorem 3.2 is -gamma(delta_a+delta_c)-delta_b, not a positive lower bound.
- [Section 5 (simulation)] The simulation fixes gamma = 5 and J = 3, while the asymptotic results require gamma_n to diverge and kappa_n to grow with n. A short discussion of how the user should choose gamma and J in finite samples would be helpful.
Circularity Check
No circularity found; the derivation is self-contained and the self-citation to Voronin (2025) is not load-bearing.
full rationale
The paper's central derivation chain does not assume its target as an input. The identification lemmas (Lemmas 2.1-2.3) reduce marginal and (conditional) independence restrictions to moment equalities using characteristic kernels and MGFs; these are equivalences proved from injectivity of the kernel embedding, not from the target bound. Theorem 3.1 establishes exact penalization with fixed w=1; its proof in Appendix B.1 is self-contained, building on Shapiro (1998) duality and the fact that dual solutions are probability measures under Assumption 1.ii. The citation to Voronin (2025) is used only as an 'exact penalization idea' and is superseded by the paper's own proof, so it is not load-bearing. Theorem 3.2 is a deterministic perturbation bound whose sieve-bias term is controlled by Jackson-type inequalities from Ditzian-Totik and Kolomoitsev-Tikhonov, with no fitted input being relabeled as a prediction. Theorem 3.3's rates combine this bound with empirical-process estimates; the uniform Donsker step relies on Nickl and Pötscher (2007), an external result, and Assumption 4 only imposes smoothness conditions rather than assuming the target convergence. The compactification bias Δ(θ,γ) is characterized separately in Proposition 4.1 and Lemma 4.1, not assumed to vanish. The self-citation in Remark 3.1 merely contrasts the fixed penalty level with the simple-LP case and is not used as evidence for a claim. Concerns about the unverified uniform Donsker step or the narrower scope for conditional independence noted in Remark 3.3 are correctness or scope issues, not circularity: nothing in the paper's equations makes the claimed bound equal to its assumptions by construction. Therefore, no significant circularity is present.
Assumptions & free parameters
free parameters (4)
- norm bound gamma_n
- sieve degree kappa_n
- kernel choice and smoothness
- asymmetry tuning eta in (0,1]
assumptions (6)
- standard math Riesz representation theorem: the dual of C(T) is the space of finite Borel measures
- domain assumption T is compact and a, b are continuous (Assumption 1.i)
- ad hoc to paper Existence of atom s* with a(t)(s*)=c(s*)=1 (Assumption 1.ii)
- ad hoc to paper Slater condition in the primary program (P)
- domain assumption Sobolev smoothness of c and a(t) (Assumption 3)
- domain assumption Uniform Donsker property of the kernel function class
Cite this review
Pith. "Pith review of Generalized Optimal Transport." pith.science (2026). https://pith.science/paper/NYTZAZZQ
@misc{pith2026250722422,
author = {Pith},
title = {Pith review of: Generalized Optimal Transport},
year = {2026},
howpublished = {\url{https://pith.science/paper/NYTZAZZQ}},
note = {Machine review of arXiv:2507.22422}
}
abstract
Many causal and structural parameters in economics can be identified and estimated by computing the value of an optimization program over all distributions consistent with the model and the data. Existing tools apply when the data is discrete, or when only disjoint marginals of the distribution are identified, which is restrictive in many applications. We develop a general framework that yields sharp bounds on a linear functional of the unknown true distribution under i) an arbitrary collection of identified joint subdistributions and ii) structural conditions, such as (conditional) independence. We encode the identification restrictions as a continuous collection of moments of characteristic kernels, and use duality and approximation theory to rewrite the infinite-dimensional program over Borel measures as a finite-dimensional program that is simple to compute. Our approach yields a consistent estimator that is $\sqrt{n}$-uniformly valid for the sharp bounds. In the special case of empirical optimal transport with Lipschitz cost, where the minimax rate is $n^{2/d}$, our method yields a uniformly consistent estimator with an asymmetric rate, converging at $\sqrt{n}$ uniformly from one side.
Figures
Forward citations
Cited by 1 Pith paper
-
Inference in partially identified moment models via regularized optimal transport
Entropic optimal transport gives computable bounds and bootstrap confidence regions for partially identified moment models, but the promised uniform CLT is not proved in the appendix and fixed regularization changes t...
Reference graph
Works this paper leans on
-
[1]
The Lifetime Impacts of the New Deal’s Youth Employment Program,
Aizer, A., N. Early, S. Eli, G. Imbens, K. Lee, A. Lleras-Muney, and A. Strand (2024): “The Lifetime Impacts of the New Deal’s Youth Employment Program,”The Quarterly Journal of Economics, 139, 2579–2635. Aliprantis, C. and K. Border (2007): Infinite Dimensional Analysis: A Hitchhiker’s Guide, Springer. Andrews, D. W. and X. Shi (2013): “Inference based o...
arXiv 2024
-
[2]
The conclusion then follows by e.g
to a tight, Borel measurable version of the Brownian bridgeGP, uniformly in P∈P . The conclusion then follows by e.g. bounding√nδa n asymptotically with a supremum over a Brownian bridge. The argument forδa n is identical ifK is the exponential kernel, observing that it gives rise to analytic functions. The result forˆc is established similarly. The resul...
work page 1999
-
[5]
■ A.3. Proof of Lemma 3 The argument is the same, as in the Proof of Lemma 2, where one observes that under anyP∈M 1(U), by law of iterated expectation, the joint MGF of(UIj)3 j=1 rewrites as MPI1∪I2∪I3 (s) = EP [exp( 3∑ j=2 s′ IjUIj)EP [exp(s′ I1UI1)|UI2∪I3]], and ifUI1 |=PUI2|UI3, then EP [exp( 3∑ j=2 s′ IjUIj)EP [exp(s′ I1UI1)|UI2∪I3]] = EP [exp( 3∑ j=...
work page 2002
-
[6]
Per usual arguments involved in the constructions of a Bochner integral (see Hytönen et al. (2016)), each of the sequences ∫ fndµ±, ∫ fndν±, ∫ fnd(ν +µ)± then converges to a respective element inV. Because linearity holds for simple functions, we have ∫ fndµ± + ∫ fndν± = ∫ fnd(ν +µ)± for eachn∈ N. The continuity ofg(x,y,z ) =x +y−z in V 3 yields the resul...
work page 2016
-
[7]
Finally, the discussion on page 15 in Hytönen et al
So,A∗∈L (Y∗,X∗). Finally, the discussion on page 15 in Hytönen et al. (2016) allows to conclude that, indeed, ⟨A∗y∗,x⟩ =⟨ ∫ a(t)dµ(t),x⟩ = ∫ ⟨a(t),x⟩dµ(t) =⟨y∗,Ax⟩. Finally, it follows from equation 2.6 in Shapiro (1998)12, that the dual to inf x∈X ⟨c,x⟩, s.t.: b−Ax∈K, is sup y∗∈K− ⟨y∗,b⟩, s.t.: A∗y∗ =c. Plugging in the definitions yields the claim of the...
work page 2016
-
[8]
Showing this amounts to proving that∃x∈X, such that ⟨a(t),x⟩− b(t)> 0∀t∈T . 11if µ is a signed measure, define the integral as ∫ fdµ+− ∫ fdµ− 12This equation contains a typo:min should be max, see the rest of the chapter and the enclosing book. 28 Observe that, for anyx∈X,⟨a(t),x⟩ =x(s∗)+ ∫ S\{s∗}a(t)(s)x(s)dν(s). Considerx = 1{s∗}x∗. As s∗∈S , suchx is m...
work page 1998
-
[9]
We have L∞(x,θ ) =⟨c,x⟩ for anyx∈ ΘI(θ)
Denote the feasible set of(P ) by ΘI(θ). We have L∞(x,θ ) =⟨c,x⟩ for anyx∈ ΘI(θ). Because ΘI(θ)⊆X, we have β∞(θ) = inf x∈X L∞(x;θ)≤ inf x∈ΘI(θ) L∞(x;θ) = inf x∈ΘI(θ) ⟨c,x⟩ =β(θ). (8) Now suppose thatA∗(θ) is non-empty, which is equivalent to Assumption 1.iii by the above result. By definition of the dual program (see e.g. Shapiro (1998)), the dual problem...
work page 1998
-
[10]
Part ii) follows directly from the definition of∆(·)
BecauseL∞(x;θ) is a continuous convex functional of x∈X, andX is a Hilbert space (thus reflexive), part i) follows from Proposition 1.2 on p.35 in Ekeland and Temam (1999), which is a consequence of Banach-Alaoglu Theorem. Part ii) follows directly from the definition of∆(·). For Part iii), consider the problem in(1). By Theorem 3.1, there exists{xn}⊂ X, ...
work page 1999
Show all 14 references
-
[11]
Because the cube is a Lipschitz domain, we can use the Sobolev extension theorem to extendf to ˜f on Rd, with|| ˜f||Ws(Rd)≤C||f||Ws((−1,1)d) for an absolute constantC >0
Proof. Because the cube is a Lipschitz domain, we can use the Sobolev extension theorem to extendf to ˜f on Rd, with|| ˜f||Ws(Rd)≤C||f||Ws((−1,1)d) for an absolute constantC >0. Since ˜f coincides withf on (−1, 1)d, it follows thatωs(f,t )2≤ωs( ˜f,t )2, where the latter is upp...
2020
-
[12]
Thus, for anys,t 1 the mapt2→K(s, (t1,t 2)) is in Hd2/2+ν(Rd2). By continuity ofK(s,t ), the non-empty set F≡{ t2→ξ(t2)K(s, (t1,t 2)) :s∈Hd(0),t 1∈T 1}⊆ Hd2/2+ν(Rd2) is bounded for an appropriate analytic mollifierξ(t2)∈ [0, 1] that vanishes onT ε 2\T 2, and becomes ξ(t2) = 1 ...
2007
-
[14]
First, observe that we can treatA as an operatorA :X→L 2(λ)
It will suffice to concentrate on a single-coordinateγ, as an extension to the procedure in Section 3.5 is straightforward. First, observe that we can treatA as an operatorA :X→L 2(λ). By Assumption 1.i, and because an equivalence class in anyL2 space has at most one continuou...
2023
-
[173]
The multistochastic Monge–Kantorovich problem,
——— (2022): “The multistochastic Monge–Kantorovich problem,” Journal of Mathematical Analysis and Applications, 506, 125666. Gu, J., T. M. Russell, and T. Stringham (2025): “Counterfactual Identification and Latent Space Enumeration in Discrete Outcome Models,”The Review of Ec...
2022 arXiv
-
[1192]
Inference on causal and structural 21 parameters using many moment inequalities,
Chernozhukov, V., D. Chetverikov, and K. Kato (2019): “Inference on causal and structural 21 parameters using many moment inequalities,”The Review of Economic Studies, 86, 1867–1900. Chernozhukov, V., H. Hong, and E. Tamer (2007): “Estimation and Confidence Regions for Paramet...
2019
-
[2024]
Set Identification in Models with Multiple Equilibria,
Galichon, A. (2016): Optimal transport methods in economics, Princeton University Press. Galichon, A. and M. Henry (2011): “Set Identification in Models with Multiple Equilibria,” The Review of Economic Studies, 78, 1264–1298. Gladkov, N. A., A. V. Kolesnikov, and A. P. Zimin ...
2016
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.