Pith. sign in

REVIEW 2 major objections 3 minor 1 cited by

Generalized Optimal Transport

T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A generalized optimal transport program claims to deliver sharp bounds at parametric √n speed for partial identification problems with overlapping marginals and independence restrictions.

desk verdict A serious new framework for partial identification via characteristic-kernel moment restrictions, with a real gap in the √n proof for independence and an abstract that overstates the conditional-independence scope. read the letter →

arxiv 2507.22422 v1 pith:NYTZAZZQ submitted 2025-07-30 econ.EM math.STstat.TH

classification econ.EMmath.STstat.TH MSC 62G0590C3449Q2262P20
keywords generalizedoptimaltransportpartialidentificationsharpboundscharacteristickernelssemi-infiniteprogrammingdualityuniforminferencetreatmenteffects
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a very broad class of partial identification problems—any parameter that is the expectation of a known function under an unknown distribution, constrained by an arbitrary collection of identified joint subdistributions and by independence or conditional-independence assumptions—can be solved by one optimization program, called Generalized Optimal Transport (GOT). The paper encodes each identification restriction as a continuum of moment equalities against characteristic kernels, then replaces the infinite-dimensional program over probability measures by a finite-dimensional program over polynomials. The main payoff is an estimator of the sharp upper or lower bound that is uniformly valid from one side at the parametric √n rate: it never underestimates an upper bound in large samples, even in settings where two-sided estimation is known to suffer from a curse of dimensionality. If correct, the framework gives a single computational template for combining experimental and observational datasets, matching models, entry games, auctions, and network models.

What carries the argument

The objects that carry the argument are characteristic kernels: integral operators P ↦ ∫ K(s,t)dP(t) that are injective on probability measures, so that equality of kernel means becomes equality of distributions. This converts identified marginals of subvectors, and unconditional or conditional independence, into a continuum of moment equalities E_P[a(T)(s)] = c(s). The resulting infinite-dimensional linear program is rewritten, through Banach-space duality and exact penalization with a fixed penalty level, as an unconstrained convex program over L2, and then projected onto spaces of polynomials on each cube. The projection is tractable because the orthogonal projection operator is self-adjoint, so the sieve bias factors through the projection applied to the known functions c and a(t), yielding a Jackson-type bound $Cκ^{{-r}}$ under Sobolev smoothness; a compactification term Δ(θ,γ) remains, and its rate controls the slower, upward approach to the bound.

What would settle it

Take a feasible problem with identified moment functions c and a(t) that lie outside the Sobolev class of Assumption 3, compute the sieve gap δ_p = ||(I-p_J)c|| + max_t ||(I-p_J)a(t)|| as the polynomial degree J grows, and check whether it decays at the predicted polynomial rate; then simulate the one-sided coverage of the inequality (3). If the sieve gap still decays polynomially, or if the one-sided uniform validity holds despite the nonsmooth moments, the smoothness assumption is not load-bearing; if both fail, the theorem's rate mechanism is confirmed to depend on it.

Watch

Extended reading notes

Core claim

The central claim is that the sharp upper bound β(θ) = sup_P E_P[b(T)] over all Borel probability measures on a compact set T satisfying a ν-almost-everywhere continuum of moment equalities E_P[a(T)(s)] = c(s) is captured by a penalized L2 program, and that projecting that program onto polynomials of degree κ yields a finite semi-infinite linear program whose approximation error is bounded by $Cκ^{{-r}}$ when c and a(t) have Sobolev smoothness of order r. Theorem 3.2 turns this into a deterministic perturbation bound in the estimation errors δ_a, δ_b, δ_c and the sieve gap δ_p. Theorem 3.3 then states that, when the restrictions are estimated by sample analogues, the resulting estimator satisfies $$O_P(\gamma_n/\sqrt{n}) \le \beta_{J_n}(\hat{\$\theta$};\gamma_n) - \$\beta$(\$\theta$) \le O_P(\gamma_n/\sqrt{n}) + o_P(1)$$ uniformly over the class, so the bound is approached from the conservative side at a parametric rate; the upward error is governed by a compactification bias Δ(θ,γ) that is related to the ill-posed inverse problem and vanishes as γ grows. In the optimal-transport special case, this asymmetry is exactly what permits one-sided √n validity despite the two-sided minimax rate $n^{{2/d}}$.

Load-bearing premise

The clean √n result rests on the identified moment functions being smooth enough that a polynomial sieve approximates them at the known rate $Cκ^{{-r}}$; if that smoothness is absent, the sieve bias need not vanish at the assumed rate and the one-sided uniform √n validity is lost.

Editorial extensions

If this is right

  • Any problem fitting the format—identified marginals, possibly overlapping, plus independence restrictions, with continuous or discrete variables—gets a single computational procedure instead of bespoke linear-programming or optimal-transport solvers.
  • The estimator is conservative by construction: for an upper bound it converges from below at rate √n/γ_n uniformly, so it does not understate the identified set's upper limit in large samples.
  • For empirical optimal transport with a Lipschitz cost, one obtains a one-sided √n (up to logs) estimator despite the two-sided minimax rate n^{2/d}, and tuning a parameter η trades the two one-sided rates against each other.
  • A growing number of discrete moment conditions can be added—for example, spline moments of identified bounded functions—at rate (γ_n + γ_n log l_n)/√n, as long as the added moments are redundant for identification.
  • Conditional-independence restrictions can be embedded by the same construction, though the paper notes they require nonparametric estimation of conditional moment generating functions and therefore only nonparametric rates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit consequence is that problems of combining experimental and observational data, such as the long-term treatment-effect example in the paper, become a routine computation once the variable vector, support, and characteristic kernels are specified.
  • The one-sided nature of the estimator suggests a practical reporting rule: use the conservative bound as the policy number, and treat the vanishing upward bias as a diagnostic for how much regularization (the choice of γ) is doing.
  • A natural extension the paper leaves open is fully data-driven choice of the kernel, sieve order, and tuning parameter; the theorems only require κ_n to diverge fast enough relative to n^{1/(2r*)}.
  • The link to the ill-posed inverse problem hints that the framework could be dualized to deliver inference for nonparametric instrumental-variable-like objects, but that direction is not developed in the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper develops a method, Generalized Optimal Transport (GOT), for computing sharp bounds on a linear functional E_P[b(T)] when the econometric model identifies only certain joint subdistributions of T and/or imposes (conditional) independence restrictions. The identification restrictions are encoded as a continuum of moment equalities using characteristic kernels. The resulting infinite-dimensional linear program is dualized, exactly penalized with a fixed penalty, and solved over a finite-dimensional polynomial sieve. The main theoretical results are a deterministic perturbation bound (Theorem 3.2), a one-sided uniform convergence rate for the sieve estimator (Theorem 3.3), an extension to many discrete moments (Theorem 3.5), and an analysis of the compactification bias Delta(theta; gamma), including a sharp characterization in the optimal transport special case (Section 4).

Significance. If the results hold, this is a valuable contribution that unifies and extends LP-based partial identification and empirical optimal transport. The one-sided uniform validity result (Theorem 3.3, eq. (3)) is an interesting new phenomenon: the estimator does not underestimate an upper bound, and in the OT case it achieves sqrt(n)-type one-sided rates despite the n^{2/d} two-sided minimax rate. The paper also delivers a computable sieve formulation and a simulation illustration. The main chain from moment encoding to duality to exact penalization to sieve approximation is coherent and largely proved in the appendix.

major comments (2)
  1. The step establishing delta_a^n = O_P(1/sqrt(n)) uniformly in P for the independence restrictions of Section 2.4 is not proved. The class F is asserted to be bounded in H^{d2/2+nu}(R^{d2}) 'by continuity' and is then declared uniformly Donsker by citing Corollary 2 in Nickl and Potscher (2007). The proof does not verify two prerequisites: (i) a specific norm computation or a detailed continuity argument for the Sobolev norm of t2 -> xi(t2)K(s,(t1,t2)) as a function of (s,t1); for the exponential kernel one needs an explicit mollifier argument because exp(s2' t2) is not in L2(R^{d2}); and (ii) that the bracketing entropy bound is uniform in the law P2 of T2, since a bound with respect to Lebesgue measure does not automatically transfer to arbitrary probability measures. Because Theorem 3.3's rate depends on this Donsker step, the proof is incomplete. Please supply a self-contained lemma with a full verification of the uniform Donsker property.
  2. [Abstract and Remark 3.3 (scope)] The abstract lists '(conditional) independence' among the structural conditions and claims the approach yields a 'sqrt(n)-uniformly valid' estimator for the sharp bounds. Remark 3.3, however, states that conditional independence is not covered by Theorem 3.3 and that estimating the conditional MGFs in Lemma 2.3 would only give a nonparametric rate for delta_a^n. Thus the headline sqrt(n)/uniform-validity result applies to identified marginals and unconditional independence, not to conditional independence. The abstract and introduction should be revised to state this restriction, or the paper should extend Theorem 3.3 to conditional independence with the appropriate nonparametric rates.
minor comments (3)
  1. [Lemma 2.2] The text contains a duplicated word: 'for for s in H, lambda-a.e.' Also, the proof in Appendix A.2 uses the notation 'E_{P~}[integral K dP~_{I2}]', which is confusing; the argument is correct but should be written with explicit integrals over P_{I1} and P_{I2} to avoid ambiguity.
  2. [Theorem 3.3] The displayed relation 'OP(gamma_n/sqrt(n)) <= beta_{J_n}(theta_hat;gamma_n) - beta(theta) <= OP(gamma_n/sqrt(n)) + oP(1)' is notationally nonstandard. Combined with (3) the intended meaning appears to be that the difference is at most O_P(gamma_n/sqrt(n)) from above and bounded away from zero at the same rate from below; please clarify this statement, since the lower bound in Theorem 3.2 is -gamma(delta_a+delta_c)-delta_b, not a positive lower bound.
  3. [Section 5 (simulation)] The simulation fixes gamma = 5 and J = 3, while the asymptotic results require gamma_n to diverge and kappa_n to grow with n. A short discussion of how the user should choose gamma and J in finite samples would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the derivation is self-contained and the self-citation to Voronin (2025) is not load-bearing.

full rationale

The paper's central derivation chain does not assume its target as an input. The identification lemmas (Lemmas 2.1-2.3) reduce marginal and (conditional) independence restrictions to moment equalities using characteristic kernels and MGFs; these are equivalences proved from injectivity of the kernel embedding, not from the target bound. Theorem 3.1 establishes exact penalization with fixed w=1; its proof in Appendix B.1 is self-contained, building on Shapiro (1998) duality and the fact that dual solutions are probability measures under Assumption 1.ii. The citation to Voronin (2025) is used only as an 'exact penalization idea' and is superseded by the paper's own proof, so it is not load-bearing. Theorem 3.2 is a deterministic perturbation bound whose sieve-bias term is controlled by Jackson-type inequalities from Ditzian-Totik and Kolomoitsev-Tikhonov, with no fitted input being relabeled as a prediction. Theorem 3.3's rates combine this bound with empirical-process estimates; the uniform Donsker step relies on Nickl and Pötscher (2007), an external result, and Assumption 4 only imposes smoothness conditions rather than assuming the target convergence. The compactification bias Δ(θ,γ) is characterized separately in Proposition 4.1 and Lemma 4.1, not assumed to vanish. The self-citation in Remark 3.1 merely contrasts the fixed penalty level with the simple-LP case and is not used as evidence for a claim. Concerns about the unverified uniform Donsker step or the narrower scope for conditional independence noted in Remark 3.3 are correctness or scope issues, not circularity: nothing in the paper's equations makes the claimed bound equal to its assumptions by construction. Therefore, no significant circularity is present.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

The framework introduces no new physical entities and estimates no free parameters from data. Its load-bearing inputs are the characteristic-kernel construction, the atom normalization, the Sobolev smoothness of the moments, and standard duality/approximation results. The main tuning parameters, gamma_n, kappa_n, kernel choice, and eta, are chosen by the researcher and are not fitted to data.

free parameters (4)
  • norm bound gamma_n
    Researcher-chosen sequence that defines the compact ball B_gamma; Theorem 3.3 requires gamma_n/√n to diverge and gamma_n ≤ √n M w.h.p.
  • sieve degree kappa_n
    Researcher-chosen polynomial degree for each cube; the approximation bias is O(kappa_n^{-r}) and the rates require kappa_n/n^{1/(2r*)} to diverge.
  • kernel choice and smoothness
    Exponential or Matérn kernels are used; the Sobolev smoothness r* and the entropy bounds depend on this choice, affecting the allowed rates.
  • asymmetry tuning eta in (0,1]
    In Corollary 1, eta chooses the tradeoff between the fast (from below) and slow (from above) sides in the OT special case.
assumptions (6)
  • standard math Riesz representation theorem: the dual of C(T) is the space of finite Borel measures
    Used in Appendix B.1 to identify the dual problem (D).
  • domain assumption T is compact and a, b are continuous (Assumption 1.i)
    Ensures maxima are attained and the dual space identification holds; compactness is noted as not easily relaxed.
  • ad hoc to paper Existence of atom s* with a(t)(s*)=c(s*)=1 (Assumption 1.ii)
    This normalization makes every feasible dual measure a probability measure, enabling w=1 exact penalization.
  • ad hoc to paper Slater condition in the primary program (P)
    Guaranteed by the atom construction in the proof of Theorem 3.1; needed for strong duality.
  • domain assumption Sobolev smoothness of c and a(t) (Assumption 3)
    Yields the Jackson-type polynomial approximation bound; if moments are nonsmooth the sieve bias may not vanish at the stated rate.
  • domain assumption Uniform Donsker property of the kernel function class
    Used in the proof of Theorem 3.3; proved for Matérn kernels via bracketing entropy, asserted for exponential kernels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generalized Optimal Transport." pith.science (2026). https://pith.science/paper/NYTZAZZQ

@misc{pith2026250722422,
  author       = {Pith},
  title        = {Pith review of: Generalized Optimal Transport},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NYTZAZZQ}},
  note         = {Machine review of arXiv:2507.22422}
}
abstract

Many causal and structural parameters in economics can be identified and estimated by computing the value of an optimization program over all distributions consistent with the model and the data. Existing tools apply when the data is discrete, or when only disjoint marginals of the distribution are identified, which is restrictive in many applications. We develop a general framework that yields sharp bounds on a linear functional of the unknown true distribution under i) an arbitrary collection of identified joint subdistributions and ii) structural conditions, such as (conditional) independence. We encode the identification restrictions as a continuous collection of moments of characteristic kernels, and use duality and approximation theory to rewrite the infinite-dimensional program over Borel measures as a finite-dimensional program that is simple to compute. Our approach yields a consistent estimator that is $\sqrt{n}$-uniformly valid for the sharp bounds. In the special case of empirical optimal transport with Lipschitz cost, where the minimax rate is $n^{2/d}$, our method yields a uniformly consistent estimator with an asymmetric rate, converging at $\sqrt{n}$ uniformly from one side.

Figures

Figures reproduced from arXiv: 2507.22422 by the authors.

Figure 1
Figure 1. Example (5), KDEs computed over a 100 simulations for oracle (solid) and GOT (dashed), fixed γ = 5, and J = 3. GOT solved using cutting-plane method with tol = 10−6 . of degree J ∈ N for each dimension of T, we obtain the problem βJ (θ; γ) ≡ min (x,y,x∗)∈Bγ,a≥0  a + x ∗+ (6) X (i,j,k)∈0,J3 xijk Z H M1(s1, s2)ψijk(s)ds + yijk Z H M2(s2, s3)ψijk(s)ds , s.t. for T ∈ [0, 1]3 , a + x ∗ ≥ T3 − T2− X (i,j,k)∈0,J3  xijk … view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inference in partially identified moment models via regularized optimal transport

    econ.EM 2025-12 reject novelty 6.0 of 10

    Entropic optimal transport gives computable bounds and bootstrap confidence regions for partially identified moment models, but the promised uniform CLT is not proved in the appendix and fixed regularization changes t...

Reference graph

Works this paper leans on

14 extracted references · 13 canonical work pages · cited by 1 Pith paper

  1. [1]

    The Lifetime Impacts of the New Deal’s Youth Employment Program,

    Aizer, A., N. Early, S. Eli, G. Imbens, K. Lee, A. Lleras-Muney, and A. Strand (2024): “The Lifetime Impacts of the New Deal’s Youth Employment Program,”The Quarterly Journal of Economics, 139, 2579–2635. Aliprantis, C. and K. Border (2007): Infinite Dimensional Analysis: A Hitchhiker’s Guide, Springer. Andrews, D. W. and X. Shi (2013): “Inference based o...

  2. [2]

    The conclusion then follows by e.g

    to a tight, Borel measurable version of the Brownian bridgeGP, uniformly in P∈P . The conclusion then follows by e.g. bounding√nδa n asymptotically with a supremum over a Brownian bridge. The argument forδa n is identical ifK is the exponential kernel, observing that it gives rise to analytic functions. The result forˆc is established similarly. The resul...

  3. [5]

    ■ A.3. Proof of Lemma 3 The argument is the same, as in the Proof of Lemma 2, where one observes that under anyP∈M 1(U), by law of iterated expectation, the joint MGF of(UIj)3 j=1 rewrites as MPI1∪I2∪I3 (s) = EP [exp( 3∑ j=2 s′ IjUIj)EP [exp(s′ I1UI1)|UI2∪I3]], and ifUI1 |=PUI2|UI3, then EP [exp( 3∑ j=2 s′ IjUIj)EP [exp(s′ I1UI1)|UI2∪I3]] = EP [exp( 3∑ j=...

  4. [6]

    (2016)), each of the sequences ∫ fndµ±, ∫ fndν±, ∫ fnd(ν +µ)± then converges to a respective element inV

    Per usual arguments involved in the constructions of a Bochner integral (see Hytönen et al. (2016)), each of the sequences ∫ fndµ±, ∫ fndν±, ∫ fnd(ν +µ)± then converges to a respective element inV. Because linearity holds for simple functions, we have ∫ fndµ± + ∫ fndν± = ∫ fnd(ν +µ)± for eachn∈ N. The continuity ofg(x,y,z ) =x +y−z in V 3 yields the resul...

  5. [7]

    Finally, the discussion on page 15 in Hytönen et al

    So,A∗∈L (Y∗,X∗). Finally, the discussion on page 15 in Hytönen et al. (2016) allows to conclude that, indeed, ⟨A∗y∗,x⟩ =⟨ ∫ a(t)dµ(t),x⟩ = ∫ ⟨a(t),x⟩dµ(t) =⟨y∗,Ax⟩. Finally, it follows from equation 2.6 in Shapiro (1998)12, that the dual to inf x∈X ⟨c,x⟩, s.t.: b−Ax∈K, is sup y∗∈K− ⟨y∗,b⟩, s.t.: A∗y∗ =c. Plugging in the definitions yields the claim of the...

  6. [8]

    11if µ is a signed measure, define the integral as ∫ fdµ+− ∫ fdµ− 12This equation contains a typo:min should be max, see the rest of the chapter and the enclosing book

    Showing this amounts to proving that∃x∈X, such that ⟨a(t),x⟩− b(t)> 0∀t∈T . 11if µ is a signed measure, define the integral as ∫ fdµ+− ∫ fdµ− 12This equation contains a typo:min should be max, see the rest of the chapter and the enclosing book. 28 Observe that, for anyx∈X,⟨a(t),x⟩ =x(s∗)+ ∫ S\{s∗}a(t)(s)x(s)dν(s). Considerx = 1{s∗}x∗. As s∗∈S , suchx is m...

  7. [9]

    We have L∞(x,θ ) =⟨c,x⟩ for anyx∈ ΘI(θ)

    Denote the feasible set of(P ) by ΘI(θ). We have L∞(x,θ ) =⟨c,x⟩ for anyx∈ ΘI(θ). Because ΘI(θ)⊆X, we have β∞(θ) = inf x∈X L∞(x;θ)≤ inf x∈ΘI(θ) L∞(x;θ) = inf x∈ΘI(θ) ⟨c,x⟩ =β(θ). (8) Now suppose thatA∗(θ) is non-empty, which is equivalent to Assumption 1.iii by the above result. By definition of the dual program (see e.g. Shapiro (1998)), the dual problem...

  8. [10]

    Part ii) follows directly from the definition of∆(·)

    BecauseL∞(x;θ) is a continuous convex functional of x∈X, andX is a Hilbert space (thus reflexive), part i) follows from Proposition 1.2 on p.35 in Ekeland and Temam (1999), which is a consequence of Banach-Alaoglu Theorem. Part ii) follows directly from the definition of∆(·). For Part iii), consider the problem in(1). By Theorem 3.1, there exists{xn}⊂ X, ...

Show all 14 references
  1. [11]

    Because the cube is a Lipschitz domain, we can use the Sobolev extension theorem to extendf to ˜f on Rd, with|| ˜f||Ws(Rd)≤C||f||Ws((−1,1)d) for an absolute constantC >0

    Proof. Because the cube is a Lipschitz domain, we can use the Sobolev extension theorem to extendf to ˜f on Rd, with|| ˜f||Ws(Rd)≤C||f||Ws((−1,1)d) for an absolute constantC >0. Since ˜f coincides withf on (−1, 1)d, it follows thatωs(f,t )2≤ωs( ˜f,t )2, where the latter is upp...

  2. [12]

    Thus, for anys,t 1 the mapt2→K(s, (t1,t 2)) is in Hd2/2+ν(Rd2). By continuity ofK(s,t ), the non-empty set F≡{ t2→ξ(t2)K(s, (t1,t 2)) :s∈Hd(0),t 1∈T 1}⊆ Hd2/2+ν(Rd2) is bounded for an appropriate analytic mollifierξ(t2)∈ [0, 1] that vanishes onT ε 2\T 2, and becomes ξ(t2) = 1 ...

  3. [14]

    First, observe that we can treatA as an operatorA :X→L 2(λ)

    It will suffice to concentrate on a single-coordinateγ, as an extension to the procedure in Section 3.5 is straightforward. First, observe that we can treatA as an operatorA :X→L 2(λ). By Assumption 1.i, and because an equivalence class in anyL2 space has at most one continuou...

  4. [173]

    The multistochastic Monge–Kantorovich problem,

    ——— (2022): “The multistochastic Monge–Kantorovich problem,” Journal of Mathematical Analysis and Applications, 506, 125666. Gu, J., T. M. Russell, and T. Stringham (2025): “Counterfactual Identification and Latent Space Enumeration in Discrete Outcome Models,”The Review of Ec...

  5. [1192]

    Inference on causal and structural 21 parameters using many moment inequalities,

    Chernozhukov, V., D. Chetverikov, and K. Kato (2019): “Inference on causal and structural 21 parameters using many moment inequalities,”The Review of Economic Studies, 86, 1867–1900. Chernozhukov, V., H. Hong, and E. Tamer (2007): “Estimation and Confidence Regions for Paramet...

  6. [2024]

    Set Identification in Models with Multiple Equilibria,

    Galichon, A. (2016): Optimal transport methods in economics, Princeton University Press. Galichon, A. and M. Henry (2011): “Set Identification in Models with Multiple Equilibria,” The Review of Economic Studies, 78, 1264–1298. Gladkov, N. A., A. V. Kolesnikov, and A. P. Zimin ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.