{"id":"62e61594-aa36-43f2-ad5b-e3e398fa05a3","arxiv_id":"2507.22422","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Generalized Optimal Transport encodes identified marginals and independence as kernel moments, then uses duality and polynomial sieves to compute sharp bounds with a one-sided sqrt(n)-consistent estimator.","lead":"Economists often need bounds on causal effects when only some parts of the data distribution are observed. This paper develops a general optimization framework, called Generalized Optimal Transport, that computes such bounds with continuous variables, overlapping marginals, and independence restrictions, and it provides an estimator that is conservative at near-sqrt(n) rates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.3's √n-uniform validity rests on an unproved uniform Donsker step for estimated moment functions in the independence case, and conditional independence is explicitly outside the √n theory; the abstract overstates the scope.","rationale":"The reader's weakest_assumption (Assumption 3) is less damaging than it first appears: for the exponential and Matérn kernels used in Sections 2.3-2.4, c(s) and a(t)(s) are analytic or Sobolev-smooth in s, so Assumption 3 holds with the stated r. The genuinely load-bearing point is Theorem 3.3's dependence on an unproved uniform Donsker claim. The proof of δa_n = OP(1/√n) for the independence restriction (Appendix B.4) is the only place where the √n rate is shown for a structural condition beyond identified marginals, and it is asserted rather than demonstrated. The class F|T2 is a parameter-dependent family, and the cited entropy corollary requires a fixed Sobolev ball; the paper does not show the boundedness in Sobolev norm uniformly over the compact parameter set. This is fixable, but until verified the theorem is incomplete. Additionally, the abstract's 'conditional independence' is contradicted by Remark 3.3, which states that only nonparametric rates would follow; the paper is transparent about this, but the scope claim needs revision. These concerns do not invalidate the core duality/sieve argument for identified marginals, but they support the reader's CONDITIONAL verdict. A complete proof of the Donsker step would resolve the main uncertainty.","tokens_in":27901,"tokens_out":26137,"duration_ms":299295,"concrete_test":"Write out the bracketing entropy verification for F|T2: for the exponential kernel K(s,t) = exp(s' t) and a Matérn kernel with smoothness ν, compute the Sobolev norm ||ξ(t2)K((s1,s2),(t1,·))||_{W^{d2/2+ν,2}} as a function of (s1,s2,t1) and check it is uniformly bounded on the compact parameter space; then apply Corollary 2 in Nickl and Pötscher (2007) to obtain log N_{[]}(ε,F|T2) = O(ε^{-d2/(d2/2+ν)}) and verify the Donsker entropy integral converges. If the bound fails, Theorem 3.3's √n rate for independence restrictions is not established. Separately, derive the convergence rate of the conditional MGF estimator in Lemma 2.3 to confirm it is slower than n^{-1/2}; this would require qualifying the abstract's conditional-independence claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of √n one-sided uniform validity is delivered by Theorem 3.3. Its proof requires, uniformly in P, δa_n = OP(1/√n) for the Section 2.4 independence restrictions. Appendix B.4 asserts this by claiming the class F = { t2 → K((s1,s2),(t1,t2)) : s ∈ Hd(0), t1 ∈ T1 } is uniformly Donsker. The proof cites Corollary 2 in Nickl and Pötscher (2007) for a bracketing entropy bound, but it does not verify the two prerequisites: (i) that F is contained in a fixed Sobolev ball of order d2/2+ν, and (ii) that the entropy bound is uniform in P. The parameter set (s,t1) is compact, but the map to the Sobolev norm is only asserted to be bounded 'by continuity'; no norm computation is given for the exponential or Matérn kernels. If this Donsker step fails, the rate for δa_n is not √n and the one-sided validity for independence restrictions is unsupported. Furthermore, Remark 3.3 explicitly excludes conditional independence (Section 2.5) from Theorem 3.3, requiring nonparametric rates, while the abstract lists '(conditional) independence' as part of the scope. Thus the headline √n result covers a strictly smaller class than announced.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a method, Generalized Optimal Transport (GOT), for computing sharp bounds on a linear functional E_P[b(T)] when the econometric model identifies only certain joint subdistributions of T and/or imposes (conditional) independence restrictions. The identification restrictions are encoded as a continuum of moment equalities using characteristic kernels. The resulting infinite-dimensional linear program is dualized, exactly penalized with a fixed penalty, and solved over a finite-dimensional polynomial sieve. The main theoretical results are a deterministic perturbation bound (Theorem 3.2), a one-sided uniform convergence rate for the sieve estimator (Theorem 3.3), an extension to many discrete moments (Theorem 3.5), and an analysis of the compactification bias Delta(theta; gamma), including a sharp characterization in the optimal transport special case (Section 4).","tokens_in":28181,"tokens_out":15594,"duration_ms":170898,"significance":"If the results hold, this is a valuable contribution that unifies and extends LP-based partial identification and empirical optimal transport. The one-sided uniform validity result (Theorem 3.3, eq. (3)) is an interesting new phenomenon: the estimator does not underestimate an upper bound, and in the OT case it achieves sqrt(n)-type one-sided rates despite the n^{2/d} two-sided minimax rate. The paper also delivers a computable sieve formulation and a simulation illustration. The main chain from moment encoding to duality to exact penalization to sieve approximation is coherent and largely proved in the appendix.","major_comments":[{"comment":"The step establishing delta_a^n = O_P(1/sqrt(n)) uniformly in P for the independence restrictions of Section 2.4 is not proved. The class F is asserted to be bounded in H^{d2/2+nu}(R^{d2}) 'by continuity' and is then declared uniformly Donsker by citing Corollary 2 in Nickl and Potscher (2007). The proof does not verify two prerequisites: (i) a specific norm computation or a detailed continuity argument for the Sobolev norm of t2 -> xi(t2)K(s,(t1,t2)) as a function of (s,t1); for the exponential kernel one needs an explicit mollifier argument because exp(s2' t2) is not in L2(R^{d2}); and (ii) that the bracketing entropy bound is uniform in the law P2 of T2, since a bound with respect to Lebesgue measure does not automatically transfer to arbitrary probability measures. Because Theorem 3.3's rate depends on this Donsker step, the proof is incomplete. Please supply a self-contained lemma with a full verification of the uniform Donsker property.","section":null},{"comment":"The abstract lists '(conditional) independence' among the structural conditions and claims the approach yields a 'sqrt(n)-uniformly valid' estimator for the sharp bounds. Remark 3.3, however, states that conditional independence is not covered by Theorem 3.3 and that estimating the conditional MGFs in Lemma 2.3 would only give a nonparametric rate for delta_a^n. Thus the headline sqrt(n)/uniform-validity result applies to identified marginals and unconditional independence, not to conditional independence. The abstract and introduction should be revised to state this restriction, or the paper should extend Theorem 3.3 to conditional independence with the appropriate nonparametric rates.","section":"Abstract and Remark 3.3 (scope)"}],"minor_comments":[{"comment":"The text contains a duplicated word: 'for for s in H, lambda-a.e.' Also, the proof in Appendix A.2 uses the notation 'E_{P~}[integral K dP~_{I2}]', which is confusing; the argument is correct but should be written with explicit integrals over P_{I1} and P_{I2} to avoid ambiguity.","section":"Lemma 2.2"},{"comment":"The displayed relation 'OP(gamma_n/sqrt(n)) <= beta_{J_n}(theta_hat;gamma_n) - beta(theta) <= OP(gamma_n/sqrt(n)) + oP(1)' is notationally nonstandard. Combined with (3) the intended meaning appears to be that the difference is at most O_P(gamma_n/sqrt(n)) from above and bounded away from zero at the same rate from below; please clarify this statement, since the lower bound in Theorem 3.2 is -gamma(delta_a+delta_c)-delta_b, not a positive lower bound.","section":"Theorem 3.3"},{"comment":"The simulation fixes gamma = 5 and J = 3, while the asymptotic results require gamma_n to diverge and kappa_n to grow with n. A short discussion of how the user should choose gamma and J in finite samples would be helpful.","section":"Section 5 (simulation)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for econ.EM. The central idea is promising and the deterministic parts are solid; the main risk is the uniform Donsker step in Appendix B.4, which is likely fixable but must be supplied with a complete proof. If the authors can provide that lemma and correct the abstract's scope claim, I would be willing to accept in a later round."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth taking seriously. The GOT framework—encoding overlapping identified marginals and independence as moment equalities through characteristic kernels, then solving a finite-dimensional semi-infinite LP via duality and exact penalization—is a genuine modeling contribution. It generalizes LP/OT partial identification to continuous variables and overlapping marginals, and the one-sided uniform results are new. Theorem 3.2 (deterministic sieve bound) is clean, and Corollary 1's asymmetric rate for empirical OT is a real insight about how the n^{2/d} curse is one-sided.\n\nThe main soft spot is the proof of Theorem 3.3 for the independence restrictions. The step that delivers δa_n = O_P(1/√n) requires the class of functions t2→K(s,(t1,t2)) to be uniformly bounded in a Sobolev ball of order d2/2+ν. Appendix B.4 asserts this by continuity and cites Harbrecht et al. (2024b) and Nickl–Pötscher (2007), but no norm computation is given. Compactness of (s,t1) alone doesn't give the bound if the map into the Sobolev norm isn't shown continuous. This is probably repairable—the Matérn kernel is translation-invariant and the degeneracy at s1=t1 is manageable—but right now the √n claim for independence is conditional on a calculation the reader can't check.\n\nSecond issue: conditional independence (Section 2.5) is explicitly outside Theorem 3.3. Remark 3.3 says it needs nonparametric rates. The abstract lists \"(conditional) independence\" as if both were covered at √n. That's an overstatement that should be fixed.\n\nMinor: the simulation is a point-identified toy; it doesn't demonstrate the method on a genuinely set-identified problem with overlapping marginals. The computational complexity claim for the semi-infinite LP is plausible but unanalyzed.\n\nNo circularity, no invented entities; the self-citation to Voronin (2025) is legitimate. The appendix is serious and most of the proof chain checks out.\n\nThis paper deserves a real referee. I'd send it out, with an instruction to ask the author for the missing Donsker calculation and a revised abstract.","headline":"A serious new framework for partial identification via characteristic-kernel moment restrictions, with a real gap in the √n proof for independence and an abstract that overstates the conditional-independence scope.","tokens_in":28683,"tokens_out":4468,"would_cite":true,"duration_ms":53462,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","90C34","49Q22","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A generalized optimal transport program claims to deliver sharp bounds at parametric √n speed for partial identification problems with overlapping marginals and independence restrictions.","keywords":["generalized optimal transport","partial identification","sharp bounds","characteristic kernels","semi-infinite programming","duality","uniform inference","treatment effects"],"falsifier":"Take a feasible problem with identified moment functions c and a(t) that lie outside the Sobolev class of Assumption 3, compute the sieve gap δ_p = ||(I-p_J)c|| + max_t ||(I-p_J)a(t)|| as the polynomial degree J grows, and check whether it decays at the predicted polynomial rate; then simulate the one-sided coverage of the inequality (3). If the sieve gap still decays polynomially, or if the one-sided uniform validity holds despite the nonsmooth moments, the smoothness assumption is not load-bearing; if both fail, the theorem's rate mechanism is confirmed to depend on it.","tokens_in":27632,"feed_emoji":"🎯","tokens_out":5756,"duration_ms":68602,"temperature":0.7,"pith_summary":"This paper claims that a very broad class of partial identification problems—any parameter that is the expectation of a known function under an unknown distribution, constrained by an arbitrary collection of identified joint subdistributions and by independence or conditional-independence assumptions—can be solved by one optimization program, called Generalized Optimal Transport (GOT). The paper encodes each identification restriction as a continuum of moment equalities against characteristic kernels, then replaces the infinite-dimensional program over probability measures by a finite-dimensional program over polynomials. The main payoff is an estimator of the sharp upper or lower bound that is uniformly valid from one side at the parametric √n rate: it never underestimates an upper bound in large samples, even in settings where two-sided estimation is known to suffer from a curse of dimensionality. If correct, the framework gives a single computational template for combining experimental and observational datasets, matching models, entry games, auctions, and network models.","feed_headline":"Sharp bounds for any identified subdistribution, at √n speed","feed_subtitle":"Generalized optimal transport turns overlapping marginals and independence restrictions into one computable, conservative program.","key_machinery":"The objects that carry the argument are characteristic kernels: integral operators P ↦ ∫ K(s,t)dP(t) that are injective on probability measures, so that equality of kernel means becomes equality of distributions. This converts identified marginals of subvectors, and unconditional or conditional independence, into a continuum of moment equalities E_P[a(T)(s)] = c(s). The resulting infinite-dimensional linear program is rewritten, through Banach-space duality and exact penalization with a fixed penalty level, as an unconstrained convex program over L2, and then projected onto spaces of polynomials on each cube. The projection is tractable because the orthogonal projection operator is self-adjoint, so the sieve bias factors through the projection applied to the known functions c and a(t), yielding a Jackson-type bound $Cκ^{{-r}}$ under Sobolev smoothness; a compactification term Δ(θ,γ) remains, and its rate controls the slower, upward approach to the bound.","core_discovery":"The central claim is that the sharp upper bound β(θ) = sup_P E_P[b(T)] over all Borel probability measures on a compact set T satisfying a ν-almost-everywhere continuum of moment equalities E_P[a(T)(s)] = c(s) is captured by a penalized L2 program, and that projecting that program onto polynomials of degree κ yields a finite semi-infinite linear program whose approximation error is bounded by $Cκ^{{-r}}$ when c and a(t) have Sobolev smoothness of order r. Theorem 3.2 turns this into a deterministic perturbation bound in the estimation errors δ_a, δ_b, δ_c and the sieve gap δ_p. Theorem 3.3 then states that, when the restrictions are estimated by sample analogues, the resulting estimator satisfies $$O_P(\\gamma_n/\\sqrt{n}) \\le \\beta_{J_n}(\\hat{\\$\\theta$};\\gamma_n) - \\$\\beta$(\\$\\theta$) \\le O_P(\\gamma_n/\\sqrt{n}) + o_P(1)$$ uniformly over the class, so the bound is approached from the conservative side at a parametric rate; the upward error is governed by a compactification bias Δ(θ,γ) that is related to the ill-posed inverse problem and vanishes as γ grows. In the optimal-transport special case, this asymmetry is exactly what permits one-sided √n validity despite the two-sided minimax rate $n^{{2/d}}$.","pith_inferences":["An implicit consequence is that problems of combining experimental and observational data, such as the long-term treatment-effect example in the paper, become a routine computation once the variable vector, support, and characteristic kernels are specified.","The one-sided nature of the estimator suggests a practical reporting rule: use the conservative bound as the policy number, and treat the vanishing upward bias as a diagnostic for how much regularization (the choice of γ) is doing.","A natural extension the paper leaves open is fully data-driven choice of the kernel, sieve order, and tuning parameter; the theorems only require κ_n to diverge fast enough relative to n^{1/(2r*)}.","The link to the ill-posed inverse problem hints that the framework could be dualized to deliver inference for nonparametric instrumental-variable-like objects, but that direction is not developed in the paper."],"forward_implications":["Any problem fitting the format—identified marginals, possibly overlapping, plus independence restrictions, with continuous or discrete variables—gets a single computational procedure instead of bespoke linear-programming or optimal-transport solvers.","The estimator is conservative by construction: for an upper bound it converges from below at rate √n/γ_n uniformly, so it does not understate the identified set's upper limit in large samples.","For empirical optimal transport with a Lipschitz cost, one obtains a one-sided √n (up to logs) estimator despite the two-sided minimax rate n^{2/d}, and tuning a parameter η trades the two one-sided rates against each other.","A growing number of discrete moment conditions can be added—for example, spline moments of identified bounded functions—at rate (γ_n + γ_n log l_n)/√n, as long as the added moments are redundant for identification.","Conditional-independence restrictions can be embedded by the same construction, though the paper notes they require nonparametric estimation of conditional moment generating functions and therefore only nonparametric rates."],"supporting_citations":[{"why":"Supplies the theory of characteristic kernels, which makes equality of kernel means equivalent to equality of distributions.","marker":"Steinwart and Ziegel (2021)"},{"why":"Provides Banach-space duality for conic linear programs, used to derive the dual program and strong duality under the Slater condition.","marker":"Shapiro (2001)"},{"why":"Contributes the exact penalization idea that lets the infinite-dimensional program be rewritten with a fixed penalty level.","marker":"Voronin (2025)"},{"why":"Provides the Jackson-type modulus-of-smoothness estimates behind the polynomial sieve bias bound Cκ^{-r}.","marker":"Ditzian and Totik (2012)"},{"why":"Establishes the minimax two-sided rate n^{2/d} for empirical optimal transport, the benchmark that the paper's one-sided √n result must be reconciled with.","marker":"Manole and Niles-Weed (2024)"},{"why":"Provides the optimal-transport dual-potential bounds and empirical OT setup used in the special-case analysis of Δ(θ,γ).","marker":"Ober-Reynolds (2023)"},{"why":"Supplies bracketing entropy bounds for Sobolev-type function classes used to prove the uniform empirical-process rate for the estimated moment functions.","marker":"Nickl and Pötscher (2007)"},{"why":"Provides the uniform Donsker theorem used to translate entropy bounds into the √n rates for δ_a and δ_c.","marker":"Van Der Vaart and Wellner (1996)"},{"why":"Documents cutting-plane methods for semi-infinite programs, which make the finite-dimensional sieve program computationally practical.","marker":"Reemtsen and Görner (1998)"}],"fun_headline_variants":["√n sharp bounds for any identified subdistribution","Generalized optimal transport yields one-sided √n bounds","Sharp bounds with overlapping marginals at √n speed","One-sided √n consistency for sharp bounds in OT","Conservative √n estimation for arbitrary moment restrictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The clean √n result rests on the identified moment functions being smooth enough that a polynomial sieve approximates them at the known rate $Cκ^{{-r}}$; if that smoothness is absent, the sieve bias need not vanish at the assumed rate and the one-sided uniform √n validity is lost.","fun_headline_variants_meta":{"raw":{"variants":["√n sharp bounds for any identified subdistribution","Generalized optimal transport yields one-sided √n bounds","Sharp bounds with overlapping marginals at √n speed","One-sided √n consistency for sharp bounds in OT","Conservative √n estimation for arbitrary moment restrictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1384,"prompt_tokens":1010,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":301}},"tokens_in":626,"tokens_out":374,"duration_ms":4876,"temperature":1.0,"reasoning_tokens":301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:42:54.260515+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a feasible problem with identified moment functions c and a(t) that lie outside the Sobolev class of Assumption 3, compute the sieve gap δ_p = ||(I-p_J)c|| + max_t ||(I-p_J)a(t)|| as the polynomial degree J grows, and check whether it decays at the predicted polynomial rate; then simulate the one-sided coverage of the inequality (3). If the sieve gap still decays polynomially, or if the one-sided uniform validity holds despite the nonsmooth moments, the smoothness assumption is not load-bearing; if both fail, the theorem's rate mechanism is confirmed to depend on it.","supporting_citations":[],"review_version":1}