REVIEW 3 major objections 5 minor 13 references
Inference in partially identified moment models via regularized optimal transport
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read The paper shows that entropic optimal transport makes partially identified moment models estimable and testable, with a uniform central limit theorem powering confidence regions that control size locally uniformly.
desk verdict Useful and genuinely novel framework, but the main theorem is unproved and the inference target is the regularized set, not the sharp identified set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the entropic optimal transport value c_{θ,ε}(u)=min_π E_π[u'φ(X,Y,θ)] + ε·KL(π||μ⊗ν), a strictly convex regularization of the classical OT problem that restores √n-convergence and can be computed by the standard iterative scaling algorithm. The theorem that carries the inference is a uniform CLT for this regularized OT value as a process on C(B×Θ); it is the first CLT of such generality for entropic OT with smooth costs and is what turns the sample analogue of the max-min distance into a Gaussian process. The remaining machinery is a bootstrap for directionally differentiable functionals, applied with an enlarged argmax set Û_n={u∈B: ĉ(u) ≥ max_v ĉ(v) − ι_n} where
What would settle it
In the paper's Gaussian toy example (μ=N(0,1), ν=N(2,1), moment function 1{Y1>Y0}−θ), compute the entropic distance D_ε(θ) at θ=0.69, the sharp lower bound. With ε>0 fixed, the entropy penalty makes D_ε(0.69)>0, so 0.69 is outside the regularized identified set; if the bootstrap confidence region built from the regularized statistic excludes 0.69 with high probability, that is direct evidence that the procedure does not cover the sharp identified set.
Extended reading notes
Core claim
The paper's central claim is that in a GMM model where the data distribution is identified only up to its marginals, the sharp identified set equals {θ∈Θ : max_{u∈B} min_{π∈Π(μ,ν)} E_π[u'φ(X,Y,θ)] = 0}; the inner minimization is an optimal transport problem with linear cost u'φ, and the outer maximization finds the direction of the largest moment violation. To estimate this object, the paper replaces classical OT by the entropic regularized value c_{θ,ε}(u)=min_π E_π[u'φ]+ε KL(π||μ⊗ν), and proves a uniform CLT: under smooth moment functions and compact convex supports, √n(ĉ_θ(u)−c_θ(u)) converges weakly in C(B×Θ) to a tight Gaussian process G(u,θ). The delta method applied to the max functio
Load-bearing premise
The inference is formally about the regularized identified set Θ_{I,ε} that depends on the entropy penalty ε; for the confidence regions to cover the sharp identified set Θ_{I,0}, the regularization bias must vanish (via ε→0 at a suitable rate or explicit debiasing), which the paper's main theorems do not establish.
Editorial extensions
If this is right
- Researchers can build confidence regions for the identified set in partially identified GMM models with marginal data, and the regions control size locally uniformly rather than only at a fixed parameter.
- The procedure remains valid at parameter values on the boundary of the identified set, where the argmax of the cost function is not a singleton and the ordinary bootstrap is known to fail.
- The entropic regularization plus iterative scaling computation makes the method feasible in realistic sample sizes, as the Monte Carlo with n=15,000 shows.
- In the leading application, the method yields identified sets for the common slope parameter and the average marginal effects in a fixed-effects panel logit with attrition and refreshment; it tightens the bounds by fixing the observed retainer joint distribution and applying OT only to the attriters.
- The same framework covers nonlinear treatment effects, nonparametric instrumental variables without large-support conditions, and Euler equations estimated from repeated cross-sections.
Reading between the lines
- The main gap is the regularization bias: the formal inference targets the regularized set Θ_{I,ε}, and the paper does not prove ε-asymptotics showing that Θ_{I,ε} approaches the sharp set Θ_{I,0} at a rate compatible with the CLT; supplying that (or explicit debiasing) would turn the confidence regions into confidence regions for the sharp identified set.
- The uniform CLT should extend to Cramér–von Mises type statistics and to multi-marginal optimal transport, which the paper names as future work, so the framework could handle panels with more than two waves and specification tests of Θ_I≠∅.
- The paper's partitioned-attrition device—fixing the joint distribution for the subset of units with complete data and solving OT only for the incomplete units—is a general template for missing-data problems where partial joint information exists.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers partially identified GMM models in which the joint distribution of (X,Y) is unknown but the marginals are observed. It characterizes the sharp identified set as the set of parameters for which the distance from zero to the set of moment predictions is zero, and represents this distance through a max-min optimal transport problem. To make estimation and inference feasible, the authors replace the classical OT problem by an entropically regularized version, yielding a smooth criterion that can be computed with Sinkhorn iterations. They propose a plug-in estimator of the identified set, a test statistic based on the regularized distance, a bootstrap procedure using directionally differentiable functionals, and a uniform CLT for the regularized OT value. The paper also discusses several empirical applications and reports a Monte Carlo simulation for a fixed-effects panel logit with attrition and refreshment.
Significance. If the main results were correct, the paper would make a useful contribution: a general smooth-cost CLT for entropic OT values would extend the existing literature, and the proposed inference procedure would provide a practical route to testing and confidence regions in a broad class of partially identified models with marginal data. The paper also correctly identifies an interesting set of applications, and the computational strategy is sensible. However, the central statistical target is not the sharp identified set but a regularized subset of it, and no asymptotics are provided to remove the regularization bias. In addition, the proof of the key uniform CLT in Appendix A.3 is not a coherent proof. These are load-bearing problems, not presentation issues.
major comments (3)
- [§3.1, Theorem 2, Abstract] The inference target is Θ_{I,ε}, not the sharp identified set Θ_{I,0}. After defining Θ_{I,ε} = {θ : D_ε(θ)=0}, the paper states 'we drop the subscript ε for brevity', and all subsequent results — Theorem 2, Corollary 1, Algorithm 1 — are stated for Θ_I and D, i.e., for the regularized criterion with fixed ε. Since entropic regularization adds a nonnegative KL penalty, D_ε(θ) ≥ D_0(θ) for every θ, and Θ_{I,ε} ⊆ Θ_{I,0} with strict inclusion in general (the toy example at θ=0.69 is a case in point). No ε→0 asymptotics or debiasing is supplied in the main theory. Remark 4 offers only a conservative sample-level adjustment and is not incorporated into Algorithm 1. Therefore the abstract's claim of confidence regions for the sharp identified set, and Theorem 2's Hausdorff consistency for Θ_I, are not established for the procedure actually implemented. The paper could be reframed as targeting
- [Appendix A.3, Theorem 3] The proof of Theorem 3, the paper's principal theoretical contribution, is not a proof. After defining the map c and a class F, the section jumps into 'Proof. Part (i). Since s > d/2, ...' and proceeds with statements about Donsker classes, semiparametric efficiency, and a Proposition 4, without ever establishing finite-dimensional convergence, tightness on C(B×Θ), or the Gaussian process G. The text then states several propositions 'follow[ing]' Mena and Niles-Weed (2019) and Goldfeld et al. (2024), but the connection to the uniform CLT claimed in Theorem 3 is not shown. Given that Theorem 3 is the basis for Corollary 1 and the bootstrap, this missing argument is load-bearing. The reader cannot verify the central claim from the manuscript.
- [§3.2, Corollary 1 and Assumption 4] The bootstrap validity result is asserted by reference to 'Theorem 4 of Franguridi and Moon (2025)', a companion paper, rather than proved or even stated. More importantly, Corollary 1 and the bootstrap procedure rely on the null hypothesis being D(θ0)=0 and on the argmax set U_c(θ0) being well-defined. Under the regularized criterion actually used, the relevant null is D_ε(θ0)=0, i.e., θ0 ∈ Θ_{I,ε}, not θ0 ∈ Θ_{I,0}. Thus even if Theorem 3 were fully proved, the procedure would control size for a hypothesis about the regularized set, not the sharp set claimed in the paper. The reliance on an external result for the bootstrap may be acceptable in principle, but it leaves a central step unverified in this manuscript.
minor comments (5)
- [Assumption 4(ii)] The condition 'ι_n ↓ 0 and n^{-1/2} ι_n ↑ ∞' is impossible as written since the second expression tends to 0 if ι_n → 0. The intended condition is presumably n^{1/2} ι_n → ∞. The simulation uses ι_n = 0.05 n^{-1/2} log n, which satisfies n^{1/2} ι_n → ∞ but not n^{-1/2} ι_n ↑ ∞. Please fix the statement and align it with the implementation.
- [Notation, §3.1] The practice of dropping the subscript ε after defining D_ε and Θ_{I,ε} makes Assumption 2(i) ambiguous: it states D(θ) ≥ m(d(θ,Θ_I)), but if D is the regularized distance, the correct set in the separation condition is Θ_{I,ε}. This ambiguity should be removed, either by keeping ε explicit or by clearly redefining the target set.
- [Remark 4] Remark 4 proposes an adjusted statistic based on the bound ĉ_{θ0,ε}(u) − ĉ_{θ0,0}(u) ≤ ε(log n − KL(...)), but no formal theorem states the coverage properties of this adjusted test. The phrase 'potentially conservative' is not enough. Either prove a claim or omit the suggestion from the main text.
- [Appendix B.1, Algorithm 4] Line 4 uses a kernel estimator for f_{2|ret} but no assumptions or rate conditions on the bandwidth h are given. The Monte Carlo uses discrete covariates where this is not an issue, but the general algorithm as described would require a theory for the first-step nonparametric estimator that is not supplied.
- [Figure 2] The right-panel caption says 'coverage probability' but the color scale is not defined; the text says the confidence region 'performs reasonably well' without reporting numerical coverage rates. Since the main inference claim is at stake, the simulation section should report actual rejection rates for points inside and outside the identified set.
Circularity Check
No definitional circularity in the uniform CLT; the load-bearing bootstrap-validity claim is delegated to a companion paper by the first author.
-
self citation load bearing
[Section 3.2, after Assumption 4 and Algorithm 1]
"It is straightforward to establish that under Assumption 4, our testing procedure controls size locally uniformly in the sense of Corollary 3.2 of Fang and Santos (2019). We refer the reader to Theorem 4 of Franguridi and Moon (2025) for details."
The local-uniform size control of the bootstrap test is the paper's main inferential guarantee and the basis for the confidence regions advertised in the abstract. The manuscript does not prove this result; it states that it follows from Assumption 4 and then delegates the verification to Theorem 4 of Franguridi and Moon (2025), a companion paper by the first author. A load-bearing step of the inference chain is therefore carried by a self-citation rather than by a derivation contained in this paper or by an independently machine-checked theorem. This is not a definitional equality, and the CLT itself is external, so the circularity is partial rather than total.
full rationale
The core theoretical result, Theorem 3, is not circular: its proof is built on Mena and Niles-Weed (2019), Goldfeld et al. (2024), and the empirical-process delta method, all external to this paper, and the uniform CLT is used in Corollary 1 via Shapiro's directional delta method. The main self-citation concern is the bootstrap-validity step, where the paper delegates the locally-uniform size control to Theorem 4 of Franguridi and Moon (2025), a companion paper by the first author. Theorem 1's proof also says it follows Proposition 1 of the same companion, but a full proof is reproduced in Appendix A.1, so that citation is less load-bearing. I also note a non-circular but important limitation: Section 3.1 defines the target as Theta_{I,epsilon} and says 'we drop the subscript epsilon for brevity,' so Theorem 2 and Algorithm 1 concern the regularized identified set, not the sharp set Theta_{I,0}; no epsilon-to-zero or debiasing asymptotics are supplied, and Remark 4 only gives a conservative adjustment. This is an identification/scope gap, not a circular reduction. Overall score 4: some self-citation at a load-bearing joint, but the central CLT has independent content.
Assumptions & free parameters
free parameters (4)
- ε (entropic regularization parameter) =
0.1 in Monte Carlo; otherwise user-specified
- η_n (threshold for set estimator)
- ι_n (bootstrap argmax enlargement) =
0.05 n^{-1/2} log n in Monte Carlo
- bandwidth h for kernel estimator of f2|ret
assumptions (7)
- domain assumption Assumption 1(i)-(v): compact Θ, compact supports, nonempty Θ_{I,0}, continuity of φ
- domain assumption Assumption 2(i): existence of separation function m with D(θ)≥m(d(θ,Θ_I))
- domain assumption Assumption 3(i)-(iii): bounded convex supports, φ_j ∈ C^s with s>d/2, independent samples
- domain assumption Assumption 4(i): well-separated argmax, cθ0(u) ≤ max - κ d_H(u,U_c)
- standard math Sion's minimax theorem and norm duality
- standard math Berge's maximum theorem and compactness results from Aliprantis-Border
- domain assumption Theorem 4 of Franguridi and Moon (2025)
Cite this review
Pith. "Pith review of Inference in partially identified moment models via regularized optimal transport." pith.science (2026). https://pith.science/paper/PIWMMYQA
@misc{pith2026251218084,
author = {Pith},
title = {Pith review of: Inference in partially identified moment models via regularized optimal transport},
year = {2026},
howpublished = {\url{https://pith.science/paper/PIWMMYQA}},
note = {Machine review of arXiv:2512.18084}
}
read the original abstract
Many statistical and econometric problems involve parameters defined by moments of a joint distribution when only marginal distributions are observed, leading naturally to partial identification. We develop a methodology for identification, estimation, and inference in the corresponding partially identified GMM model. We characterize the sharp identified set for the parameter of interest via a support-function/optimal-transport (OT) representation. To estimate the identified set, we employ entropic regularization, which yields a smooth approximation to the classical OT problem that can be computed efficiently using the Sinkhorn algorithm. We also propose a test statistic for hypothesis testing and the construction of confidence regions for the identified set. To derive its asymptotic distribution, we establish a novel central limit theorem for the entropic OT value under general smooth cost functions. We then obtain valid critical values using the bootstrap for directionally differentiable functionals of Fang and Santos (2019). The resulting testing procedure controls size locally uniformly, including at parameter values on the boundary of the identified set. We demonstrate good finite-sample performance of our methodology in Monte Carlo simulations. Finally, as an empirical illustration, we estimate a panel logit model of self-reported happiness with attrition and refreshment, using data from the Understanding America Study.
Figures
Reference graph
Works this paper leans on
-
[1]
Inconsistency of the bootstrap when a parameter is on the boundary of the parameter space,
Aliprantis, C. D. and K. C. Border(2006):Infinite dimensional analysis: a hitchhiker’s guide, Springer. Andrews, D. W.(2000): “Inconsistency of the bootstrap when a parameter is on the boundary of the parameter space,”Econometrica, 399–405. Andrews, D. W. and P. Guggenberger(2009): “Validity of subsampling and “plug-in asymp- 22 totic” inference for param...
2006
-
[2]
30 Finally,V µ(φ) +V ν(ψ) is the semiparametric variance bound due to Corollary 2 of Goldfeld et al
=θ(ρ 1, ρ2) andF=F ⊕ yields √n(θ(ˆµn,ˆνn)−θ(µ, ν)) = √n(δ(ˆµn ⊗ˆνn)−δ(µ⊗ν)) ⇝δ ′ µ⊗ν(Gµ⊗ν) =G µ⊗ν(φ⊕ψ)∼N(0,V µ(φ) +V ν(ψ)). 30 Finally,V µ(φ) +V ν(ψ) is the semiparametric variance bound due to Corollary 2 of Goldfeld et al. (2024). The following result and its proof follow Proposition A.1 in Mena and Niles-Weed (2019). Proposition 1.For anyu∈Bandθ∈Θ, the...
2024
-
[6]
, α2d)∈N 2d 0 or order|α|:=α 1 + · · ·+α2d ≤s, and the partial derivative operator ∇αf(x, y) := ∂|α| ∂xα1 1 · · ·∂xαd d ∂y αd+1 1 · · ·∂yα2d d f(x, y)
A.3 Proof of Theorem 3 For a functionf∈C s(X × Y), denote its H¨ older norm by ∥f∥ Cs(X ×Y)= max 0≤|α|≤s sup (x,y)∈X ×Y |∇αf(x, y)|, where the maximum is taken over all multi-indicesα= (α 1, . . . , α2d)∈N 2d 0 or order|α|:=α 1 + · · ·+α2d ≤s, and the partial derivative operator ∇αf(x, y) := ∂|α| ∂xα1 1 · · ·∂xαd d ∂y αd+1 1 · · ·∂yα2d d f(x, y). Assumpti...
2023
-
[7]
Similarly, (16) implies θ(µ1, ν1)−θ(µ 0, ν0)≤ Z (φ11 ⊕ψ 01)d(µ 1 ⊗ν 1 −µ 0 ⊗ν 0)
= Z (φ01 ⊕ψ 00)d(µ 1 ⊗ν 1 −µ 0 ⊗ν 0). Similarly, (16) implies θ(µ1, ν1)−θ(µ 0, ν0)≤ Z (φ11 ⊕ψ 01)d(µ 1 ⊗ν 1 −µ 0 ⊗ν 0). Sinceφ ij ⊕ψ kℓ ∈F ⊕, we obtain (7). Moreover, arguing as in the proof of Proposition 3, we obtain lim t↓0 θ(µ0 +t(µ 1 −µ 0), µ0 +t(ν 1 −ν 0))−θ(µ 0, ν0) t = Z (φ00 ⊕ψ 00)d(µ 1 ⊗ν 1 −µ 0 ⊗ν 0). DefineP 0 as the set of probability measure...
2024
-
[9]
Therefore,φ u,θ, ψu,θ are optimal potentials
Since (φ0 u,θ, ψ0 u,θ) maximizes the dual objective, so does (φu,θ, ψu,θ). Therefore,φ u,θ, ψu,θ are optimal potentials. The following result and its proof follow Proposition 1 in Mena and Niles-Weed (2019). Proposition 2.For anyu∈Bandθ∈Θ, there exist optimal dual potentialsφ u,θ, ψu,θ such that for any multi-indexα∈N d 0 of order|α| ≤s, |∇αφu,θ(x)| ≤Cs,d...
2019
-
[10]
Denotek=|α|. By the multivariate Fa` a di Bruno’s formula (see, e.g., Corollary 2.10 in Constantine and Savits (1996)), ∇αφu,θ(x) = X β1+···+βk=α λα,β1,...,βk kY j=1 Z e−ψ0 u,θ(y)∇βj e−u′ϕ(x,y,θ) dν(y),(12) where the summation is over multi-indicesβ 1, . . . , βk ∈N d 0 such thatβ 1 +· · ·+βk =αandλ α,β1,...,βk are combinatorial quantities that only depen...
1996
-
[11]
SinceFis independent of (u, θ), the proof is completed
≤ ∥µ1 −µ 0∥F , where the last inequality is due toφ 0 u,θ, φ1 u,θ ∈F. SinceFis independent of (u, θ), the proof is completed. Proposition 4.There exists a tight Gaussian processG µ⊗ν inℓ ∞(F ⊕)such that √n(ˆµn ⊗ˆνn −µ⊗ν)⇝G µ⊗ν inℓ ∞(F ⊕). Proof.See the proof of part (ii) of Theorem 1 in Goldfeld et al. (2024). A.4 Proof of Corollary 1 Theorem 3.1 in Shapi...
2024
-
[12]
Generalized Optimal Transport,
van der V aart, A. and J. Wellner(2023):Weak Convergence and Empirical Processes: With Applications to Statistics, Springer. V an der V aart, A. W.(2000):Asymptotic statistics, vol. 3, Cambridge university press. Voronin, A.(2025): “Generalized Optimal Transport,”arXiv preprint arXiv:2507.22422. 26 Appendices A Proofs of theoretical results A.1 Proof of T...
arXiv 2023
Show all 13 references
-
[13]
We partition the population into retainers (observed in both periods) and attriters (observed only in period 1)
We embed the event{Y i1 +Y i2 = 1}in the cost function ϕ(y 1, y2, x1, x2;θ) =s(y 1, y2, x1, x2;θ)1{y 1 +y 2 = 1}. We partition the population into retainers (observed in both periods) and attriters (observed only in period 1). This partitioned approach yields tighter bounds by...
2024
-
[14]
Using these explicit forms, we can simplify the expressions forp(x, s, θ) anda(x, s, θ)
Then, expanding the defining identity gives λ1(x, θ)u+λ2(x, θ)u2 +λ 3(x, θ)u3 =θ ju(1−u) 1 +u exp (x2 −x 1)′ θ −1 , which implies λ1(x, θ) =θj, λ 2(x, θ) =θj exp (x2 −x 1)′ θ −2, λ 3(x, θ) =θj 1−exp (x2 −x 1)′ θ . Using these explicit forms, we can simplify the expressions for...
2024
-
[26]
Identification and estimation of average marginal effects in fixed effects logit models,
Davezies, L., X. D’Haultfoeuille, and L. Laage(2024): “Identification and estimation of average marginal effects in fixed effects logit models,”arXiv preprint arXiv:2105.00879. Dynan, K. E., J. Skinner, and S. P. Zeldes(2004): “Do the Rich Save More?”Journal of Political Econo...
2024 arXiv
-
[32]
Debiaser beware: Pitfalls of 25 centering regularized transport maps,
Pooladian, A.-A., M. Cuturi, and J. Niles-Weed(2022): “Debiaser beware: Pitfalls of 25 centering regularized transport maps,” inInternational conference on machine learning, PMLR, 17830–17847. Romano, J. P. and A. M. Shaikh(2010): “Inference for the identified set in partially...
2022
-
[709]
Invalidity of the bootstrap and the m out of n bootstrap for confidence interval endpoints defined by moment inequalities,
Andrews, D. W. and S. Han(2009): “Invalidity of the bootstrap and the m out of n bootstrap for confidence interval endpoints defined by moment inequalities,”The Econometrics Journal, 12, S172–S199. Andrews, D. W. and X. Shi(2013): “Inference based on conditional moment inequal...
2009 arXiv
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.