REVIEW 2 major objections 5 minor 18 references
A plug-in estimator of the upper endpoint of the identified set of a counterfactual choice probability is asymptotically normal under a nondegeneracy condition, enabling standard interval construction without point identification.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:40 UTC pith:PBQFT5FR
load-bearing objection Solid theory paper on set-identified choice probabilities with a sound normality theorem, but the nondegeneracy condition makes the central claim narrower than the abstract suggests. the 2 major comments →
Semiparametric inference on identification sets in choice modeling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central result is that the upper (and, symmetrically, lower) endpoint of the identified set of a counterfactual choice probability is a pathwise-differentiable functional of the observed distribution, despite being defined as the solution of a non-smooth optimization. Formally, for a sequence of designs with stable choice-set frequencies, if the true distribution is realizable by some mixing distribution and the endpoint linear program has a nondegenerate maximizer—a maximizer supported on exactly one more preference atom than there are nonredundant moment restrictions, with strictly positive weights and an invertible moment matrix—then the plug-in endpoint estimator satisfies a central
What carries the argument
The load-bearing mechanism is the pair of primal and dual linear programs that define each endpoint. The primal maximizes the target functional over mixing distributions subject to finitely many moment equalities; the dual minimizes an affine function of the observed moments subject to a pointwise constraint that the affine function majorizes the target. Nondegeneracy guarantees that the dual optimizer is unique and locally constant, so in a neighborhood of the true distribution the endpoint is exactly an affine function of the moment vector with that fixed slope. This local linearity is what converts a max-of-affine functional into a smooth parameter, yielding an influence-function expansio
Load-bearing premise
The claim stands or falls on the nondegeneracy condition: the true distribution must not sit on a tie where several equally good preference distributions define the same endpoint; if it does, the local linearity that drives the normal limit is absent.
What would settle it
Take a two-choice-set design and a target so that two distinct one-point preference distributions reproduce the observed moments and give the same upper bound. Simulate from the corresponding mixture at the tie, estimate the endpoint with the plug-in, and studentize with the paper's variance formula. If Theorem 3 is correct away from ties, the studentized statistic should be standard normal in nondegenerate cases and non-normal on the tie; if normality also holds on the tie, the paper's condition is not actually necessary.
If this is right
- Analysts can compute confidence intervals for counterfactual choice probabilities even when the latent preference distribution is not identified; only the nondegeneracy condition needs to hold.
- Because the endpoint is linear in finitely many moments, the inference procedure does not require approximating the preference space by grids, avoiding the bias grid approximations introduce.
- The plug-in variance estimator is consistent, so standard normal-based intervals for the endpoint are valid in large samples.
- The finite-support theorem implies the identified set can be traced by checking candidate values supported on at most the number of restrictions plus two preference atoms.
- The expectation-maximization membership algorithm gives a practical way to certify whether a given counterfactual probability lies in the identified set.
Where Pith is reading between the lines
- One could construct a valid confidence interval for the entire identified interval by combining the endpoint normality with a correction that accounts for the distance between endpoints; the paper focuses on endpoint coverage only.
- The nondegeneracy failure at ties suggests a promising extension: a uniformly valid confidence region that adapts to degeneracy by smoothing the endpoint or using a bootstrap calibrated on the local geometry; the paper does not offer such a procedure.
- The local affine representation suggests a design-of-experiments objective: choose which choice sets to observe and with what frequencies to minimize the resulting asymptotic variance; the paper mentions experimental design as motivation but does not solve the optimization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies inference on the identified set of a counterfactual choice probability in a discrete choice model where only finitely many choice-set probabilities are observed. The identified set is characterized through linear programs over mixing distributions, and the endpoints are shown to be pathwise differentiable under a nondegeneracy condition. The main theoretical result is asymptotic normality of plug-in endpoint estimators, along with a consistent variance estimator. The paper also provides a profile-NPMLE representation, a finite-support representation of the identified set, and an EM-like alternating projection algorithm with local convergence guarantees, specialized to mixed multinomial logit models.
Significance. If Theorem 3 is correct, this is the first general result giving Wald-type inference on identification sets of counterfactual choice probabilities without point-identification assumptions. The proof chain from realizability to the moment LP, to strong duality, to the local linear expansion, and finally to the CLT is coherent and mostly transparent. The finite-support representation and the membership-certification algorithm are also useful contributions. The main limitation is that the inference result relies on a nondegeneracy (no-tie) condition that is not verified for realistic DGPs and is not empirically checkable; this narrows the advertised scope but does not invalidate the internal logic.
major comments (2)
- [Definition 3 and Theorem 3] Nondegeneracy in Definition 3 is load-bearing: it implies strong duality, uniqueness of the dual optimizer, and the local linear expansion in Theorem 2. On a tie face, Theorem 2 fails and Theorem 3's normality does not follow, as the paper acknowledges. However, no primitive conditions or examples show that nondegeneracy holds for the mixed-MNL or transformer models that motivate the paper, and the condition is not verifiable from finite data. The scope of the inference claim should be explicitly qualified, or primitive sufficient conditions should be supplied.
- [Appendix D, Lemma 17 and Lemma 19] The proof of Lemma 17 is incomplete as written. Approximating the integral by simple functions shows only that the integral lies in the closure of Conv(H), not necessarily in Conv(H) itself. Lemma 19 then cites Lemma 17 to conclude Conv(H) membership, making the induction unjustified. This is load-bearing for the finite-support representation in Theorem 5 and for Proposition 2. The result may be repairable using the cited Schäfer-Ullrich theorem or a supporting-hyperplane argument, but the current text needs correction.
minor comments (5)
- [Section 4] The notation P_N = delta_(Y_1^N,...,Y_N^N) conflicts with the true distribution P_N. Recommend using P_hat_N or empirical measure notation throughout the inference section.
- [Appendix D, Lemma 17] Even if the lemma is true, the proof should distinguish between Conv(H) and its closure and provide a separate argument for the finite-convex-combination conclusion.
- [Lemma 14 proof] There is a typo near the beginning: 'P P_N' should be 'P_{P_N}'.
- [Section 6, Algorithm 1] The abstract says the algorithm certifies membership, but Algorithm 1 only certifies approximate membership up to tolerance epsilon and only from initializations in a set of positive probability. This qualification should appear in the abstract or early in Section 6.
- [Assumptions 3-4] Assumption 3 fixes b_N(P_N) and lambda(P_N) to be eventually constant. This is natural for a fixed-DGP asymptotic, but the paper should state explicitly that Theorem 3 is a sequence result under such an interpretation.
Circularity Check
No circularity: Theorem 3's normality follows from a proved local linearity of a dual-defined endpoint, not from a fitted input; self-citations are contextual only.
full rationale
The derivation chain is self-contained. Theorem 1 equates the identified-set endpoint to an LP value using Lemma 5's equivalence between moment matching and marginal matching under realizability — a proved representation, not a definitional identity. Proposition 1 proves strong duality and dual uniqueness from nondegeneracy (Definition 3) and Assumption 1; Theorem 2 derives the local affine expansion with constant slope λ(P_N) from those results. Theorem 3's asymptotic normality is a Lindeberg CLT on the linearized statistic (21), with variance σ_N^2 defined from the same dual coefficient; Lemma 14's consistency of σ_N(P̂_N)/σ_N(P_N) follows from Assumption 4 / Lemma 1, which are themselves proved from nondegeneracy, not assumed as the conclusion. The EM certification result (Theorem 6) is a conditional local-convergence theorem: Assumption 10 assumes the existence of the zero-KL point and Theorem 6 proves d^(h)→0 from it; this is a standard local convergence result, not a prediction that is forced by the assumption. Self-citations (Zielnicki et al. 2025, van der Laan et al. 2025, Aridor 2025, Kallus-Udell 2016) appear only in examples and related-work context and none is load-bearing for the main theorems. No fitted parameter is renamed as a prediction; the dual λ(P_N) is a function of the true moment vector and is estimated plug-in. Therefore no circular step is present.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Assumption 1: g ∉ Span{1, t_1^N, …, t_{L̃}^N}
- domain assumption Definition 3 (nondegeneracy): primal maximizer supported on L̃+1 atoms, positive weights, full-rank M(Q)
- domain assumption Assumption 6 (realizability with common Q_β for all N)
- domain assumption Assumption 7 (B compact metric, continuous kernel/g) and Assumption 9 (log-concavity)
- domain assumption Assumptions 10–15 (interior fixed point, local identifiability, PD Hessians, C^1 M-step)
- standard math Supporting hyperplane theorem, Carathéodory/Tchakaloff, Lindeberg CLT, implicit function theorem
read the original abstract
In a discrete choice model, choice probabilities observed for a finite collection of choice sets may not identify a counterfactual choice probability under an unobserved choice set. We represent this counterfactual probability as a linear functional of a mixing distribution. Because the target is a functional of a distribution whose support is not restricted to a finite set, the parameter space is infinite-dimensional, while the data impose only finitely many moment restrictions. Therefore, observed choice probabilities need not point identify such a target. The identified set is defined as the set of target values compatible with observed choice probabilities. Rather than imposing conditions to ensure point identification, we characterize the identified set, and conduct inference on its lower and upper endpoints. We represent each endpoint as the value of a linear program over probability measures, and give conditions to obtain pathwise differentiability of the identification bounds. As a consequence, we are able to prove asymptotic normality of plug-in endpoint estimators. Finally, we provide an Expectation-Maximization-like algorithm for certifying membership of candidate values in the identified set and establish local convergence guarantees.
Figures
Reference graph
Works this paper leans on
-
[1]
On the event ∥bN (PN )−b N (P N )∥< η, Assumption 4 gives λ(PN ) =λ(P N )
From Lemma 11 and (24), we have that σ2 N (P N )→ X a∈{a1,...,aM } 1 e(a) X l:a(l)=a blλ2 l − X l:a(l)=a blλl 2 >0 asN→ ∞,(25) From Lemma 13, we have that∥b N (PN )−b N (P N )∥ →0 in probability asN→ ∞, hence for anyη >0, we haveP P N (∥bN (PN )−b N (P N )∥< η)→1 asN→ ∞. On the event ∥bN (PN )−b N (P N )∥< η, Assumption 4 gives λ(PN ) =...
2000
-
[3]
Proof of Proposition 1.ConsiderSandTas defined in (16)
ThenPrimal(P N )andDual(P N )have the same value, andDual(P N ) admits a unique minimizer. Proof of Proposition 1.ConsiderSandTas defined in (16). LetP N ∈ Mnp,N and ΨD,N (P N ) denote the value ofDual(P N ). From weak duality, Ψ D,N (P N )≥ ˜Ψ+,N (P N ). Therefore, it remains to show that Ψ D,N (P N )≤ ˜Ψ+,N (P N ). Step 1: Supporting hyperplane.By Lemma...
2004
-
[11]
Lars van der Laan, Nathan Kallus, and Aur´ elien Bibaut. Inverse reinforcement learning using just classification and a few regressions.arXiv preprint arXiv:2509.21172,
-
[12]
For anyNsufficiently large, by realizability ofP N under (Qβ, s), we have that{XN i }i∈[N] is a sequence of independent random variables
whereE[X N i ] = 0 for anyN≥1, i∈[N] (Lemma 11). For anyNsufficiently large, by realizability ofP N under (Qβ, s), we have that{XN i }i∈[N] is a sequence of independent random variables. Therefore, by the Lindeberg central limit theorem [Billingsley, 1986, Theorem 27.2], we have that PN i=1 X N iqPN i=1 VarP N (X N i ) d− → N(0,1).(26) Using Lemma 11 one ...
1986
-
[16]
LetR={r∈Rsuch thatx 0+rv∈ H}.Hbeing bounded implies thatRis bounded
Let x0 ∈H, there existsv∈R L such that aff(H) =x 0+Rv. LetR={r∈Rsuch thatx 0+rv∈ H}.Hbeing bounded implies thatRis bounded. Therefore, there existsr:B →R measurable such that for anyβ∈ B, h(β) =x0 +r(β)v. Let ¯r= R B r(β)dπ(β), we have that Z B h(β)dπ(β) = Z B (x0 +r(β)v)dπ(β) =x 0 + Z B r(β)dπ(β)v=x 0 + ¯rv . Therefore, Lemma 18 gives that ¯r∈Conv(R). He...
2004
-
[17]
Sinceh(B 0)⊆H, we have Conv(h(B 0))⊆Conv(H), thus Z B h dπ∈Conv(H), hence the induction step and the result
gives that Z B0 h dπ∈Conv(h(B 0)). Sinceh(B 0)⊆H, we have Conv(h(B 0))⊆Conv(H), thus Z B h dπ∈Conv(H), hence the induction step and the result. A more general result than Lemma 19 is given in Sch¨ afer and Ullrich [2025, Lemma 2.16.]. We now define h:B →R ˜LN +1, β7→(t N (β)⊤, g(β))⊤, which is bounded and measurable with respect toB. For anyP N ∈ Mnp,N , ...
2025
-
[18]
Therefore, by the inverse function theorem [Lee, 2003, Theorem 4.5], there exists a neighborhood ˜Vk ofβ ⋆ k on whichFis injective
Consequently,∇F(β ⋆ k) has full column rankd. Therefore, by the inverse function theorem [Lee, 2003, Theorem 4.5], there exists a neighborhood ˜Vk ofβ ⋆ k on whichFis injective. Sinceβ ⋆ k ∈int(B), we can choose open neighborhoodsV k satisfying β⋆ k ∈V k, Vk ⊆ ˜Vk ∩int(B). Let Ω = n u∈R K−1 :u i >0, PK−1 i=1 ui <1 o . Sincew ⋆ ∈int(∆ K), there exists an o...
2003
-
[1951]
Susan Athey and Guido W. Imbens. Identification of average treatment effects in nonpara- metric panel models.arXiv preprint arXiv:2503.19873,
-
[1975]
A semiparametric discrete choice model: An application to hospital mergers.Economic Inquiry, 55(4):1919–1944,
Devesh Raval, Ted Rosenbaum, and Steven A Tenn. A semiparametric discrete choice model: An application to hospital mergers.Economic Inquiry, 55(4):1919–1944,
1919
-
[1982]
doi: 10.1007/978-1-4612-5769-1. Robin L Plackett. The analysis of permutations.Journal of the Royal Statistical Society Series C: Applied Statistics, 24(2):193–202,
-
[1983]
Alfred O Hero and Jeffrey A Fessler
doi: 10.1287/mksc.2.3.203. Alfred O Hero and Jeffrey A Fessler. Convergence in norm for alternating expectation- maximization (em) type algorithms.Statistica Sinica, pages 41–54,
-
[1985]
Eli Ben-Michael. Partial identification via conditional linear programs: estimation and policy learning.arXiv preprint arXiv:2506.12215,
-
[1987]
Martin Sch¨ afer and Tino Ullrich. Beyond tchakaloff quadrature: Positive functionals, frames and widths.arXiv preprint arXiv:2511.15425,
-
[1992]
Identification in some discrete choice models: A computational approach
Eric Mbakop. Identification in some discrete choice models: A computational approach. arXiv preprint arXiv:2305.15691,
-
[2010]
The value of personalized recommendations: Evidence from netflix.arXiv preprint arXiv:2511.07280,
Kevin Zielnicki, Guy Aridor, Aur´ elien Bibaut, Allen Tran, Winston Chou, and Nathan Kallus. The value of personalized recommendations: Evidence from netflix.arXiv preprint arXiv:2511.07280,
-
[2016]
Revealed preference at scale: Learning personalized preferences from assortment choices
Nathan Kallus and Madeleine Udell. Revealed preference at scale: Learning personalized preferences from assortment choices. InProceedings of the 2016 ACM Conference on Economics and Computation, pages 821–837,
2016
-
[2021]
Constantin Carath´ eodory.¨Uber den variabilit¨ atsbereich der fourier’schen konstanten von positiven harmonischen funktionen.Rendiconti Del Circolo Matematico di Palermo (1884-1940), 32(1):193–217,
1940
-
[2024]
On a theorem of karhunen and related moment problems and quadrature formulae
Georg Berschneider and Zolt´ an Sasv´ ari. On a theorem of karhunen and related moment problems and quadrature formulae. InSpectral Theory, Mathematical System Theory, Evolution Equations, Differential and Difference Equations: 21st International Workshop on Operator Theory and Applications, Berlin, July 2010, pages 173–187. Springer,
2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.