REVIEW 4 major objections 4 minor 1 cited by
This paper shows that the bias that imperfect product proxies inject into demand counterfactuals can be removed with a post-estimation correction whose standard errors ignore how the proxies were built.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 11:40 UTC pith:HJCRCXZR
load-bearing objection A genuinely useful debiasing toolkit for demand counterfactuals with ML proxies; the theory is coherent, but the key local-misspecification condition needs primitive support before the application claims are fully load-bearing. the 4 major comments →
From Unstructured Data to Demand Counterfactuals: Theory and Practice
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that mismeasurement of product attributes by proxies is a model misspecification, not a classical measurement-error problem, and can be neutralized by reparameterizing the demand model in terms of a composite parameter γ that bundles structural parameters and latent attributes. Once the naive estimate γ̂ = γ(θ̂, ẽ) is in hand, the counterfactual estimator is adjusted by subtracting weighted averages of the estimation moments, with weights chosen to make the adjusted estimator's first-order dependence on γ̂ vanish. Under the condition that γ̂ is within a neighbourhood of γ0 whose radius is negligible relative to sampling error, the adjusted estimator is asymptoti
What carries the argument
The key machinery is the composite parameter γ(θ,e), the low-dimensional combination of structural parameters and latent product attributes through which all choice probabilities and counterfactuals enter the model. By expressing both the naive estimator and the estimation moments as functions of γ̂, the paper chooses correction weights so that the adjusted counterfactual has zero first-order sensitivity to γ̂; this makes the choice of proxy irrelevant to the estimator's leading bias term. The same object drives the diagnostics: LM1 compares γ̂ to the set of composite parameters spanned by the proxies, and LM2 augments the proxy space with an extra direction to test whether the proxy dimensi
Load-bearing premise
The load-bearing assumption is that the composite parameter derived from the proxies lands close enough to the true value — closer than the fourth root of the sample size; if the proxies are too noisy, nothing in the data forces this, and the correction loses its centering property.
What would settle it
Run a controlled simulation with true latent attributes known and proxy noise at roughly ρ=0.5 or higher, increasing sample size; if the bias-corrected estimator's bias does not shrink at the claimed rate or does not remain far below the naive estimator's bias, the local condition is not doing the work in that regime.
If this is right
- Bias-corrected counterfactuals are centered at the true value and come with closed-form standard errors, so no bootstrap or re-estimation is needed after the correction.
- Standard errors remain valid when embeddings are fine-tuned on the choice data, because the asymptotic distribution does not depend on the proxy to first order.
- The LM1 and LM2 diagnostics give a practical answer to which unstructured-data proxies to use and how many principal components to keep.
- In the e-book experiment, the correction raises closest-substitute hit rates from 40% to 60–70% and improves or ties 11 of 13 specifications ruled in by the dimension diagnostic.
- Even when mismeasurement is not a concern, the estimator offers efficient counterfactual inference with simple standard errors.
Where Pith is reading between the lines
- A natural testable extension is to calibrate the LM1 threshold against a small validation set with known attribute values, rather than the heuristic χ² log T cutoff, to see whether better proxy-selection decisions result.
- The paper's logic implies that the cost of searching over many embedding choices is lower than usually assumed: model selection should emphasize the LM diagnostics, since the target counterfactual variance is proxy-independent to first order.
- The method is best suited to attributes that are fixed product characteristics; applying it to time-varying or context-dependent attributes would require the composite-parameter mapping to hold within each market or individual setting, which the paper does not claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops post-estimation bias corrections and diagnostics for demand counterfactuals when product attributes are proxied by embeddings or other imperfect measures. The key device is a reparameterization in which attributes and structural parameters enter choice probabilities only through a composite parameter γ; the naive counterfactual estimator is corrected by adding a linear combination of the model's own moment conditions, with weights chosen to remove first-order dependence on γ̂ and to minimize asymptotic variance. The main theoretical results (Propositions 1, 3, 5) state that, provided γ̂ = γ₀ + o_p(T^{-1/4}), the corrected estimator is asymptotically centered at the true counterfactual with a variance independent of γ̂, θ̂, and the proxy. Two LM-type diagnostics are proposed to assess whether γ̂ is sufficiently close to γ₀ and whether the proxy dimension is adequate. Simulations and an application to ebook choice data illustrate the method, with closest-substitute hit rates improving from 40% to 70% in the preferred specification.
Significance. If the maintained rate condition holds, the paper gives a practically useful, computationally light way to debias counterfactual inference in demand models with imperfect proxies. The closed-form standard errors, efficiency property within a natural class, and accommodation of data-dependent proxies are genuine strengths. The empirical demonstration against ground-truth second choices is valuable. However, the central theoretical guarantee is conditional on the proxy-induced estimator γ̂ converging to γ₀ at o_p(T^{-1/4}), and the paper does not provide primitive conditions ensuring this for the leading case of fixed embeddings. The diagnostics are useful heuristics but, as formalized, cannot certify the required rate. These issues are load-bearing for the paper's headline claim of valid inference for unstructured-data proxies, though they are fixable by a more careful statement of scope and by adding rate conditions or a local-misspecification framework.
major comments (4)
- [§6.1, Propositions 1 and 3; §6.2, Proposition 5; Remark 8] The condition γ̂ = γ₀ + o_p(T^{-1/4}) is maintained without a primitive justification. For a fixed, off-the-shelf embedding ẽ, the natural limit is γ̂ → γ(θ*(ẽ), ẽ), where θ*(ẽ) is the pseudo-true value; nothing guarantees γ(θ*(ẽ), ẽ) = γ₀. Remark 8 appeals to fine-tuning on the same data, but no rate for the convergence of a data-dependent ẽ to e₀ is established, so the required o_p(T^{-1/4}) shrinkage is not derived. Since the paper's leading case is exactly such proxies, the central asymptotic claim is not supported for the main application without additional assumptions. I recommend either providing concrete primitive conditions on proxy construction (e.g., embeddings refined at a controlled rate) or explicitly recasting the theory as local-misspecification asymptotics and stating that the debiasing guarantee applies only when the proxy error is smaller than the sampling error at the
- [§2.3.1 and §3.3.1, Proposition 4 and Proposition 6] The LM1 diagnostic is presented as validating the condition that γ̂ is sufficiently close to γ₀. But Proposition 4 is proved under γ̂ = γ₀ + o_p(1). Under fixed proxy mismeasurement γ̂ is inconsistent, so LM1 diverges at rate T and the proposed threshold C_T² = χ²_{dim γ,0.95} log T will reject with probability approaching one. A finite-sample non-rejection then only indicates that the sample is too small to detect the misspecification; it cannot certify the required o_p(T^{-1/4}) rate. The wording in the practitioner's guides (e.g., "conclude that ẽ is sufficiently close to e₀") overstates what the diagnostic can establish. The diagnostic is still useful as a specification check, but the paper should state this limitation and provide either a formal local-power analysis or a bound that is valid under fixed misspecification.
- [§6.1, proof of Proposition 1; §4 simulations] Even when the bias correction removes the first-order term, the remaining bias is of order O(∥γ̂−γ₀∥²) under fixed proxy error. If γ̂ fails to converge to γ₀, this second-order bias can be non-negligible; the simulations in Figure 2 indeed show that for ρ > 0.5 the corrected estimator's bias increases. The paper acknowledges this behavior informally, but the theoretical sections do not make explicit the consequence that the distributional results are only approximate for fixed mismeasurement and that the approximation degrades as ∥γ̂−γ₀∥ grows. Please state this as a formal caveat near Propositions 1 and 5, and quantify the remainder as a function of ∥γ̂−γ₀∥.
- [§6.1.2, Proposition 4; §2.3.1, threshold choice] The choice C_T² = χ²_{dim(γ),0.95} log T is heuristic. Proposition 4 only gives statements 'wpa1' conditional on a chi-square random variable being below ϵ²C_T²; it does not provide the distributional approximation needed to calibrate the threshold for controlling a false-acceptance probability. The paper should either derive the large-deviation or local-alternative properties of LM1 that justify this threshold, or present it as an ad hoc rule. This matters because the threshold is used in the application to select among specifications.
minor comments (4)
- [Proposition 3 (p. 35)] The statement says 'Let Assumption 3 hold', but the result is for the no-microdata case and should refer to Assumption 2.
- [Proposition 7 (p. 39)] There is a typo: 'Let Assumptions 3 and 4 bold hold' should be 'both hold'. Also the display uses 'LM' without a subscript; it should be 'LM1'.
- [§2.2, Remark 2 and eq. (15)] The variance estimator (15) is stated without derivation. It would help readers to see that it is the sample analogue of the variance expression in Proposition 1, including the cross-term between k_t and the moments.
- [§4, Figure 2] The text says 'when mismeasurement becomes very large (ρ>0.5), the bias correction starts to also perform worse,' but the figure appears to show the onset around ρ=0.4–0.5. Please align the text with the simulation grid.
Circularity Check
No significant circularity: the bias-correction construction is self-contained and the application is validated against an external ground truth.
full rationale
The target κ0 is not an input to the construction of κ̂bc. The weights ĉ and d̂t in eqs. (12)-(14) are chosen from sample moments and derivatives of kt, ξt, and mt at γ̂, with probability limits solving the orthogonality constraint (33); Proposition 1 then shows the first-order term in (γ̂−γ0) vanishes by the model's own moment condition E[Ztξt(γ0)]=0. This is a standard influence-function debiasing argument, not a tautology: the centering at κ0 follows from those moment conditions, not from plugging κ0 into the estimator. The application uses second-choice data only to score the closest-substitute prediction, so the 40%-to-70% improvement is an external falsification exercise, not a fitted prediction. The only self-citation, Compiani et al. (2025), supplies the data and naive estimates for comparison; it is not used to derive Propositions 1 or 5 and is validated against an external ground truth, so it is not load-bearing. The rate condition γ̂=γ0+o_p(T^{-1/4}) (Remark 8; Propositions 1,3,5) is a substantive regularity/identification assumption limiting applicability when fixed embeddings are mismeasured, but an unsatisfied assumption is a correctness/robustness concern, not a by-construction circularity; LM1 is a diagnostic for this condition, not a proof of it. No step in the derivation chain reduces to its own input.
Axiom & Free-Parameter Ledger
free parameters (3)
- LM1 threshold C_T² = χ²_{dim(γ),0.95} log T =
χ²_{0.95} log T (target rate √(log T)/T)
- Number of principal components (dimension of ẽ) in the application
- Mismeasurement design parameter ρ in simulations =
0, 0.1, 0.2, 0.3, 0.4, >0.5
axioms (8)
- domain assumption IV exogeneity E[ξ_jt | z_jt] = 0 (Eq. 2)
- domain assumption Composite-parameter structure: (θ, e) enter choice probabilities, moments, and counterfactuals only through γ(θ, e)
- domain assumption Latent attributes e do not vary across markets
- domain assumption Same dimension and rank: e₀ and ẽ are J×r with full row rank; Γ convex and open
- ad hoc to paper Local misspecification: γ̂ = γ₀ + o_p(T^{-1/4})
- domain assumption Proxies admit a probability limit e* (fine-tuning case)
- standard math Smoothness, moment, and rank conditions (Assumptions 1(i)-(iii), 2, 3)
- domain assumption Assumption 4(iii): ∥θ̂−θ*∥ ≤ C∥ẽ−e*∥ wpa1 and first-order condition for θ̂
read the original abstract
Empirical models of multi-product demand rely on low-dimensional product representations to capture substitution patterns, increasingly using proxies built from unstructured data. When proxies are imperfect, standard workflows yield biased counterfactuals and invalid inference. We develop a practical toolkit to address these issues. Our methods apply to market-level and/or individual data, require minimal additional computation, provide simple standard-error formulas, and accommodate proxies from fine-tuned models. Further, we propose diagnostics to assess proxy quality. Our methods yield meaningful improvements in predicting substitution in empirically calibrated simulations and in an application where we assess counterfactual prediction performance against a ground truth.
Figures
Forward citations
Cited by 1 Pith paper
-
Econometrics with Pre-Trained Embeddings for Unstructured Data
Pre-trained embeddings are valid in double machine learning when the target nuisance function lies in the span of the source-task representation; under that condition the downstream estimator can converge faster than ...
Reference graph
Works this paper leans on
-
[1]
Ai, C. and X. Chen (2012): The semiparametric efficiency bound for models of sequential moment restrictions containing unknown functions, Journal of Econometrics, 170, 442--457
2012
-
[2]
Allcott, H. and N. Wozny (2014): Gasoline prices, fuel economy, and the energy paradox, Review of Economics and Statistics, 96, 779--795
2014
-
[3]
Allon, G., D. Chen, Z. Jiang, and D. Zhang (2023): Machine learning and prediction errors in causal inference, The Wharton School Research Paper
2023
-
[4]
Andrews, D. W. (1994 a ): Asymptotics for semiparametric econometric models via stochastic equicontinuity, Econometrica: Journal of the Econometric Society, 62, 43--72
1994
-
[5]
--- -.1pt --- -.1pt --- (1994 b ): Empirical process methods in econometrics, Handbook of econometrics, 4, 2247--2294
1994
-
[6]
--- -.1pt --- -.1pt --- (2005): Cross-section regression with common shocks, Econometrica, 73, 1551--1585
2005
-
[7]
Angelopoulos, A. N., S. Bates, C. Fannjiang, M. I. Jordan, and T. Zrnic (2023): Prediction-powered inference, Science, 382, 669--674
2023
-
[8]
Bach, P., V. Chernozhukov, S. Klaassen, M. Spindler, J. Teichert-Kluge, and S. Vijaykumar (2024): Adventures in demand analysis using AI, arXiv preprint arXiv:2501.00382
arXiv 2024
-
[9]
Conlon, and M
Backus, M., C. Conlon, and M. Sinkinson (2021): Common Ownership and Competition in the Ready-To-Eat Cereal Industry, NBER Working Paper 28350
2021
-
[10]
Battaglia, L., T. Christensen, S. Hansen, and S. Sacher (2024): Inference for Regression with Variables Generated by AI or Machine Learning, arXiv preprint arXiv:2402.15585
Pith/arXiv arXiv 2024
-
[11]
Ferreira, and R
Bayer, P., F. Ferreira, and R. McMillan (2007): A unified framework for measuring preferences for schools and neighborhoods, Journal of Political Economy, 115, 588--638
2007
-
[12]
Levinsohn, and A
Berry, S., J. Levinsohn, and A. Pakes (1995): Automobile Prices in Market Equilibrium, Econometrica, 63, 841--890
1995
-
[13]
--- -.1pt --- -.1pt --- (2004): Differentiated Products Demand System from a Combination of Micro and Macro Data: The New Car Market, Journal of Political Economy, 112, 68--105
2004
-
[14]
Berry, S. T. and P. A. Haile (2014): Identification in differentiated products markets using market level data, Econometrica, 82, 1749--1797
2014
-
[15]
4, 1--62
--- -.1pt --- -.1pt --- (2021): Foundations of demand estimation, in Handbook of industrial organization, Elsevier, vol. 4, 1--62
2021
-
[16]
--- -.1pt --- -.1pt --- (2024): Nonparametric identification of differentiated products demand using micro data, Econometrica, 92, 1135--1162
2024
-
[17]
Brown, B. W. and W. K. Newey (1998): Efficient semiparametric estimation of expectations, Econometrica, 66, 453--464
1998
-
[18]
Carlson, J. and M. Dell (2025): A Unifying Framework for Robust and Efficient Inference with Unstructured Data, arXiv preprint arXiv:2505.00282
arXiv 2025
-
[19]
Hong, and E
Chen, X., H. Hong, and E. Tamer (2005): Measurement error models with auxiliary data, The Review of Economic Studies, 72, 343--366
2005
-
[20]
Hong, and A
Chen, X., H. Hong, and A. Tarozzi (2008): Semiparametric efficiency in GMM models with auxiliary data, The Annals of Statistics, 36, 808--843
2008
-
[21]
Compiani, G., I. Morozov, and S. Seiler (2025): Demand estimation with text and image data, arXiv preprint arXiv:2503.20711
arXiv 2025
-
[22]
Conlon, C. and J. Gortmaker (2025): Incorporating Micro Data into Differentiated Products Demand Estimation with PyBLP, Journal of Econometrics, 105926
2025
-
[23]
Dub \'e , J.-P. and P. E. Rossi (2019): Handbook of the Economics of Marketing, vol. 1, North Holland
2019
-
[24]
Hinck, B
Egami, N., M. Hinck, B. Stewart, and H. Wei (2023): Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large language models, Advances in Neural Information Processing Systems, 36, 68589--68601
2023
-
[25]
(2013): Ownership consolidation and product characteristics: A study of the US daily newspaper market, American Economic Review, 103, 1598--1628
Fan, Y. (2013): Ownership consolidation and product characteristics: A study of the US daily newspaper market, American Economic Review, 103, 1598--1628
2013
-
[26]
Fong, C. and M. Tyler (2021): Machine learning predictions as regression covariates, Political Analysis, 29, 467--484
2021
-
[27]
(2015): Asymptotic theory for differentiated products demand models with many markets, Journal of Econometrics, 185, 162--181
Freyberger, J. (2015): Asymptotic theory for differentiated products demand models with many markets, Journal of Econometrics, 185, 162--181
2015
-
[28]
Goldberg, P. K. (1995): Product differentiation and oligopoly in international markets: The case of the US automobile industry, Econometrica, 891--951
1995
-
[29]
Grieco, P. L., C. Murry, J. Pinkse, and S. Sagl (2025): Optimal Estimation of Discrete Choice Demand Models with Consumer and Product Data, NBER Working Paper 33397
2025
-
[30]
Grieco, P. L., C. Murry, and A. Yurukoglu (2024): The evolution of market power in the us automobile industry, The Quarterly Journal of Economics, 139, 1201--1253
2024
-
[31]
Kuersteiner, and M
Hahn, J., G. Kuersteiner, and M. Mazzocco (2022): Joint time-series and cross-section limit theory under mixingale assumptions, Econometric Theory, 38, 942--958
2022
-
[32]
Han, S. and K. Lee (2025): Copyright and Competition: Estimating Supply and Demand with Unstructured Data, arXiv preprint arXiv:2501.16120
arXiv 2025
-
[33]
Hansen, B. E. (1996): Inference when a nuisance parameter is not identified under the null hypothesis, Econometrica, 64, 413--430
1996
-
[34]
Hausman, J. A. (1994): Valuation of new goods under perfect and imperfect competition, National Bureau of Economic Research Cambridge, Mass., USA
1994
-
[35]
(2025): Generative brand choice, Working Paper
Lee, K. (2025): Generative brand choice, Working Paper
2025
-
[36]
Lee, R. S. (2013): Vertical integration and exclusivity in platform and two-sided markets, American Economic Review, 103, 2960--3000
2013
-
[37]
McClure, and A
Magnolfi, L., J. McClure, and A. Sorensen (2025): Triplet embeddings for demand estimation, American Economic Journal: Microeconomics, 17, 282--307
2025
-
[38]
(2017): Targeted vouchers, competition among schools, and the academic achievement of poor students, Working Paper
Neilson, C. (2017): Targeted vouchers, competition among schools, and the academic achievement of poor students, Working Paper
2017
-
[39]
Neuman, A. M., Y. Xie, and Q. Sun (2023): Restricted Riemannian geometry for positive semidefinite matrices, Linear Algebra and its Applications, 665, 153--195
2023
-
[40]
(2000): Mergers with differentiated products: The case of the ready-to-eat cereal industry, The RAND Journal of Economics, 395--421
Nevo, A. (2000): Mergers with differentiated products: The case of the ready-to-eat cereal industry, The RAND Journal of Economics, 395--421
2000
-
[41]
--- -.1pt --- -.1pt --- (2001): Measuring market power in the ready-to-eat cereal industry, Econometrica, 69, 307--342
2001
-
[42]
Newey, W. K. (1994): The asymptotic variance of semiparametric estimators, Econometrica: Journal of the Econometric Society, 62, 1349--1382
1994
-
[43]
Newey, W. K. and D. McFadden (1994): Large sample estimation and hypothesis testing, Handbook of econometrics, 4, 2111--2245
1994
-
[44]
(2002): Quantifying the benefits of new products: The case of the minivan, Journal of political Economy, 110, 705--729
Petrin, A. (2002): Quantifying the benefits of new products: The case of the minivan, Journal of political Economy, 110, 705--729
2002
-
[45]
Zhang, J., W. Xue, Y. Yu, and Y. Tan (2023): Debiasing ML-or AI-Generated Regressors in Partial Linear Models, SSRN Working Paper 4636026
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.