{"id":"30b322a1-4b3b-49a2-a4e4-4b38384e46cb","arxiv_id":"2601.05374","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A demand-economics toolkit that debiases counterfactuals computed from imperfect ML-generated product attributes, using the model's own moments, plus two diagnostics for proxy quality — shown to lift closest-substitute prediction from 40% to 70% in an e-books experiment.","lead":"This paper gives economists a cheap post-estimation fix for demand models whose product attributes come from machine-learning embeddings: add a correction built from the model's own estimation equations so the answer no longer depends on which imperfect proxy was used. In a 10-product e-books experiment, the corrected model predicts consumers' closest substitutes correctly 70% of the time versus 40% without correction.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Local-misspecification condition o_p(T^{-1/4}) is not satisfied by fixed imperfect proxies; central debiasing guarantee needs a primitive condition or a modified asymptotic experiment.","rationale":"The reader identifies the local-misspecification condition γ̂ = γ₀ + o_p(T^{-1/4}) as the weakest assumption. My stress test agrees and sharpens the concern: for fixed proxies—the leading case in the paper—this condition is not merely hard to verify; it is generically false because γ̂ has a probability limit γ(θ*(ẽ), ẽ) different from γ₀. The paper's reliance on fine-tuning to justify the condition is not supported by any rate result, and the LM1 diagnostic cannot certify the required rate; it is a specification test whose power grows with T, while the local condition requires a rate statement. The simulations' degradation at ρ > 0.5 confirms that the assumption is not innocuous. This does not falsify the conditional theorems, which are coherent under the stated high-level assumption, but it does mean the practical claim of bias removal for imperfect embeddings is only conditionally established. The reader's CONDITIONAL verdict already captures this; my concern does not move the verdict, so I recommend UNCHANGED.","tokens_in":32964,"tokens_out":3850,"duration_ms":45359,"concrete_test":"Re-run the Section 4 simulation with a fixed proxy mismeasurement level ρ = 0.5 and sample sizes n = 10^3, 10^4, and 10^5, holding the latent attributes e₀ and the proxy ẽ fixed across replications (or drawing ẽ once and keeping it fixed). Compute the bias of κ̂_bc at each n. If the bias does not decay at roughly n^{-1/2}—or does not decay at all—then the condition γ̂ = γ₀ + o_p(n^{-1/4}) is violated for fixed proxies, and Proposition 5's conclusion does not cover the leading application. A complementary analytical check: in the scalar-e, two-product special case of Example 1, compute γ(θ*(ẽ), ẽ) − γ(θ₀, e₀) explicitly and verify it is nonzero for ẽ ≠ e₀.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 1 and Proposition 5 both hinge on the assumption that γ̂ = γ(θ̂, ẽ) satisfies γ̂ = γ₀ + o_p(T^{-1/4}). For the leading case in which ẽ is a fixed embedding computed from unstructured product data, this condition is not a mild regularity condition but a substantive restriction that will generally fail. Because θ̂ converges to some pseudo-true value θ*(ẽ) under the misspecified proxy model, γ̂ converges to γ(θ*(ẽ), ẽ), which is not equal to γ₀ = γ(θ₀, e₀) unless the proxy is asymptotically perfect or there is an unlikely coincidence in the composite parameter mapping. Nothing in the paper supplies primitive conditions on proxy construction that would deliver γ̂ − γ₀ = o_p(T^{-1/4}) for fixed ẽ. Remark 8 appeals to the possibility that ẽ is fine-tuned on the same data so that γ̂ approaches γ₀ as the sample size grows, but no rate for fine-tuning convergence is established, and off-the-shelf embeddings clearly do not satisfy it. The LM1 diagnostic does not rescue the assumption: Proposition 4 is proved under γ̂ = γ₀ + o_p(1), but under fixed proxy mismeasurement γ̂ is inconsistent, so LM1 grows at rate T (not log T) and the proposed C_T² = χ²_{0.95} log T threshold will reject with probability approaching one. A finite-sample below-threshold LM1 only indicates that T is not large enough for the misspecification to be detected; it cannot certify the required o_p(T^{-1/4}) rate. The simulations confirm the practical bite: in Section 4, the corrected estimator's bias increases markedly once mismeasurement exceeds ρ ≈ 0.5, exactly the regime in which the local condition fails. Thus the headline claim that the estimator 'removes first-order proxy bias' is only established in an asymptotic experiment where the proxy error itself shrinks with T, which is not the experiment corresponding to fixed black-box embeddings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops post-estimation bias corrections and diagnostics for demand counterfactuals when product attributes are proxied by embeddings or other imperfect measures. The key device is a reparameterization in which attributes and structural parameters enter choice probabilities only through a composite parameter γ; the naive counterfactual estimator is corrected by adding a linear combination of the model's own moment conditions, with weights chosen to remove first-order dependence on γ̂ and to minimize asymptotic variance. The main theoretical results (Propositions 1, 3, 5) state that, provided γ̂ = γ₀ + o_p(T^{-1/4}), the corrected estimator is asymptotically centered at the true counterfactual with a variance independent of γ̂, θ̂, and the proxy. Two LM-type diagnostics are proposed to assess whether γ̂ is sufficiently close to γ₀ and whether the proxy dimension is adequate. Simulations and an application to ebook choice data illustrate the method, with closest-substitute hit rates improving from 40% to 70% in the preferred specification.","tokens_in":33237,"tokens_out":3944,"duration_ms":49517,"significance":"If the maintained rate condition holds, the paper gives a practically useful, computationally light way to debias counterfactual inference in demand models with imperfect proxies. The closed-form standard errors, efficiency property within a natural class, and accommodation of data-dependent proxies are genuine strengths. The empirical demonstration against ground-truth second choices is valuable. However, the central theoretical guarantee is conditional on the proxy-induced estimator γ̂ converging to γ₀ at o_p(T^{-1/4}), and the paper does not provide primitive conditions ensuring this for the leading case of fixed embeddings. The diagnostics are useful heuristics but, as formalized, cannot certify the required rate. These issues are load-bearing for the paper's headline claim of valid inference for unstructured-data proxies, though they are fixable by a more careful statement of scope and by adding rate conditions or a local-misspecification framework.","major_comments":[{"comment":"The condition γ̂ = γ₀ + o_p(T^{-1/4}) is maintained without a primitive justification. For a fixed, off-the-shelf embedding ẽ, the natural limit is γ̂ → γ(θ*(ẽ), ẽ), where θ*(ẽ) is the pseudo-true value; nothing guarantees γ(θ*(ẽ), ẽ) = γ₀. Remark 8 appeals to fine-tuning on the same data, but no rate for the convergence of a data-dependent ẽ to e₀ is established, so the required o_p(T^{-1/4}) shrinkage is not derived. Since the paper's leading case is exactly such proxies, the central asymptotic claim is not supported for the main application without additional assumptions. I recommend either providing concrete primitive conditions on proxy construction (e.g., embeddings refined at a controlled rate) or explicitly recasting the theory as local-misspecification asymptotics and stating that the debiasing guarantee applies only when the proxy error is smaller than the sampling error at the","section":"§6.1, Propositions 1 and 3; §6.2, Proposition 5; Remark 8"},{"comment":"The LM1 diagnostic is presented as validating the condition that γ̂ is sufficiently close to γ₀. But Proposition 4 is proved under γ̂ = γ₀ + o_p(1). Under fixed proxy mismeasurement γ̂ is inconsistent, so LM1 diverges at rate T and the proposed threshold C_T² = χ²_{dim γ,0.95} log T will reject with probability approaching one. A finite-sample non-rejection then only indicates that the sample is too small to detect the misspecification; it cannot certify the required o_p(T^{-1/4}) rate. The wording in the practitioner's guides (e.g., \"conclude that ẽ is sufficiently close to e₀\") overstates what the diagnostic can establish. The diagnostic is still useful as a specification check, but the paper should state this limitation and provide either a formal local-power analysis or a bound that is valid under fixed misspecification.","section":"§2.3.1 and §3.3.1, Proposition 4 and Proposition 6"},{"comment":"Even when the bias correction removes the first-order term, the remaining bias is of order O(∥γ̂−γ₀∥²) under fixed proxy error. If γ̂ fails to converge to γ₀, this second-order bias can be non-negligible; the simulations in Figure 2 indeed show that for ρ > 0.5 the corrected estimator's bias increases. The paper acknowledges this behavior informally, but the theoretical sections do not make explicit the consequence that the distributional results are only approximate for fixed mismeasurement and that the approximation degrades as ∥γ̂−γ₀∥ grows. Please state this as a formal caveat near Propositions 1 and 5, and quantify the remainder as a function of ∥γ̂−γ₀∥.","section":"§6.1, proof of Proposition 1; §4 simulations"},{"comment":"The choice C_T² = χ²_{dim(γ),0.95} log T is heuristic. Proposition 4 only gives statements 'wpa1' conditional on a chi-square random variable being below ϵ²C_T²; it does not provide the distributional approximation needed to calibrate the threshold for controlling a false-acceptance probability. The paper should either derive the large-deviation or local-alternative properties of LM1 that justify this threshold, or present it as an ad hoc rule. This matters because the threshold is used in the application to select among specifications.","section":"§6.1.2, Proposition 4; §2.3.1, threshold choice"}],"minor_comments":[{"comment":"The statement says 'Let Assumption 3 hold', but the result is for the no-microdata case and should refer to Assumption 2.","section":"Proposition 3 (p. 35)"},{"comment":"There is a typo: 'Let Assumptions 3 and 4 bold hold' should be 'both hold'. Also the display uses 'LM' without a subscript; it should be 'LM1'.","section":"Proposition 7 (p. 39)"},{"comment":"The variance estimator (15) is stated without derivation. It would help readers to see that it is the sample analogue of the variance expression in Proposition 1, including the cross-term between k_t and the moments.","section":"§2.2, Remark 2 and eq. (15)"},{"comment":"The text says 'when mismeasurement becomes very large (ρ>0.5), the bias correction starts to also perform worse,' but the figure appears to show the onset around ρ=0.4–0.5. Please align the text with the simulation grid.","section":"§4, Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is well executed conditional on the rate condition, but the central unresolved point is whether that condition is plausible or verifiable for fixed embeddings. The authors should either add primitives that deliver γ̂−γ₀ = o_p(T^{-1/4}) for a data-dependent proxy construction or carefully restrict the claims to local misspecification. The LM1 diagnostic as currently formulated cannot carry the weight placed on it. I would be comfortable with acceptance after this is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a solid methodology paper for empirical IO and marketing people who feed ML embeddings into demand models and worry that the embeddings are mismeasured proxies. The new pieces are real: the composite-parameter reparameterization gamma(theta,e), first-order debiasing of counterfactuals using the model's own moments, the efficiency result for the corrected estimator, and the two LM diagnostics, especially LM2's dimension test with simulated critical values. The proofs in Appendix A are coherent; I checked the key expansions in Propositions 1 and 5. The in-sample orthogonality conditions (36)/(37) are what make the correction work, and the argument is standard but well executed. The simulations are honest: bias and RMSE improve for rho up to about 0.5, and the paper flags the deterioration beyond that. The empirical application uses held-out second choices as a ground truth and gets the closest-substitute hit rate from 40% to 70% in the preferred review-text specification. That is meaningful.\n\nThe main soft spot is the one in the stress-test: the entire guarantee rests on gamma-hat = gamma_0 + o_p(T^{-1/4}), and for a fixed off-the-shelf embedding gamma-hat typically converges to a pseudo-true value gamma(theta*(e-tilde), e-tilde) that is not gamma_0. The paper does not supply primitive conditions on proxy construction that would deliver the required rate. It is transparent about this in Remark 8 and in the practitioner discussion, but the LM1 diagnostic with the chi-square log T threshold is heuristic; it cannot certify the o_p(T^{-1/4}) rate, only reject in large samples. That is a genuine gap between the theory and the leading application. It is not fatal: the paper claims a first-order correction in a local-misspecification regime, and the simulations confirm that regime is practically relevant. But a referee should push for either primitive conditions (e.g., fine-tuning rates) or a reformulated asymptotic experiment that treats proxy error as drifting with T.\n\nOther soft spots are minor. The BLP/instrumental-variables case is derived but never simulated, so the flagship market-level setting lacks finite-sample evidence. The empirical hit-rate functional is non-smooth and the theory covers smooth counterfactuals, and there is no uncertainty quantification on the 40% to 70% improvement over 10 books. No code or data are shipped, and the application reuses the authors' own prior dataset. None of this undercuts the core contribution.\n\nWho should read it: anyone estimating demand with unstructured data, and econometricians working on two-step debiasing. It deserves a serious referee, with the expectation that the local-misspecification condition gets tightened. I would send it to review.","headline":"A genuinely useful debiasing toolkit for demand counterfactuals with ML proxies; the theory is coherent, but the key local-misspecification condition needs primitive support before the application claims are fully load-bearing.","tokens_in":33937,"tokens_out":2578,"would_cite":true,"duration_ms":28418,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that the bias that imperfect product proxies inject into demand counterfactuals can be removed with a post-estimation correction whose standard errors ignore how the proxies were built.","keywords":["demand estimation","unstructured data","product embeddings","bias correction","counterfactuals","proxy variables","Lagrange multiplier diagnostics","discrete choice"],"falsifier":"Run a controlled simulation with true latent attributes known and proxy noise at roughly ρ=0.5 or higher, increasing sample size; if the bias-corrected estimator's bias does not shrink at the claimed rate or does not remain far below the naive estimator's bias, the local condition is not doing the work in that regime.","tokens_in":32656,"feed_emoji":"📈","tokens_out":7406,"duration_ms":76163,"temperature":0.7,"pith_summary":"Demand models increasingly use embeddings of images, text, and reviews as product attributes, but these are proxies for the true, latent dimensions that drive substitution. This paper argues that treating proxies as truth biases counterfactual predictions, and shows how to correct that bias after estimation with a simple additive adjustment. The corrected estimator is asymptotically centered on the true counterfactual and is efficient, and its variance does not depend on how the proxies were constructed, so fine-tuning or data-dependent embeddings require no standard-error adjustment. Two LM diagnostics help practitioners choose which proxies and how many dimensions to use. In an e-book choice experiment, the correction raises the rate of correctly predicted closest substitutes from 40% to 60–70% for the best review-text specification.","feed_headline":"Bias correction lifts substitute prediction from 40% to 70%","feed_subtitle":"A post-estimation adjustment purges proxy mismeasurement from demand counterfactuals and shows which embeddings to trust.","key_machinery":"The key machinery is the composite parameter γ(θ,e), the low-dimensional combination of structural parameters and latent product attributes through which all choice probabilities and counterfactuals enter the model. By expressing both the naive estimator and the estimation moments as functions of γ̂, the paper chooses correction weights so that the adjusted counterfactual has zero first-order sensitivity to γ̂; this makes the choice of proxy irrelevant to the estimator's leading bias term. The same object drives the diagnostics: LM1 compares γ̂ to the set of composite parameters spanned by the proxies, and LM2 augments the proxy space with an extra direction to test whether the proxy dimensi","core_discovery":"The paper's central claim is that mismeasurement of product attributes by proxies is a model misspecification, not a classical measurement-error problem, and can be neutralized by reparameterizing the demand model in terms of a composite parameter γ that bundles structural parameters and latent attributes. Once the naive estimate γ̂ = γ(θ̂, ẽ) is in hand, the counterfactual estimator is adjusted by subtracting weighted averages of the estimation moments, with weights chosen to make the adjusted estimator's first-order dependence on γ̂ vanish. Under the condition that γ̂ is within a neighbourhood of γ0 whose radius is negligible relative to sampling error, the adjusted estimator is asymptoti","pith_inferences":["A natural testable extension is to calibrate the LM1 threshold against a small validation set with known attribute values, rather than the heuristic χ² log T cutoff, to see whether better proxy-selection decisions result.","The paper's logic implies that the cost of searching over many embedding choices is lower than usually assumed: model selection should emphasize the LM diagnostics, since the target counterfactual variance is proxy-independent to first order.","The method is best suited to attributes that are fixed product characteristics; applying it to time-varying or context-dependent attributes would require the composite-parameter mapping to hold within each market or individual setting, which the paper does not claim."],"forward_implications":["Bias-corrected counterfactuals are centered at the true value and come with closed-form standard errors, so no bootstrap or re-estimation is needed after the correction.","Standard errors remain valid when embeddings are fine-tuned on the choice data, because the asymptotic distribution does not depend on the proxy to first order.","The LM1 and LM2 diagnostics give a practical answer to which unstructured-data proxies to use and how many principal components to keep.","In the e-book experiment, the correction raises closest-substitute hit rates from 40% to 60–70% and improves or ties 11 of 13 specifications ruled in by the dimension diagnostic.","Even when mismeasurement is not a concern, the estimator offers efficient counterfactual inference with simple standard errors."],"fun_headline_variants":["Proxy errors no longer distort demand counterfactuals","Fix for biased substitution predictions from noisy proxies","Post-hoc adjustment purges proxy bias in demand models","Diagnose and correct proxy mismeasurement in demand","Better substitution forecasts from flawed product data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the composite parameter derived from the proxies lands close enough to the true value — closer than the fourth root of the sample size; if the proxies are too noisy, nothing in the data forces this, and the correction loses its centering property.","fun_headline_variants_meta":{"raw":{"variants":["Proxy errors no longer distort demand counterfactuals","Fix for biased substitution predictions from noisy proxies","Post-hoc adjustment purges proxy bias in demand models","Diagnose and correct proxy mismeasurement in demand","Better substitution forecasts from flawed product data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1316,"prompt_tokens":620,"completion_tokens":696,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":364,"completion_tokens_details":{"reasoning_tokens":625}},"tokens_in":364,"tokens_out":696,"duration_ms":8226,"temperature":1.0,"reasoning_tokens":625,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T11:40:54.185303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled simulation with true latent attributes known and proxy noise at roughly ρ=0.5 or higher, increasing sample size; if the bias-corrected estimator's bias does not shrink at the claimed rate or does not remain far below the naive estimator's bias, the local condition is not doing the work in that regime.","supporting_citations":[],"review_version":1}