REVIEW 4 major objections 7 minor 29 references
Agentic Economic Modeling
T0 review · 4 major / 7 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Synthetic LLM choices, corrected with 10% of human data, reproduce experimental treatment effects
desk verdict A useful proof-of-concept for bias-corrected LLM choice data, with one genuinely held-out regional validation and an honest appendix that undermines the time-wise headline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the mixture bias-correction model: given a region z, each (persona, order) pair yields an LLM choice vector cl_i; an attention-style weighting (Equation 5) computes a weighted average of those one-hot choices to form a predicted region-level share sl_z, and the weights are trained by minimizing KL divergence between sl_z and the observed human share sh_z (Equation 6). The weight structure is the same form as a random-coefficient discrete choice mixture: the LLM agents play the role of decision rules, and the learned weights recover the community's composition. In the conjoint setting the correction is instead a learned conditional distribution P(y|x,z) (a logistic
What would settle it
Run the region-wise procedure on a second regional experiment with a known treatment effect and check whether the ZOOD prediction falls within the full human experiment's 95% CI; specifically, if holding out 10% of ZIP3s and training the mixture model produces an estimate outside [-52, -68] bps (or opposite sign) while the human experiment is -60±8 bps, the claimed generalization fails. A quicker probe: permute the mapping between personas and ZIP3s in ZOOD and see if predicted shares change; if they do, the correction is using region identity rather than transferable persona structure.
Extended reading notes
Core claim
The central claim is that bias-corrected synthetic choices are an adequate substitute for the bulk of human experimental data in recovering economic parameters. On the paper's own terms: a mixture model that reweights LLM persona choices to match observed regional shares, trained on 10% of geographic regions, identifies the same treatment effect in unseen regions as the full human experiment, and in conjoint analysis a correction model trained on 10% of respondents yields more accurate part-worth estimates than using only the primary human sample. The paper frames this as a step toward LLM-based counterfactual generation: rather than treating an LLM as a perfect human stand-in, it uses the L
Load-bearing premise
The learned bias-correction mapping must generalize out-of-domain: weights trained on 10% of regions (or one day of data) must identify counterfactual shares in regions and times where no human data exist, and the paper itself states this as an assumption ('We assume and hope...').
Editorial extensions
If this is right
- If AEM is correct, a regional experiment could be run in 10% of areas, with LLM simulation filling in the rest, and still yield treatment effects statistically indistinguishable from the full experiment (e.g., -65±10 bps vs -60±8 bps).
- Conjoint demand estimation no longer requires surveying every respondent: a 10% human sample used to learn the correction gives lower part-worth error than using the full small primary sample, while uncorrected LLM choices actively distort estimates.
- The pipeline-wide bootstrap (20 resamples of personas/regions, retraining, re-inference) shows that confidence intervals in LLM experiments must come from the whole Generation–Correction–Inference chain, not from sampling variance alone.
- One day of human data plus LLM augmentation yields a significant treatment effect (-24 bps, p<1e-5) where the human-only day-one estimate is insignificant (-17 bps, p=0.20), implying early termination of RCTs may be feasible.
- The same Generation–Correction–Inference structure is claimed to apply to any causal inference problem with missing counterfactuals, not just the two RCT settings validated.
Reading between the lines
- If the learned correction transfers across regions, AEM implies a new pre-registration option: simulate a policy change on held-out geographies before committing to a full rollout, with calibration data from a small pilot.
- The success of the weighted-mixture correction over a single black-box model suggests that heterogeneity in LLM personas, not the LLM itself, is what matches real communities; one testable extension is whether increasing persona diversity monotonically improves correction accuracy.
- A natural stress test the paper does not run: train the correction on ZID regions with deliberately shifted demographic compositions (e.g., income or SSD availability) and check whether the held-out estimate degrades, which would reveal the limits of the 'hope' stated in Section 5.2.
- Extending AEM to temporal dynamics—modeling day/week trends as an additional persona or covariate—is a direct next step, since Section C.4 concedes the current LLM experiment cannot capture temporal changes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Agentic Economic Modeling (AEM), a Generation–Correction–Inference pipeline in which LLM-generated choice data from diverse personas are bias-corrected using a small sample of human labels, and then standard econometric estimators are applied to the corrected synthetic data. The framework is validated in two real-world settings: a large conjoint study, where a correction model fitted on 10% of the data reportedly lowers demand-parameter error relative to using primary data alone, and a regional field experiment with a delivery-option treatment, where a mixture model calibrated on 10% of ZIP3 regions produces a treatment-effect estimate of -65±10 bps on hold-out regions, reportedly matching the full human experiment (-60±8 bps). The paper also reports a time-wise extrapolation using only day-one human data, yielding -24 bps with a tight confidence interval. The central claims are that AEM can substantially reduce the scale and duration of RCTs while recovering the same economic parameters.
Significance. If the claims were fully established, AEM would be a practically important contribution to experimental economics and empirical industrial organization: it offers a concrete method for combining cheap LLM-generated choices with small human samples and standard econometric inference, with validation on a real field experiment and a large conjoint survey. The paper has several genuine strengths: it explicitly separates generation, correction, and inference; it bootstraps the full pipeline rather than reporting nominal sampling error alone; and it candidly states key assumptions, including the generalization assumption in §5.2 and the temporal-dynamics limitation in Appendix C.4. These strengths are, however, undermined by the fact that the headline 'out-of-domain' regional result rests on an untested interpolation assumption, the conjoint validation uses one product without uncertainty quantification, and the time-wise result is openly acknowledged not to recover the full-experiment parameter. The framework is promising, but the current evidence does not support the paper's strongest claims.
major comments (4)
- [§5.2, Table 4] The region-wise headline ('out-of-domain treatment effect of -65±10 bps') rests on an assumption that is stated but not tested. The ZID/ZOOD split is a random 10/90 partition within a single national experiment, in the same time window, using the same 3,806-order pool (Appendix C.1.2: each ZIP3's 42-order set is mostly drawn from other ZIP3s). This tests interpolation across randomly similar regions, not invariance of the correction mapping to a genuine domain shift (e.g., different SSD availability, local price sensitivity, or time period). The bootstrap in §5.5 resamples within ZID/ZOOD and cannot quantify this unmodeled shift. Because the DiD estimate is computed on predicted shares, differential prediction error in treatment vs. control can masquerade as a treatment effect. The paper's own sentence in §5.2 — 'We assume and hope that the learned relationship can generalize' — confirms
- [§4.4-4.5, Table 2] The conjoint validation is based on a single product and reports point MAPE changes without confidence intervals. Appendix B.1 says 'We randomly select one testing product for initial evaluation.' With one product, the claim that 10% human data 'lowers the error' is a single draw; no statement of uncertainty or replication across products is given, and the table's entries are changes relative to β_primary rather than raw MAPEs, despite the heading 'MAPE(β)'. The sign pattern is also unexplained: raw LLM choices alone worsen MAPE by +0.11, while naively combining primary human choices with raw LLM choices improves MAPE by -8.42. Please report product-level variation and CIs, or mark the conjoint result as exploratory and explain the sign pattern.
- [§5.6 and Appendix C.4] The time-wise result is internally acknowledged not to recover the full-experiment parameter. Appendix C.4 states that the model is 'unable to model the temporal changes of the regional experiment' and that a gap remains between the short-period estimate and the full experiment. The reported -24 bps (95% CI [-26,-22]) is 36 bps away from the full-experiment -60 bps; it only tightens the CI around a day-one estimate that is itself far from the target. Reporting this as an 'improvement' over the human-only baseline is misleading unless the goal is stated as variance reduction conditional on a biased anchor. Please either remove the time-wise claim from the headline or reframe it explicitly as precision improvement, not parameter recovery.
- [§5.4] Same-Day share is selected as the primary outcome because the full human experiment showed it to be significant ('Other outcomes in the experiment showed no significant change. Therefore, we prioritize...'). Since the evaluation criterion is chosen after observing the full data, the reported success on this outcome is vulnerable to selection on the target. Please report all pre-specified or analogous outcomes (including Next-Day/Second-Day, which the paper says moved by +25 bps) or justify the selection ex ante. Without this, the match in Table 4 cannot be interpreted as out-of-sample confirmation.
minor comments (7)
- [Abstract] The abstract says 'millions of observations' for the conjoint study, but the actual validation uses one test product. Please clarify the scope in the abstract or main text.
- [Eq. (5)] The index notation in the attention weight is inconsistent (h_i vs. h_k in numerator/denominator). Please define the dimensions of U and V and the batch indexing.
- [Table 2] Specify in the table caption that entries are changes in MAPE relative to β_primary, not MAPEs themselves. The current heading is misleading.
- [Appendix C.1.2] State in the main text that most of each ZIP3's 42 orders come from other ZIP3s. This is important for the region-wise interpretation: the ZOOD 'regions' are not independent order pools.
- [§5.5] Report the actual 20 bootstrap estimates and the degrees of freedom used for the t-based CI. A t-distribution with 19 df is sensitive to outliers; a percentile or bias-corrected CI would be more robust.
- [Appendix C.1.1] The text refers to 'the prompt used to generate personas is provided in Appendix A.1. of the paper,' but Appendix A of this manuscript is a summary table, not the prompt. Include the prompt or fix the reference.
- [Appendix C.2.3] The acronym 'SSD' is defined as 'Sub Same-Day' but the intended meaning is unclear. Please spell out the term in full.
Circularity Check
No material circularity: the pipeline is a supervised imputation/calibration procedure evaluated on held-out human labels and regions; the OOD generalization caveat in §5.2 is an external-validity limitation, not a derivation-level circularity.
full rationale
The derivation chain is not circular. The correction mapping f:(x,z)->yhat is estimated on primary human labels and applied to auxiliary (x,z) pairs whose human labels are withheld; conjoint performance is measured against the held-out 90% of human choices, and regional performance is measured on disjoint ZOOD ZIP3s. Thus no predicted treatment effect is equal to its training target by construction. Equations (5)-(6) fit mixture weights to ZID human shares and then reuse the learned weight function on ZOOD triplets, which is a genuine extrapolation step. The paper explicitly flags this assumption in §5.2 ('We assume and hope that the learned relationship can generalize across variations in persona distributions') and §C.4 admits time-wise models cannot capture temporal dynamics; these are validity threats about out-of-domain transfer, not circular definitions. The cited prior work (e.g., [27] for the MAPE benchmark, [7] for persona generation) is not load-bearing, and no self-citation chain forces the conclusions. If the ZID-to-ZOOD mapping shifts, the estimates could be biased, but that would be an identification or generalization failure, not a tautology.
Assumptions & free parameters
free parameters (5)
- Mixture model attention weights U, V =
not reported
- Integrated model MLP weights =
not reported
- Logistic regression correction coefficients (conjoint) =
not reported
- Architecture hyperparameters =
hidden dim 64, attention dim 32, LR 1e-5
- Sampling design choices =
n_z=42, 3,806 orders, r=0.1, first x=7 or 1 days, 17k personas
assumptions (5)
- domain assumption LLM-generated choices are informative about human choices and personas span the space of human decision rules.
- domain assumption The correction mapping learned on ZID regions generalizes to ZOOD regions.
- domain assumption The difference-in-differences parallel trends assumption holds for delivery option shares.
- ad hoc to paper Temporal dynamics are ignorable for the time-wise extrapolation.
- ad hoc to paper Same-Day delivery share is the correct primary outcome.
invented entities (1)
-
LLM-generated customer personas (e.g., 'a 44-year-old retail sales associate')
Cite this review
Pith. "Pith review of Agentic Economic Modeling." pith.science (2026). https://pith.science/paper/J5Y4M3KP
@misc{pith2026251025743,
author = {Pith},
title = {Pith review of: Agentic Economic Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/J5Y4M3KP}},
note = {Machine review of arXiv:2510.25743}
}
abstract
We introduce Agentic Economic Modeling (AEM), a framework that aligns synthetic LLM choices with small-sample human evidence for econometric inference. AEM first generates task-conditioned synthetic choices via LLMs, then learns a bias-correction mapping from task features and raw LLM choices to human-aligned choices, upon which standard econometric estimators perform inference to recover demand elasticities and treatment effects. We validate AEM in two experiments. In a large scale conjoint study, using only 10% of the original data to fit the correction model lowers the error of the demand-parameter estimates, while uncorrected LLM choices increase the errors. In a regional field experiment, a mixture model calibrated on 10% of geographic regions estimates a treatment effect of -65$\pm$10 bps on the hold-out regions, closely matching the full human experiment (-60$\pm$8 bps). These results demonstrate AEM's potential to improve RCT efficiency and represent a step toward LLM-based counterfactual generation.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[27]
Mengxin Wang, Dennis J Zhang, and Heng Zhang. Large language models for market research: A data- augmentation approach.arXiv preprint arXiv:2412.19363, 2024
arXiv 2024
-
[1]
Ai–human hybrids for marketing research: Leveraging large language models (llms) as collaborators.Journal of Marketing, 89(2):43–70, 2025
Neeraj Arora, Ishita Chakraborty, and Yohei Nishimura. Ai–human hybrids for marketing research: Leveraging large language models (llms) as collaborators.Journal of Marketing, 89(2):43–70, 2025
2025
-
[2]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate.arXiv preprint arXiv:1409.0473, 2014
arXiv 2014
-
[3]
Machine bias
Julien Boelaert, Samuel Coavoux, Etienne Ollion, Ivaylo Petev, and Patrick Präg. Machine bias. how do generative language models answer opinion polls?Sociological Methods & Research, page 00491241251330582, 2025
2025
-
[4]
The mixed subjects design: Treating large language models as potentially informative observations.Sociological Methods & Research, page 00491241251326865, 2025
David Broska, Michael Howes, and Austin van Loon. The mixed subjects design: Treating large language models as potentially informative observations.Sociological Methods & Research, page 00491241251326865, 2025
2025
-
[5]
Mixture-of-personas language models for population simulation.arXiv preprint arXiv:2504.05019, 2025
Ngoc Bui, Hieu Trung Nguyen, Shantanu Kumar, Julian Theodore, Weikang Qiu, Viet Anh Nguyen, and Rex Ying. Mixture-of-personas language models for population simulation.arXiv preprint arXiv:2504.05019, 2025
arXiv 2025
-
[6]
The next generation of experimental research with llms.Nature Human Behaviour, pages 1–3, 2025
Gary Charness, Brian Jabarian, and John A List. The next generation of experimental research with llms.Nature Human Behaviour, pages 1–3, 2025
2025
-
[7]
Chaoran Chen, Weijun Li, Wenxin Song, Yanfang Ye, Yaxing Yao, and Toby Jia-Jun Li. An empathy-based sandbox approach to bridge the privacy gap among attitudes, goals, knowledge, and behaviors.Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–28, 2024
2024
Show all 29 references
-
[8]
The emergence of economic rationality of gpt
Yiting Chen, Tracy Xiao Liu, You Shan, and Songfa Zhong. The emergence of economic rationality of gpt. 2023
2023
-
[9]
Integrating generative artificial intelligence into social science research: Measurement, prompting, and simulation.Sociological Methods & Research, page 00491241251339184, 2025
Thomas Davidson and Daniel Karell. Integrating generative artificial intelligence into social science research: Measurement, prompting, and simulation.Sociological Methods & Research, page 00491241251339184, 2025
2025
-
[10]
Maria del Rio-Chanona, Marco Pangallo, and Cars Hommes
R. Maria del Rio-Chanona, Marco Pangallo, and Cars Hommes. Can generative ai agents behave like humans? evidence from laboratory market experiments, 2025
2025
-
[11]
Hashimoto
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Alpacafarm: A simulation framework for methods that learn from human feedback. 2024
2024
-
[12]
Frontiers: Can large language models capture human preferences?Marketing Science, 43(4):709–722, 2024
Ali Goli and Amandeep Singh. Frontiers: Can large language models capture human preferences?Marketing Science, 43(4):709–722, 2024
2024
-
[13]
Large language models as simulated economic agents: What can we learn from homo silicus? http://www.nber .org/papers/w31122, (31122), 2023
John J Horton. Large language models as simulated economic agents: What can we learn from homo silicus? http://www.nber .org/papers/w31122, (31122), 2023. NBER Working Paper
2023
-
[14]
Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv:2404.06654, 2024
2024 arXiv
-
[15]
Learning to be homo economicus: Can an llm learn preferences from choice, 2024
Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, and Hector Tzavellas. Learning to be homo economicus: Can an llm learn preferences from choice, 2024
2024
-
[16]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017
2017
-
[17]
Generative ai for economic research: Use cases and implications for economists.Journal of Economic Literature, 2023
Anton Korinek. Generative ai for economic research: Use cases and implications for economists.Journal of Economic Literature, 2023
2023
-
[18]
Simulating subjects: The promise and peril of artificial intelligence stand-ins for social agents and interactions.Sociological Methods & Research, page 00491241251337316, 2025
Austin C Kozlowski and James Evans. Simulating subjects: The promise and peril of artificial intelligence stand-ins for social agents and interactions.Sociological Methods & Research, page 00491241251337316, 2025
2025
-
[19]
Can llms mimic human-like mental accounting and behavioral biases?Available at SSRN 4705130, 2024
Yan Leng. Can llms mimic human-like mental accounting and behavioral biases?Available at SSRN 4705130, 2024
2024
-
[20]
Do llm agents exhibit social behavior?arXiv preprint arXiv:2312.15198, 2023
Yan Leng and Yuan Yuan. Do llm agents exhibit social behavior?arXiv preprint arXiv:2312.15198, 2023
2023 arXiv
-
[21]
Llm generated persona is a promise with a catch, 2025
Ang Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng. Llm generated persona is a promise with a catch, 2025
2025
-
[22]
Llm generated persona is a promise with a catch
Ang Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng. Llm generated persona is a promise with a catch. arXiv preprint arXiv:2503.16527, 2025
2025 arXiv
-
[23]
Balancing large language model alignment and algorithmic fidelity in social science research.Sociological Methods & Research, 54(3):1110–1155, 2025
Alex Lyman, Bryce Hepner, Lisa P Argyle, Ethan C Busby, Joshua R Gubler, and David Wingate. Balancing large language model alignment and algorithmic fidelity in social science research.Sociological Methods & Research, 54(3):1110–1155, 2025
2025
-
[24]
Econometric models for probabilistic choice among products.The Journal of Business, 53(3):S13–S29, 1980
Daniel McFadden. Econometric models for probabilistic choice among products.The Journal of Business, 53(3):S13–S29, 1980. 12 Agentic Economic ModelingA PREPRINT
1980
-
[25]
Sentence-bert: Sentence embeddings using siamese bert-networks.arXiv preprint arXiv:1908.10084, 2019
Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks.arXiv preprint arXiv:1908.10084, 2019
1908 arXiv
-
[26]
Theorizing with large language models
Matteo Tranchero, Cecil-Francis Brenninkmeijer, Arul Murugan, and Abhishek Nagaraj. Theorizing with large language models. Working Paper 33033, National Bureau of Economic Research, October 2024
2024
-
[28]
Prefer not to buy it
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025. 13 Agentic Economic ModelingA PREPRINT A Summary of the Two Tasks The summary of two task...
2025 arXiv
-
[29]
We calculate the MAPE between the ground truth share and estimated share of different groups
Shares in ZIP3 areas where SSD was launched and assigned to the treatment group 2) Shares in ZIP3 areas where SSD was launched and assigned to the control group 3) Shares in ZIP3 areas where SSD was not launched and assigned to the treatment group 4) Shares in ZIP3 areas where...
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.