{"id":"e45a929e-cca7-4052-ba18-33cf7f48d7f1","arxiv_id":"2505.17648","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LLM-based agents seeded with demographic data, prior expectations, and social media information reproduce the shape of human macroeconomic expectation distributions in three survey experiments.","lead":"This paper builds software agents from large language models that answer macroeconomic survey questions as households or experts, then compares their answers with real human responses. It finds that the agents' expectation distributions broadly match human surveys, but that feeding in respondents' prior expectations is the main driver of that match.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing prior-only baseline leaves open that PEPM inputs, not the LLM architecture, drive the distributional matches in Figures 5 and 12.","rationale":"The paper is carefully executed and reports extensive ablations and out-of-sample checks. The mechanism analysis (selective recall, DAGs) is genuinely informative and is not directly threatened by the prior-only baseline. However, the headline claim is specifically about expectation distributions and about the LLM Agent architecture outperforming prompt engineering. The PEPM is the one module that injects contemporaneous survey expectations into agents, and §6 shows it is the main driver of distributional fit. The missing control is not a small omission: it determines whether the distributional success is attributable to the proposed architecture or to the input data. I therefore agree with the reader's weakest_assumption. A prior-only baseline is straightforward to add and would settle the interpretation. Because the paper contains independent evidence—the out-of-sample 2025 MSC pre-estimation and the qualitative mental-model analysis—the appropriate verdict remains conditional pending the baseline comparison, not rejection. My read does not change the reader's conditional verdict.","tokens_in":47760,"tokens_out":4043,"duration_ms":48902,"concrete_test":"Compute a prior-only baseline in the information-provision experiment: for each Chopra et al. respondent, take their elicited pre-treatment 10-year home-price expectation as the predicted posterior (or apply a minimal Bayesian shrinkage toward the treatment forecast using their stated confidence), then build the histogram of predicted posterior means and compare it to the human posterior distribution exactly as in Figure 5(c), using the same Pearson and cosine metrics. If the prior-only baseline achieves similarity within 0.05 of the LLM Agents' reported values, the modular architecture's distributional contribution is not established. For completeness, repeat using the 2024 MSC empirical distribution as a direct forecast of the 2025 MSC target and compare with Figure 5(d).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central distributional claim rests on comparisons between LLM-Agent outputs and human distributions in §5.1 (Figure 5). The PEPM (§3.1, §3.2) feeds each agent the contemporaneous priors of the very populations used as benchmarks: 2019 MSC/SPF priors for the Andre vignettes, 2024 Chopra et al. priors for the information-provision experiment, and 2024 MSC priors for the 2025 MSC pre-estimation. No baseline is reported that outputs or statistically transforms these priors directly into an expectation distribution. The ablation in §6 (Figure 12) removes PEPM and shows degradation, but that only shows priors matter; it does not show the LLM modular architecture matters. For the information-provision experiment, the target is a posterior distribution elicited after a small informational treatment, and the prior distribution may already resemble the posterior; a prior-only baseline could therefore match Figure 5(c) almost as well. If so, the 'highly similar distributions' and the claimed superiority over prompt-engineered foundation models would be an artifact of injecting the answer-relevant input distribution, not evidence for the agent architecture. The out-of-sample MSC design is independent support, but even there the 2024 prior distribution may be a strong forecast of 2025, so the same baseline is needed. The qualitative mental-model results are less affected, but they do not carry the headline distributional claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a modular framework ('UNITE') for simulating macroeconomic expectations in survey experiments using LLM-based agents. The authors construct Household Agents equipped with a Personal Characteristics Module, a Prior Expectations & Perceptions Module, and a Social Media Information Module, and Expert Agents with a Professional Background Module, a Prior Expectations & Perceptions Module, and a Knowledge Acquisition Module. They validate the framework by replicating three survey designs: the Andre et al. (2022) hypothetical vignette experiments on inflation and unemployment expectations of households and experts, the Chopra et al. (2025) information-provision experiment on home price expectations of homeowners and renters, and a pre-estimation exercise for the 2025 Michigan Survey of Consumers. The evaluation uses distributional similarity metrics (Pearson correlation and cosine similarity) between simulated and human distributions, qualitative analysis of open-ended responses, and mental-model comparisons via causal DAGs. The paper concludes that LLM Agents generate expectation distributions highly similar to human data, outperform foundation models relying simply on prompt engineering, and that the prior expectations module is crucial for distributional matching while other modules drive human-like thought processes.","tokens_in":48065,"tokens_out":4088,"duration_ms":40948,"significance":"If the central claim were established, this framework would be a useful, low-cost, and scalable complement to traditional survey experiments on expectation formation. The paper has several strengths: it covers three distinct and representative experimental designs, includes a detailed modular architecture, provides an ablation study (Figure 12 and Appendix Figures A.20-A.23), and attempts an out-of-sample pre-estimation exercise for the 2025 MSC. The mental-model analysis via DAGs and the comparison with foundation models are also valuable contributions. However, the central distributional claim is currently under-identified. The Prior Expectations & Perceptions Module feeds each agent the contemporaneous priors of the very populations used as benchmarks, yet no baseline is reported that outputs or statistically transforms these priors directly into an expectation distribution. The ablation shows that removing PEPM degrades performance, but this only demonstrates that priors matter; it does not demonstrate that the LLM agent architecture adds value beyond those priors. This missing baseline is load-bearing for the paper's headline conclusion.","major_comments":[{"comment":"The validation of the distributional claim lacks a prior-only baseline. The PEPM supplies agents with contemporaneous prior expectations from the same populations that serve as benchmarks: 2019 MSC and SPF priors for the Andre et al. vignettes, 2024 Chopra et al. priors for the information-provision experiment, and 2024 MSC priors for the 2025 MSC pre-estimation. The ablation in Figure 12 shows that removing PEPM sharply reduces distributional similarity, but this establishes only that the injected priors are important, not that the modular LLM architecture contributes beyond those priors. A baseline that directly outputs (or applies a simple statistical transformation to) the prior distribution could plausibly match Figure 5 as well as the LLM agents do, particularly for the information-provision experiment where the posterior may be close to the prior. The authors should add such baselines for all three experiments and report the incremental improvement of LLM Agents over them. Without this, the claim that the framework 'simulates' expectations rather than merely re-issuing survey priors is not supported.","section":"Sections 3.1-3.2, 5.1 (Figure 5), 6 (Figure 12)"},{"comment":"The out-of-sample MSC test covers a single period (2025) and uses 2024 MSC priors from the same survey. Since the 2024 prior distribution may be a strong forecast of the 2025 distribution, this design does not provide independent evidence for the architecture unless compared against a prior-only baseline. The authors should report, for example, the similarity between the 2024 MSC distribution (or a simple lagged/conditional version of it) and the 2025 MSC distribution, and show that the LLM agents' pre-estimation improves on that benchmark. A multi-period or multi-horizon out-of-sample evaluation would also strengthen the claim of pre-estimation capability.","section":"Section 4.3, Figure 5(d)"},{"comment":"The distributional similarity metrics—Pearson correlation and cosine similarity on histogram-based probability vectors—are weak criteria for 'highly similar' distributions. High correlation can coexist with substantial mean shifts, and the values depend on the chosen binning rule (Freedman-Diaconis or approximate sample-size bins). The authors should supplement these metrics with more stringent comparisons such as the Kolmogorov-Smirnov statistic, Wasserstein distance, or direct comparisons of means and variances between simulated and human distributions. This is important because the headline claim of distributional fidelity rests entirely on these two similarity measures.","section":"Section 5.1, Figure 5"}],"minor_comments":[{"comment":"In the concluding remarks, 'Extending the framework to stimulate the expectations of firms' should be 'simulate'; this typo also appears in the supplementary appendix headings.","section":"Section 7"},{"comment":"The description of the bootstrap confidence intervals ('obtained by bootstrap over histogram-based probability vectors') is unclear; please specify whether the bootstrap resamples respondents, agents, or histogram bins, and how the confidence intervals are constructed.","section":"Figure 5 caption"},{"comment":"The knowledge cutoffs for Qwen3-235B-A22B-Thinking-2507 and DeepSeek-R1-0528 are inferred by querying the models rather than officially disclosed; this is a reasonable approximation but should be presented with appropriate caution, as the contamination-freeness argument depends on these cutoffs.","section":"Supplementary Appendix Table A.1"},{"comment":"The random disturbance distributions for Temperature and Top-p are introduced as free parameters, but no sensitivity analysis is reported for alternative parameterizations; a brief robustness check or discussion of how sensitive the distributional similarity results are to these choices would be useful.","section":"Section 3.1, Random Disturbances"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and interesting question, and the modular-agent framework is a potentially useful methodological contribution. The main concern is the missing prior-only baseline, which is fixable within the manuscript's scope. If the authors add that baseline and show that LLM agents improve over it, the paper could be suitable for publication. I also recommend that the editor ask for a data and code availability statement, as the manuscript currently provides no explicit replication plan despite relying on several custom agentic workflows and datasets."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2505.17648. First, it is a competent, well-documented attempt to simulate macroeconomic expectations with modular LLM agents, and it does some things better than the existing LLM-expectation papers: it distinguishes households and experts, validates on three different survey designs, and includes an unusual amount of mechanism analysis (selective recall, mental models as DAGs). Second, the central claim that the modular architecture drives the distributional matches is not actually tested, because the paper never compares against a baseline that simply uses the contemporaneous prior distributions that are fed into the agents. The ablation shows removing PEPM hurts, but that only proves priors matter. It does not prove the LLM machinery adds value over the priors themselves. That gap is the difference between 'we built a useful way to inject priors into an LLM' and 'LLM agents simulate macro expectations.'\n\nWhat is genuinely new: the UNITE framework with PCM, PEPM, SMIM, PBM, KAM; the expert-agent construction using scraped LinkedIn profiles and RAG; the open-ended response analysis with agentic workflows and causal DAGs; the ablations showing different modules contribute to different dimensions. The out-of-sample MSC pre-estimation for 2025 is a nice proof of concept even if it is only one period and the 2024 priors may be a strong forecast.\n\nThe soft spots, in order: (1) the missing prior-only baseline. For the information-provision experiment, the posterior may already resemble the prior; a prior-only baseline could match Figure 5(c) almost as well. The authors address contamination in footnote 17 but do not address this. (2) No code or data released, which makes reproducibility hard given the elaborate pipeline. (3) The 'only INITIAL' baseline strips all modules, so it is not a fair prompt-engineering competitor; the prior-only baseline is the real control. (4) Semi-synthetic expert profiles and random matching of priors to profiles add noise that is not modeled.\n\nNone of these are fatal. The paper is transparent about limitations, positions itself as complementary, and the mechanism analysis is credible. But the headline needs the prior-only baseline before the architecture's added value is established.\n\nWho is it for: economists and applied ML people working on expectations formation or AI behavioral science, and anyone thinking of using LLM agents as a low-cost pre-testing tool. I would send it to referees, but with a clear request for the baseline, artifacts, and ideally a longer out-of-sample window. If the authors provide the baseline and the architecture still wins, this becomes a solid methods paper.","headline":"Solid framework paper, but the missing prior-only baseline leaves the architecture's added value unproven; worth refereeing with that condition.","tokens_in":48550,"tokens_out":3240,"would_cite":false,"duration_ms":28914,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A modular LLM agent architecture that supplies personal characteristics, priors, and external information can simulate household and expert macroeconomic expectations closely enough to match human survey distributions and human-like…","keywords":["macroeconomic expectations","LLM agents","survey experiments","expectation formation","large language models","simulation","selective recall","mental models"],"falsifier":"Run the same three experiments with agents that receive only the prior-expectation inputs, or a simple statistical transform of those priors, and compare the distributional similarity to the full-agent results; if the prior-only baseline matches human distributions as well as the full agents do, the claim that the architecture adds simulation power would be falsified.","tokens_in":47559,"feed_emoji":"🤖","tokens_out":10938,"duration_ms":78429,"temperature":0.7,"pith_summary":"The paper proposes a framework, called UNITE, for building LLM-based economic agents that simulate how households and experts form macroeconomic expectations in survey experiments. It claims that agents equipped with modules supplying personal characteristics, prior expectations, and external or professional information generate expectation distributions that closely match human survey data across three representative experiments, while also capturing qualitative patterns in open-ended reasoning that prompt-only foundation models miss. The authors report that simulated distributions are more homogeneous than human ones and therefore position the method as a complement to, not a replacement for, traditional surveys. If the claim holds, economists gain a low-cost, scalable way to pre-test survey designs and pre-estimate future expectation distributions.","feed_headline":"Modular LLM agents reproduce human macro expectations","feed_subtitle":"With priors and external info, agents match survey distributions and beat prompt-only LLMs","key_machinery":"The central object is the LLM Agent built inside a five-step pipeline the paper calls UNITE: Construction, Initialization, Simulation, Pre-estimation, and Evaluation. Each agent combines a general-purpose large language model, treated as the 'brain,' with functional modules: a Personal Characteristics Module (PCM), a Prior Expectations & Perceptions Module (PEPM), a Social Media Information Module (SMIM), and, for experts, a Professional Background Module (PBM) and a Knowledge Acquisition Module (KAM). Initialization prompts assign each agent a role, a confidence level, a task, and a rule for trading off priors against external signals, while random normal draws on the temperature and top-p sampling parameters stand in for unobserved human heterogeneity. The framework evaluates success by comparing histogram-based probability vectors of simulated and human expectations using Pearson correlation and cosine similarity, and by comparing the causal Directed Acyclic Graphs, or mental models, extracted from open-ended responses.","core_discovery":"The central discovery is that the architecture of the agent, rather than the underlying large language model alone, is what makes expectation simulation work. Across three benchmark designs—hypothetical vignette experiments in which households and experts react to stylized macroeconomic shocks, information-provision experiments on home price expectations, and an out-of-sample pre-estimation of the 2025 Michigan Survey of Consumers—the authors report that the shape similarity between simulated and human expectation distributions, measured by Pearson correlation and cosine similarity of histogram-based probability vectors, averages around 0.8 and rarely falls below 0.5. Ablations show that removing the Prior Expectations & Perceptions Module sharply degrades the distributional match, while removing personal-characteristics or social-media modules degrades the human-likeness of the thoughts and selective-recall patterns. Foundation models given only initialization prompts produce far less human-aligned distributions and mental models. The paper concludes that priors are the main driver of distributional fidelity, whereas personal, professional, and external-information modules are what establish human-like reasoning, and that the agents therefore narrow the belief gap between generative AI and humans at the aggregate level.","pith_inferences":["The most direct test the authors leave implicit is a prior-only baseline: until agents fed only the prior-expectation inputs are compared with the full architecture, some of the credit for the distributional match could belong to the survey inputs rather than to the modules.","A natural extension would be to calibrate the random disturbance parameters to the observed dispersion of human responses, rather than using fixed normal draws, which could reduce the documented over-homogeneity.","The same modular construction could be carried over to other belief objects, such as stock market expectations, policy narratives, or firm price-setting, wherever priors and external information can be sourced.","The mental-model comparisons imply a testable connection: agents that produce more diverse causal graphs should also produce less homogeneous expectation distributions, linking the reasoning dimension to the distributional dimension."],"forward_implications":["Economists could use the LLM Agents to pre-test survey questionnaires and vignettes before fielding costly human surveys.","The same architecture can pre-estimate future expectation distributions, as demonstrated by the out-of-sample 2025 exercise, without waiting for survey data to be released.","The ablation results give design guidance: invest in prior-expectation data for distributional fidelity and in personal, social, and professional modules for human-like reasoning.","Because simulated distributions are more homogeneous than human ones, the framework should be treated as a complement or pilot, not a substitute for human subjects.","Prompt-only foundation models are not enough; the modular architecture and initialization are what carry the simulation performance."],"supporting_citations":[{"why":"Supplies the hypothetical vignette design, the human household and expert benchmarks, and the response-coding scheme used to validate simulations.","marker":"Andre et al. (2022)"},{"why":"Supplies the information-provision experiment design, the homeowner and renter survey data with priors, and the mechanism categories for open-ended responses.","marker":"Chopra et al. (2025)"},{"why":"Provides empirical evidence that perceived and expected inflation are tightly linked, grounding the Prior Expectations & Perceptions Module.","marker":"Jonung (1981)"},{"why":"Supports the claim that recent perceptions and priors are crucial determinants of future expectations.","marker":"Coibion et al. (2020)"},{"why":"Establishes that household expectations are influenced by media and professional forecasts, grounding the social media and knowledge acquisition modules.","marker":"Carroll (2003)"},{"why":"Provides the mental-model causal-DAG framework used to compare reasoning structures across humans, agents, and prompt-only models.","marker":"Andre et al. (2025)"}],"fun_headline_variants":["Modular LLM agents mirror human macro expectations","Priors key to LLM agents' survey match","Agent architecture beats model size in expectation sims","LLM agents replicate survey expectations with priors","Design, not model, drives LLM agent realism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the prior beliefs fed into the agents are not themselves almost enough to reproduce the human distributions, so the modules and initialization do real work beyond passing those priors through.","fun_headline_variants_meta":{"raw":{"variants":["Modular LLM agents mirror human macro expectations","Priors key to LLM agents' survey match","Agent architecture beats model size in expectation sims","LLM agents replicate survey expectations with priors","Design, not model, drives LLM agent realism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1186,"prompt_tokens":880,"completion_tokens":306,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":233}},"tokens_in":496,"tokens_out":306,"duration_ms":2768,"temperature":1.0,"reasoning_tokens":233,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:42:10.300202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three experiments with agents that receive only the prior-expectation inputs, or a simple statistical transform of those priors, and compare the distributional similarity to the full-agent results; if the prior-only baseline matches human distributions as well as the full agents do, the claim that the architecture adds simulation power would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the mental-model causal-DAG framework used to compare reasoning structures across humans, agents, and prompt-only models."}],"review_version":1}