REVIEW 3 major objections 6 minor 65 references
Large language models can reproduce human route-choice biases and yield cumulative prospect theory parameters that fit real traveler data without being given those parameters.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 04:14 UTC pith:JHH22HEH
load-bearing objection Solid pipeline that recovers CPT parameters competitive with human lab estimates, but the risk labels in the prompts make the “emergent bias” claim weaker than advertised. the 3 major comments →
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Without being told any prospect-theory parameters, LLM agents given only demographic and risk-preference language produce route choices whose aggregate statistics are consistent with cumulative prospect theory; the parameters fitted from those choices (α ≈ 0.4, β ≈ 0.64, λ ≈ 1.43) reproduce real human choice frequencies with lower total squared error than several classic parameter sets.
What carries the argument
A standardized three-part prompting architecture (profile, planning, action) that turns natural-language traveler descriptions into structured JSON route choices, then feeds the choice frequencies into ordinary CPT residual-sum-of-squares estimation.
Load-bearing premise
The risk-preference wording already written into each agent’s profile does not itself force the CPT-like pattern that is later recovered from the choices.
What would settle it
Re-run the same gain/loss scenarios with profiles that never mention risk, stability, or volatility; if the recovered α, β, λ still match human benchmarks, the claim holds; if they collapse toward expected-utility values, the original result was prompt-driven.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether LLMs can reproduce CPT-like route-choice biases without being given explicit prospect-theoretic parameters, and whether LLM-generated choices can support scalable CPT parameter estimation. It builds 3,000 GPT-5 agents from ten demographic–behavioral profiles via a profile/planning/action prompt architecture, runs ten gain/loss two-route cases plus multi-route and same-reference-point robustness cases, fits α, β, λ (with γ fixed at 0.74) by least squares to aggregate choice shares, infers reference points as mean stated expected travel times, and compares recovered parameters against Xu et al. (2011) human route-choice data and classic CPT parameter sets. The authors report α=0.4, β=0.64, λ=1.43, good within-sample fit, ordered type-level choice patterns, and lower total squared error on the human benchmark than several alternative parameterizations.
Significance. If the central claim holds, the work would offer a practical path around the survey bottleneck for CPT-based agent-based traffic simulation: large, heterogeneous synthetic choice panels at low marginal cost, with parameters that transfer reasonably to real route-choice data. Strengths include a carefully structured experimental design (matched-mean routes, four discrete outcomes, three independent runs, structured JSON), explicit multi-type heterogeneity, individual- and environment-level checks, a same-reference-point robustness scenario, and a direct head-to-head against published human-estimated and classic CPT parameters on Xu et al. (2011). Those elements make the manuscript a useful contribution to LLM-based behavioral modeling even if the strongest “emergence without specification” framing needs to be tightened.
major comments (3)
- [§4.1, Appendix A, Tables 7–9] Central claim vs. agent construction (§4.1, Table 1, Appendix A; Tables 7–9). The abstract and introduction state that LLMs reproduce non-rational biases “without explicit specification of prospect-theoretic parameters.” The profiles, however, explicitly encode risk attitudes and decision rules (e.g., “risk aversion,” “prefers stable routes with minimal fluctuations,” “choose routes that save more time, even if they involve higher variability”). Aggregate type-level shares then track those labels almost one-to-one (risk-averse types 1–5 prefer A in gains and shift under losses; types 9–10 almost always take B). The subsequent fit of α, β, λ (Eqs. 9–10) therefore largely re-expresses prompted heterogeneity rather than demonstrating unguided CPT structure. This is load-bearing for the paper’s strongest claim. Please either (i) add an ablation with demographically matched but risk-language-
- [§5.6, Tables 10–11] External validation and what is being claimed (§5.6, Tables 10–11, Figure 6). The comparison to Xu et al. (2011) is the main non-circular anchor and is valuable: LLM-fitted (0.40, 0.64, 1.43) yields the smallest total error among the listed sets. But parameters are estimated on LLM choices, then scored on human shares; success shows that the synthetic panel produces a useful CPT parameterization for this task, not that the LLM independently recovered human cognitive parameters. Please separate clearly (a) within-LLM CPT consistency, (b) transfer of LLM-estimated parameters to human choice shares, and (c) any claim that LLM agents “are” CPT decision makers. Also report uncertainty (bootstrap or multi-seed intervals) on α, β, λ and on the error differences in Table 10, and note that γ is fixed a priori at 0.74 (Wu & Gonzalez 1996) rather than jointly estimated.
- [§4.3 Eq. (11), Table 6] Reference-point construction and case-11 evaluation (Eq. 11; §5.3–5.5, Table 6). The reference point is defined as the mean of the same agents’ stated expected travel times. That is a reasonable behavioral proxy, but it couples the evaluation of CPT values in case 11 to quantities generated by the same model family under the same prompts. Please discuss sensitivity to alternative reference-point rules used in the literature (free-flow time, budgeted time, fixed design reference as in Scenarios 1–2), and report case-11 CPT predictions under those alternatives so that Table 6 is not solely conditioned on Eq. (11).
minor comments (6)
- [Abstract, §1] Abstract vs. methods wording: “without explicit specification of prospect-theoretic parameters” is true for numeric CPT parameters but easy to misread as “without risk-preference guidance.” Soften or qualify in abstract and §1 once the ablation/reframing decision is made.
- [§5.2, Algorithm 1] Model identity: the text uses “OpenAI gpt-5” / “GPT-5.” Confirm the exact public model identifier and API settings (temperature=1 is stated; also seed/top_p if available) for reproducibility.
- [Figure 3, Appendix B] Figure 3 / Appendix B overplotting: the dense band and the note about agents 2300–2399 are helpful; consider jitter or small-multiple panels earlier so readers do not need Appendix B.12 to resolve the artifact.
- [§4.3 Eqs. (9)–(10)] Notation: P(A_J) vs P(A_j) in Eq. (9) appears inconsistent (capital J). Unify indices in Eqs. (9)–(10).
- [§2.2] Related work: briefly situate against other LLM decision-bias / economic-preference studies beyond generative-agent sandbox papers, so the route-choice CPT contribution is clearer relative to concurrent LLM behavioral work.
- [§2.2, Appendix A] Typographical: “large-scale parameterization and pretraining on massive datas” (§2.2); “very importance” in Appendix A composed agent; duplicate “1 Traffic Environment…” block in Figure 1 caption area of the text.
Circularity Check
Prompted risk labels and decision rules force the CPT-like patterns later fitted and presented as emergent LLM biases without explicit CPT parameters.
specific steps
-
other
[§4.1 (Construction of LLM-based human agents) + Appendix A + Abstract claim]
"The decision rules specify how the agent should weigh different route attributes and how risk attitudes influence choice preferences. For example, risk-averse agents are expected to prioritize routes with lower time variability, whereas risk-seeking agents tend to prefer routes with shorter travel times but higher variability. ... [Abstract:] whether large language models (LLMs) can reproduce human behavioral biases in choice-making without explicit specification of prospect-theoretic parameters."
Agents are defined with explicit risk-preference language and decision rules that prescribe precisely the stability-seeking / risk-seeking asymmetries CPT models. Observing that the generated choices later match CPT (and can be fitted by α, β, λ) is therefore largely by construction of the prompts, not an independent emergence from unguided LLM behavior. The “without explicit specification” claim is thereby undermined at the generation step.
-
fitted input called prediction
[§5.3 (Experiment results) eqs. (9)–(10) and Table 6]
"Using the statistical results of cases 1 to 10 in table 5, the parameters α, β and λ are estimated by fitting equations (9) and (10), yielding the following estimates: α=0.4, β=0.64 and λ=1.43. ... Based on these estimated parameters, the cumulative prospect values and choice probabilities for routes A, B, and C in case 11 are computed, as reported in Table 6."
CPT parameters are fitted by least-squares to the very LLM agents’ aggregate choices in cases 1–10 (already shaped by the risk-labeled prompts). Those same parameters are then used to “predict” the same agents’ choices on the closely related three-route case 11. The reported close match is statistically forced by the shared generative process and functional-form fit rather than an out-of-sample discovery of CPT structure.
full rationale
The paper's central claim is that LLMs reproduce non-rational CPT-consistent route biases without explicit prospect-theoretic parameters, then recover α/β/λ from the resulting choices and show those parameters also fit external human data better than classics. The load-bearing generation step, however, constructs agents via prompts that already encode the risk attitudes, stability preferences, and gain/loss trade-offs that CPT is designed to capture (profile table, §4.1 decision rules, Appendix A). Aggregate choices then simply re-express those labels (Tables 7–9), so the least-squares fit of CPT (eqs. 9–10) and the subsequent “prediction” of the same agents on case 11 largely restate the prompted heterogeneity rather than independently discovering CPT structure. The external comparison to Xu et al. (2011) supplies a non-circular anchor and keeps the circularity partial rather than total; no self-citation chain or uniqueness theorem is load-bearing. An ablation that strips risk-language while retaining demographics would be required to make the “without explicit specification” claim secure.
Axiom & Free-Parameter Ledger
free parameters (5)
- α (gain sensitivity) =
0.4
- β (loss sensitivity) =
0.64
- λ (loss-aversion coefficient) =
1.43
- γ (probability weighting) =
0.74
- LLM temperature =
1
axioms (4)
- domain assumption Value function and cumulative weighting functions of CPT take the standard power and Prelec-like forms given in Eqs. (4)–(8).
- ad hoc to paper Natural-language risk-preference labels in the agent profiles induce the same qualitative risk attitudes that CPT is designed to measure.
- ad hoc to paper Reference point for each case equals the arithmetic mean of the agents’ stated expected travel times (Eq. 11).
- domain assumption Choice probabilities follow a binary logit on the difference of cumulative prospect values.
invented entities (2)
-
Standardized three-component prompting architecture (profile / planning / action)
no independent evidence
-
Ten demographic-behavioral agent types (Table 1)
no independent evidence
read the original abstract
Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely on surveys and controlled experiments to calibrate CPT parameters, yet these methods are difficult to generalize and often fail to capture the full diversity of human decision-making. To address this challenge, this paper investigates whether large language models (LLMs) can reproduce human behavioral biases in choice-making without explicit specification of prospect-theoretic parameters. Using route choice as a representative scenario, we design a behavioral evaluation framework and systematically compare LLM-generated decisions with established human behavioral patterns predicted by CPT. Experimental results demonstrate that LLMs are capable of reproducing non-rational human choice biases and can exhibit decision behaviors consistent with prospect-theoretic effects under uncertainty. These findings suggest that generative AI models may provide a scalable alternative for modeling human decision processes and offer a promising foundation for next-generation large-scale agent-based simulation and AI-driven behavioral research.
Figures
Reference graph
Works this paper leans on
-
[1]
Management Science , volume=
Hierarchical maximum likelihood parameter estimation for cumulative prospect theory: Improving the reliability of individual risk parameter estimates , author=. Management Science , volume=. 2018 , publisher=
2018
-
[2]
Nature Human Behaviour , volume=
A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures , author=. Nature Human Behaviour , volume=. 2025 , publisher=
2025
-
[3]
Nature Human Behaviour , volume=
Modelling dataset bias in machine-learned theories of economic decision-making , author=. Nature Human Behaviour , volume=. 2024 , publisher=
2024
-
[4]
The Wiley Blackwell handbook of judgment and decision making , volume=
Decision under risk: From the field to the laboratory and back , author=. The Wiley Blackwell handbook of judgment and decision making , volume=. 2015 , publisher=
2015
-
[5]
Journal of economic perspectives , volume=
Thirty years of prospect theory in economics: A review and assessment , author=. Journal of economic perspectives , volume=. 2013 , publisher=
2013
-
[6]
Advances in neural information processing systems , volume=
Attention is all you need , author=. Advances in neural information processing systems , volume=
-
[7]
Advances in neural information processing systems , volume=
Language models are few-shot learners , author=. Advances in neural information processing systems , volume=
-
[8]
arXiv preprint arXiv:2211.15661 , year=
What learning algorithm is in-context learning? investigations with linear models , author=. arXiv preprint arXiv:2211.15661 , year=
-
[9]
Proceedings of the 36th annual acm symposium on user interface software and technology , pages=
Generative agents: Interactive simulacra of human behavior , author=. Proceedings of the 36th annual acm symposium on user interface software and technology , pages=
-
[10]
Advances in neural information processing systems , volume=
Tree of thoughts: Deliberate problem solving with large language models , author=. Advances in neural information processing systems , volume=
-
[11]
Frontiers of Computer Science , volume=
A survey on large language model based autonomous agents , author=. Frontiers of Computer Science , volume=. 2024 , publisher=
2024
-
[12]
LLM-powered Autonomous Agents
Weng, Lilian. LLM-powered Autonomous Agents. lilianweng.github.io. 2023
2023
-
[13]
Proceedings of the 2024 conference on empirical methods in natural language processing , pages=
A survey on in-context learning , author=. Proceedings of the 2024 conference on empirical methods in natural language processing , pages=
2024
-
[14]
ACM computing surveys , volume=
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing , author=. ACM computing surveys , volume=. 2023 , publisher=
2023
-
[15]
Journal of Risk and uncertainty , volume=
Advances in prospect theory: Cumulative representation of uncertainty , author=. Journal of Risk and uncertainty , volume=. 1992 , publisher=
1992
-
[16]
Management science , volume=
Curvature of the probability weighting function , author=. Management science , volume=. 1996 , publisher=
1996
-
[17]
Transportation Research Part B: Methodological , volume=
Boundedly rational route choice behavior: A review of models and methodologies , author=. Transportation Research Part B: Methodological , volume=. 2016 , publisher=
2016
-
[18]
Transportation Research Part B: Methodological , volume=
Day-to-day stationary link flow pattern , author=. Transportation Research Part B: Methodological , volume=. 2009 , publisher=
2009
-
[19]
Transportation Research Part A: Policy and Practice , volume=
Investigating day-to-day route choices based on multi-scenario laboratory experiments, Part I: Route-dependent attraction and its modeling , author=. Transportation Research Part A: Policy and Practice , volume=. 2023 , publisher=
2023
-
[20]
some theoretical aspects of road traffic research
Road paper. some theoretical aspects of road traffic research. , author=. Proceedings of the institution of civil engineers , volume=. 1952 , publisher=
1952
-
[21]
Transportation Research Part B: Methodological , volume=
On the local and global stability of a travel route choice adjustment process , author=. Transportation Research Part B: Methodological , volume=. 1996 , publisher=
1996
-
[22]
Transportation science , volume=
The stability of a dynamic model of traffic assignment—an application of a method of Lyapunov , author=. Transportation science , volume=. 1984 , publisher=
1984
-
[23]
Transportation Science , volume=
Dynamic user equilibrium departure time and route choice on idealized traffic arterials , author=. Transportation Science , volume=. 1984 , publisher=
1984
-
[24]
Transportation Research Part C: Emerging Technologies , volume=
Bounded-rationality based day-to-day evolution model for travel behavior analysis of urban railway network , author=. Transportation Research Part C: Emerging Technologies , volume=. 2013 , publisher=
2013
-
[25]
European journal of operational research , volume=
Handling uncertainty in route choice models: From probabilistic to possibilistic approaches , author=. European journal of operational research , volume=. 2006 , publisher=
2006
-
[26]
Econometrica , volume=
Prospect theory: An analysis of decision under risk , author=. Econometrica , volume=
-
[27]
Transportation Research Part A: Policy and Practice , volume=
Charging mode choice of electric micromobility users under uncertainty , author=. Transportation Research Part A: Policy and Practice , volume=. 2026 , publisher=
2026
-
[28]
Transportation research part A: policy and practice , volume=
Commuter departure time choice behavior under congestion charge: Analysis based on cumulative prospect theory , author=. Transportation research part A: policy and practice , volume=. 2023 , publisher=
2023
-
[29]
Transportation Research Part A: Policy and Practice , volume=
Examining the temporary use behavior of autonomous vehicles under uncertainty: A stated preference analysis , author=. Transportation Research Part A: Policy and Practice , volume=. 2025 , publisher=
2025
-
[30]
Transportation Research Part A: Policy and Practice , volume=
An application of cumulative prospect theory to freeway drivers’ route choice behaviours , author=. Transportation Research Part A: Policy and Practice , volume=. 2013 , publisher=
2013
-
[31]
Transportation Research Part C: Emerging Technologies , volume=
Modeling effects of travel time reliability on mode choice using cumulative prospect theory , author=. Transportation Research Part C: Emerging Technologies , volume=. 2019 , publisher=
2019
-
[32]
Transportation Research Part C: Emerging Technologies , volume=
A cumulative prospect theory approach to commuters’ day-to-day route-choice modeling with friends’ travel information , author=. Transportation Research Part C: Emerging Technologies , volume=. 2018 , publisher=
2018
-
[33]
Transportation Science , volume=
Wardrop equilibrium can be boundedly rational: A new behavioral theory of route choice , author=. Transportation Science , volume=. 2024 , publisher=
2024
-
[34]
Transportation Research Part B: Methodological , volume=
A closed-form bounded route choice model accounting for heteroscedasticity, overlap, and choice set formation , author=. Transportation Research Part B: Methodological , volume=. 2025 , publisher=
2025
-
[35]
Transportation Research Part B: Methodological , volume=
A generalized rationally inattentive route choice model with non-uniform marginal information costs , author=. Transportation Research Part B: Methodological , volume=. 2024 , publisher=
2024
-
[36]
Transportation Research Part B: Methodological , volume=
Braess paradox under the boundedly rational user equilibria , author=. Transportation Research Part B: Methodological , volume=. 2014 , publisher=
2014
-
[37]
Journal of economic literature , volume=
Why bounded rationality? , author=. Journal of economic literature , volume=. 1996 , publisher=
1996
-
[38]
Transportation Research Part A: Policy and Practice , volume=
Modeling dynamic travel mode choices using cumulative prospect theory , author=. Transportation Research Part A: Policy and Practice , volume=. 2024 , publisher=
2024
-
[39]
Transportation Research Part C: Emerging Technologies , volume=
A decision-making rule for modeling travelers’ route choice behavior based on cumulative prospect theory , author=. Transportation Research Part C: Emerging Technologies , volume=. 2011 , publisher=
2011
-
[40]
Transportation Research Part C: Emerging Technologies , volume=
Development of an enhanced route choice model based on cumulative prospect theory , author=. Transportation Research Part C: Emerging Technologies , volume=. 2014 , publisher=
2014
-
[41]
Nature , volume=
Role play with large language models , author=. Nature , volume=. 2023 , publisher=
2023
-
[42]
arXiv preprint arXiv:2412.03563 , year=
From individual to society: A survey on social simulation driven by large language model-based agents , author=. arXiv preprint arXiv:2412.03563 , year=
-
[43]
Transportation Research Part C: Emerging Technologies , volume=
Agentic Large Language Models for day-to-day route choices , author=. Transportation Research Part C: Emerging Technologies , volume=. 2025 , publisher=
2025
-
[44]
Transportation Research Part C: Emerging Technologies , volume=
An explainable end-to-end autonomous driving framework based on large language model and vision modality fusion: design and application of DriveLLM-V , author=. Transportation Research Part C: Emerging Technologies , volume=. 2025 , publisher=
2025
-
[45]
Transportation Research Part C: Emerging Technologies , volume=
Traffic-IT: Enhancing traffic scene understanding for multimodal large language models , author=. Transportation Research Part C: Emerging Technologies , volume=. 2025 , publisher=
2025
-
[46]
CHI Conference on Human Factors in Computing Systems Extended Abstracts , pages=
Promptchainer: Chaining large language model prompts through visual programming , author=. CHI Conference on Human Factors in Computing Systems Extended Abstracts , pages=
-
[47]
Proceedings of the 2022 CHI conference on human factors in computing systems , pages=
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts , author=. Proceedings of the 2022 CHI conference on human factors in computing systems , pages=
2022
-
[48]
arXiv preprint arXiv:2311.17227 , year=
War and peace (waragent): Large language model-based multi-agent simulation of world wars , author=. arXiv preprint arXiv:2311.17227 , year=
-
[49]
, author=
AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors. , author=. ICLR , year=
-
[50]
Transactions on Machine Learning Research , year=
Cognitive architectures for language agents , author=. Transactions on Machine Learning Research , year=
-
[51]
Advances in Neural Information Processing Systems , volume=
Toolformer: Language models can teach themselves to use tools , author=. Advances in Neural Information Processing Systems , volume=
-
[52]
arXiv preprint arXiv:2308.05960 , year=
Bolaa: Benchmarking and orchestrating llm-augmented autonomous agents , author=. arXiv preprint arXiv:2308.05960 , year=
-
[53]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Chatdev: Communicative agents for software development , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[54]
arXiv preprint arXiv:2309.04658 , year=
Exploring large language models for communication games: An empirical study on werewolf , author=. arXiv preprint arXiv:2309.04658 , year=
-
[55]
arXiv preprint arXiv:2301.02111 , year=
Neural codec language models are zero-shot text to speech synthesizers , author=. arXiv preprint arXiv:2301.02111 , year=
-
[56]
Computer Vision and Image Understanding , volume=
Deep learning for deepfakes creation and detection: A survey , author=. Computer Vision and Image Understanding , volume=. 2022 , publisher=
2022
-
[57]
arXiv preprint arXiv:2310.11667 , year=
Sotopia: Interactive evaluation for social intelligence in language agents , author=. arXiv preprint arXiv:2310.11667 , year=
-
[58]
Transportation Research Part A: Policy and Practice , volume=
Utilizing large language models to simulate parking search , author=. Transportation Research Part A: Policy and Practice , volume=. 2025 , publisher=
2025
-
[59]
Transportation Research Part E: Logistics and Transportation , volume=
LLM4STP: A large language model-driven multi-feature fusion method for ship trajectory prediction , author=. Transportation Research Part E: Logistics and Transportation , volume=. 2026 , publisher=
2026
-
[60]
Agentsense: Benchmarking social intelligence of language agents through interactive scenarios , author=. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2025
-
[61]
International conference on machine learning , pages=
Using large language models to simulate multiple humans and replicate human subject studies , author=. International conference on machine learning , pages=. 2023 , organization=
2023
-
[62]
American economic review , volume=
Risk aversion and incentive effects , author=. American economic review , volume=. 2002 , publisher=
2002
-
[63]
Advances in Neural Information Processing Systems , volume=
Large language models as urban residents: An llm agent framework for personal mobility generation , author=. Advances in Neural Information Processing Systems , volume=
-
[64]
Travel Behaviour and Society , volume=
Valuing time in silicon: Can large language models replicate human value of travel time , author=. Travel Behaviour and Society , volume=. 2026 , publisher=
2026
-
[65]
Communications in Transportation Research , volume =
Qi, Hang and Jia, Ning and Qu, Xiaobo and He, Zhengbing , title =. Communications in Transportation Research , volume =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.