Pith. sign in

REVIEW 5 major objections 9 minor 44 references

Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study

T0 review · 5 major / 9 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper reports that LLM-simulated survey respondents, built from demographic profiles, choose more risk-averse lottery options than the human respondents they mimic, and that prompting in Chinese widens the divergence.

desk verdict Useful benchmark result, but the CRRA estimation is under-specified and the abstract overstates the findings. read the letter →

arxiv 2506.23107 v1 pith:UEHFM4V5 submitted 2025-06-29 cs.AI

classification cs.AI
keywords largelanguagemodelsriskpreferencesCRRAlotterychoicerole-playingagentscross-culturalstudystatedpreferencesurveyprompt
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether large language models can stand in for human respondents in risky-choice surveys. Across four cities, simulated respondents built from age, gender, education, income, and city cues chose more cautiously in lottery games than the real participants, and their estimated risk-aversion coefficients were generally higher. ChatGPT o1-mini tracked human choices more closely than ChatGPT 4o did. When the prompt language for Hong Kong and Nanjing was switched from English to Chinese, the distance between simulated and real risk attitudes grew, even though those respondents natively speak Chinese. If these results hold, LLM-based stand-ins for human survey respondents will need calibration before they can be trusted in policy or market research about risk.

What carries the argument

The central object is the CRRA risk-preference coefficient $r$, defined through the utility function $U(x)=x^{1-r}/(1-r)$, with $r>0$ meaning risk aversion, $r=0$ risk neutrality, and $r<0$ risk seeking. The paper adopts the linear specification $r = r_0 + \alpha_1 X_1$ from reference [44] to connect socio-demographic variables $X_1$ to risk preferences, and it builds role-playing agents whose prompts combine an individual profile (age, gender, education, income, city) with chain-of-thought planning instructions and a closed-form action rule that outputs only the chosen lottery option. Each task is simulated three times and the majority choice is kept as the simulated response.

What would settle it

Re-estimate the risk coefficients from the same human and LLM choices with an explicit stochastic choice model, such as a logistic choice rule over CRRA utilities with an error term, and compare the LLM-human gap; if the systematic conservative bias disappears or reverses under that specification, the headline result is an artifact of the unstated mapping from choices to risk scores.

Watch

Extended reading notes

Core claim

Using the Constant Relative Risk Aversion (CRRA) framework, the paper estimates risk-preference coefficients $r$ from real and LLM-simulated lottery choices. It finds that ChatGPT 4o's simulated respondents are systematically more risk-averse than the human respondents in all four city samples, and that ChatGPT o1-mini is generally more risk-averse than humans as well, though its estimates are closer to the observed human values, with Hong Kong as the one reported exception. The paper further finds that for the Hong Kong and Nanjing samples, presenting the same tasks in Chinese rather than English moves the simulated risk attitudes further from the human estimates, a bias amplification the authors attribute to the models' predominantly English-language training.

Load-bearing premise

The argument assumes that a single CRRA risk-score equation taken from an earlier transportation study converts each respondent's lottery answers into a comparable risk-aversion number across four differently designed lotteries, but the paper does not state the choice model that performs this conversion.

Editorial extensions

If this is right

  • Uncalibrated LLM stand-ins will make populations look more risk-averse than they are, pushing policy recommendations toward overly conservative measures.
  • Model choice matters: ChatGPT o1-mini is a closer proxy for human risk preferences than ChatGPT 4o in these tasks, so conclusions drawn from one LLM may not transfer to another.
  • Prompt language is an active variable: even when respondents natively speak Chinese, Chinese prompts can degrade fidelity, so localization alone does not improve simulation accuracy.
  • Demographically narrow samples, such as the Dhaka taxi-driver data, can erase measurable differences between models, suggesting that LLMs fall back on generic rather than person-specific risk priors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: vary the stakes or the gain-loss framing while holding the demographic profile fixed; if the conservative bias grows with stakes, the mechanism is closer to loss aversion or safety-oriented calibration than to CRRA curvature alone.
  • An inference not in the paper: because the mapping from lottery answers to individual $r$ estimates is never spelled out, an explicit stochastic choice model, such as a logit over CRRA utilities, could change the size or even the sign of the reported LLM-human gap.
  • A testable extension: run the same English-versus-Chinese comparison with models trained primarily on Chinese corpora; the paper's training-bias account predicts the gap would shrink or reverse.
  • An inference not in the paper: the four lotteries differ in payoff magnitudes and in risky-safe versus risky-risky framing, so a matched-design replication is needed before the cross-city differences in human risk aversion can be read as cultural rather than artifactual.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 9 minor

Summary. This paper asks whether ChatGPT 4o and ChatGPT o1-mini can mimic human risky choices by simulating respondents in four transport stated-preference surveys (Sydney, Hong Kong, Dhaka, Nanjing). The authors feed each LLM a profile built from age, gender, education, and income, have the model answer repeated binary lottery tasks in English (and, for Hong Kong and Nanjing, also in Chinese), take a majority vote over three runs, and estimate CRRA risk-aversion parameters r using the specification r = r0 + α1X1 from reference [44]. They conclude that LLM-simulated respondents are generally more risk-averse than humans, that o1-mini is closer to human estimates than 4o, and that Chinese prompts produce larger deviations. The paper positions this as evidence of an intrinsic conservative bias and warns against using LLMs for risk-sensitive policymaking without calibration.

Significance. If the estimation were fully specified and reproducible, this would be a useful cross-cultural, cross-lingual benchmark for LLM-based behavioral simulation: four diverse datasets, two models, controlled demographic conditioning, and a formal utility framework. The paper also makes concrete, falsifiable claims about systematic overestimation and language-dependent bias, and it draws policy implications that are directly testable. However, the manuscript does not provide the choice model, likelihood, standard errors, or code/data needed to verify the point estimates, and one headline claim is contradicted by the paper's own Table 7. The contribution is potentially valuable but not yet verifiable in this form.

major comments (5)
  1. [§4.2, Eq. (1)] The paper never specifies the probabilistic choice model that maps each binary lottery answer to an individual r estimate. Eq. (1) defines CRRA utility and the text invokes r = r0 + α1X1 from [44], but there is no likelihood, no statement of whether r is estimated per person or from all choices pooled, and no parameter uncertainty. Tables 6–9 therefore report means of an unstated estimator, and the paired t-tests in Section 5.1 are presented only through significance stars, with no standard errors, p-values, or test statistics. The central quantitative claims that LLMs are more risk-averse than humans and that o1-mini is closer cannot be assessed without this identification step. Please state the choice model (e.g., a logit over CRRA expected utilities with a specific noise term), the estimation method, and report standard errors or confidence intervals for each mean.
  2. [Abstract; §6.1] The headline claim that “both ChatGPT 4o and ChatGPT o1-mini systematically exhibited higher levels of risk aversion compared to actual human respondents” is contradicted by Table 7. For Hong Kong, the o1-mini mean is 0.509 versus the real mean of 0.765, so the model is less risk-averse than the human sample. The text in Section 5.1 carefully avoids saying “higher” for that cell, but the abstract and Section 6.1 overstate the result. Please qualify the claim to the datasets/cells where the direction holds, or add a statistical test that supports a general “systematic” direction.
  3. [§3.2, Table 4] The Dhaka lottery description conflicts with Table 4. The text states that in the left lottery the lower payoff 6 occurs with probability p and the higher payoff 8 with probability 1−p, and symmetrically for the right lottery (1 with p and 20 with 1−p). However, the table's expected values (e.g., 6.2 = 0.1×8 + 0.9×6 and 2.9 = 0.1×20 + 0.9×1) show that p is instead the probability of the first-listed payoff, i.e., the higher payoff in both lotteries. Because every Dhaka CRRA estimate, and hence the Dhaka rows of Tables 6–8 and Figure 1(c), depends on the correct probability mapping, this discrepancy must be resolved and corrected.
  4. [§5.2, Table 9] Table 9 does not test the English-versus-Chinese difference that Section 5.2 claims. The reported significance stars are for each simulated mean relative to the real mean (the column header “Ha: mean(diff) ≠ 0”), not for the difference between the English and Chinese simulated means. To support the statement that switching to Chinese prompts consistently amplifies the bias, the authors need a paired test of English versus Chinese simulated risk attitudes, with standard errors, and ideally a test on absolute deviations from the real values. As it stands, the linguistic effect is not statistically established.
  5. [§3.2, Tables 2–5] The four lottery designs are not directly comparable for the cross-city claim. Sydney and Hong Kong pair a risky option against a fixed payoff; Dhaka and Nanjing pair two risky options; and the number of tasks (9 versus 10), payoff magnitudes, and probability grids differ. Under CRRA expected utility, nominal r comparisons are only meaningful if the same choice rule is assumed across designs. The manuscript should include a robustness analysis using a single structural model, or an alternative model, to show that the cross-city and cross-language comparisons are not artifacts of design differences, option-order effects, or probability-weighting deviations in LLM choices.
minor comments (9)
  1. [§4.1] The heading “Planing” is a typo and should read “Planning”.
  2. [Table 4] The header row “Prob. payoff 1 2” is unclear; define the columns (e.g., Prob(payoff 1), Prob(payoff 2)) and label which payoff is lower and which is higher, consistent with the text.
  3. [§4.1] The sample prompt tells the agent that there will be “nine consecutive lottery questions,” but the Dhaka and Nanjing surveys have ten tasks; the prompt should be dataset-specific.
  4. [§4.1] The prompt text “a age-year-old” should be “an age-year-old”.
  5. [§4.2, Eq. (1)] Eq. (1) is undefined at r = 1; adding the standard limiting form would avoid a gap for estimates near risk neutrality.
  6. [Tables 6–9] The significance-star notation is used without t-statistics or p-values; even after the estimation is clarified, the tables should include measures of uncertainty.
  7. [Figure 1] The kernel density plots should state the bandwidth and the sample sizes; with n = 64 in Sydney, density estimates are noisy and may be misleading.
  8. [Data availability] For a simulation study with bespoke prompts and a proprietary LLM, releasing at least the prompts and the estimation code would be necessary for independent verification; the current statement limits access to aggregated data only.
  9. [§3.2] The text says “the primary focus should be on the differences in expected values between the left and right options,” but for CRRA estimation the full payoff distributions, not just EV differences, are needed; the text should not suggest otherwise.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: risk attitudes are refit from separately simulated LLM choices, not defined by the comparison.

full rationale

The paper's derivation chain is: (i) take human lottery choices from four prior surveys; (ii) prompt LLMs with demographic profiles and the same lottery tasks; (iii) estimate CRRA risk coefficients separately from human choices and from LLM-simulated choices; (iv) compare the resulting mean r values across cities, models, and prompt languages. None of these steps defines the reported difference in terms of the inputs by construction. The CRRA utility function U(x) = x^(1-r)/(1-r) and the linear specification r = r0 + α1X1 are standard modeling assumptions cited from [44], but the paper does not use [44]'s fitted coefficients to generate Tables 6-9; the r values are presented as estimates from the observed and simulated choices. The finding that LLM-simulated choices imply higher r than human choices is an empirical outcome of the LLMs' outputs, not an algebraic consequence of the prompts or the estimation equation. Likewise, the Chinese-vs-English difference is a measured comparison of two sets of simulated choices, not a pre-imposed relation. The self-citations to [31]-[34] supply the underlying survey datasets, which are external empirical inputs, and [44] supplies a model form; these are not cited as proof of the paper's central claim. The manuscript does under-specify the probabilistic choice/estimation model and the comparability of r across the four lottery designs, but that is an identification and transparency limitation, not circular reasoning. No equation is shown to be equivalent to its own output, and no fitted parameter is renamed as a prediction. Therefore no circular step is exhibited.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper's quantitative conclusions rest on a fitted CRRA model whose estimation procedure is not fully specified in the text. r0 and α1 are fitted to the survey choices; the choice model itself is an assumption about how people and LLMs resolve lotteries. No new physical or conceptual entities are introduced.

free parameters (2)
  • r0 (constant CRRA risk parameter) = not reported in paper
    Intercept in r = r0 + α1 X1 fitted to each dataset's choices; all reported risk attitudes (Tables 6 to 9) are outputs of this fit.
  • α1 (coefficient vector for age, gender, education, income) = not reported in paper
    Socio-demographic coefficients in the same equation; their values determine individual-level risk estimates.
assumptions (4)
  • domain assumption Choices are generated by CRRA utility U(x) = x^(1-r)/(1-r) with a single risk parameter r
    Invoked in Section 4.2; if subjects or LLMs use probability weighting or non-expected utility, the estimated r conflates risk attitude with other biases.
  • domain assumption Risk parameter shifts linearly with demographics: r = r0 + α1 X1
    Equation in Section 4.2 adopted from reference [44]; linearity and variable set are assumed, not derived.
  • domain assumption CRRA estimates from four different lottery designs are comparable within and across cities
    Assumed when pooling Tables 6 to 9 and Figure 1; different payoff magnitudes, task counts, and both-risky versus risky-safe frames may affect elicited r.
  • domain assumption LLM choices can be treated as realizations of the same CRRA decision process as human choices
    Core to interpreting simulated choices as risk attitudes; the paper provides no validity check that LLM choice patterns satisfy CRRA or transitivity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study." pith.science (2026). https://pith.science/paper/UEHFM4V5

@misc{pith2026250623107,
  author       = {Pith},
  title        = {Pith review of: Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UEHFM4V5}},
  note         = {Machine review of arXiv:2506.23107}
}
read the original abstract

Large language models (LLMs) have made significant strides, extending their applications to dialogue systems, automated content creation, and domain-specific advisory tasks. However, as their use grows, concerns have emerged regarding their reliability in simulating complex decision-making behavior, such as risky decision-making, where a single choice can lead to multiple outcomes. This study investigates the ability of LLMs to simulate risky decision-making scenarios. We compare model-generated decisions with actual human responses in a series of lottery-based tasks, using transportation stated preference survey data from participants in Sydney, Dhaka, Hong Kong, and Nanjing. Demographic inputs were provided to two LLMs -- ChatGPT 4o and ChatGPT o1-mini -- which were tasked with predicting individual choices. Risk preferences were analyzed using the Constant Relative Risk Aversion (CRRA) framework. Results show that both models exhibit more risk-averse behavior than human participants, with o1-mini aligning more closely with observed human decisions. Further analysis of multilingual data from Nanjing and Hong Kong indicates that model predictions in Chinese deviate more from actual responses compared to English, suggesting that prompt language may influence simulation performance. These findings highlight both the promise and the current limitations of LLMs in replicating human-like risk behavior, particularly in linguistic and cultural settings.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

44 extracted references · 30 canonical work pages

  1. [44]

    Travel behaviour and society 33, 100604 (2023) 20

    Liu, J., Wu, C., Jian, S., Dixit, V., Rashidi, T.H.: Understanding the impact of occasional activities on travelers’ preferences for mobility-as-a-service: A stated preference study. Travel behaviour and society 33, 100604 (2023) 20

  2. [1]

    Transportation Research Part E: Logistics and Transportation Review 197, 104075 (2025)

    Nie, T., He, J., Mei, Y., Qin, G., Li, G., Sun, J., Ma, W.: Joint estimation and prediction of city-wide delivery demand: A large language model empowered graph-based learning approach. Transportation Research Part E: Logistics and Transportation Review 197, 104075 (2025)

  3. [2]

    arXiv preprint arXiv:2402.18013 (2024)

    Yi, Z., Ouyang, J., Liu, Y., Liao, T., Xu, Z., Shen, Y.: A survey on recent advances in llm-based multi-turn dialogue systems. arXiv preprint arXiv:2402.18013 (2024)

  4. [3]

    Massachusetts Medical Society (2025) 16

    Ohde, J.W., Rost, L.M., Overgaard, J.D.: The Burden of Reviewing LLM- Generated Content. Massachusetts Medical Society (2025) 16

  5. [4]

    In: The 2024 ACM Conference on Fairness, Accountability, and Transparency, pp

    Cheong, I., Xia, K., Feng, K.K., Chen, Q.Z., Zhang, A.X.: (a) i am not a lawyer, but...: Engaging legal experts towards responsible llm policies for legal advice. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 2454–2469 (2024)

  6. [5]

    Challagundla, B.C., Ramanujam.B, G.: Financial advisory llm model for mod- ernizing financial services and innovative solutions for financial literacy in india (2024)

  7. [6]

    : The effect of using a large language model to respond to patient messages

    Chen, S., Guevara, M., Moningi, S., Hoebers, F., Elhalawani, H., Kann, B.H., Chipidza, F.E., Leeman, J., Aerts, H.J., Miller, T., et al. : The effect of using a large language model to respond to patient messages. The Lancet Digital Health 6(6), 379–381 (2024)

  8. [7]

    Evidence and Insights on Mode Choice (August 26, 2024) (2024)

    Liu, T., Li, M., Yin, Y.: Can large language models capture human travel behav- ior? evidence and insights on mode choice. Evidence and Insights on Mode Choice (August 26, 2024) (2024)

Show all 44 references
  1. [8]

    Travel Behaviour and Society 40, 101039 (2025)

    Xu, Z., Sengar, N., Chen, T., Chung, H., Oviedo-Trespalacios, O.: Where is moral- ity on wheels? decoding large language model (llm)-driven decision in the ethical dilemmas of autonomous vehicles. Travel Behaviour and Society 40, 101039 (2025)

  2. [9]

    Transactions of the Association for Computational Linguistics 12, 1011–1026 (2024)

    Tjuatja, L., Chen, V., Wu, T., Talwalkwar, A., Neubig, G.: Do llms exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics 12, 1011–1026 (2024)

  3. [10]

    Marketing Science 43(2), 254–266 (2024)

    Li, P., Castelo, N., Katona, Z., Sarvary, M.: Frontiers: Determining the validity of large language models for automated perceptual analysis. Marketing Science 43(2), 254–266 (2024)

  4. [11]

    Goli, A., Singh, A.: Frontiers: Can large language models capture human preferences? Marketing Science (2024)

  5. [12]

    In: Handbook of the Fundamentals of Financial Decision Making: Part I, pp

    Kahneman, D., Tversky, A.: Prospect theory: An analysis of decision under risk. In: Handbook of the Fundamentals of Financial Decision Making: Part I, pp. 99–127. World Scientific, ??? (2013)

  6. [13]

    Princeton university press, ??? (2005)

    Eeckhoudt, L., Gollier, C., Schlesinger, H.: Economic and Financial Decisions Under Risk. Princeton university press, ??? (2005)

  7. [14]

    offloads

    Engelmann, J.B., Capra, C.M., Noussair, C., Berns, G.S.: Expert financial advice neurobiologically “offloads” financial decision-making under risk. PLoS one 4(3), 4957 (2009)

  8. [15]

    Lafont, C.: Deliberation, participation, and democratic legitimacy: Should delib- erative mini-publics shape public policy? Journal of political philosophy 23(1), 17 40–63 (2015)

  9. [16]

    American Economic Review 94(2), 33–40 (2004)

    Greenspan, A.: Risk and uncertainty in monetary policy. American Economic Review 94(2), 33–40 (2004)

  10. [17]

    Cognitive psychology 43(1), 1–22 (2001)

    Boroditsky, L.: Does language shape thought?: Mandarin and english speakers’ conceptions of time. Cognitive psychology 43(1), 1–22 (2001)

  11. [18]

    arXiv preprint arXiv:2303.08774 (2023)

    Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  12. [19]

    arXiv preprint arXiv:2312.11805 1 (2023)

    Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., Millican, K., et al.: Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 1 (2023)

  13. [20]

    arXiv preprint arXiv:2408.04203 (2024)

    Dai, Y., Hu, H., Wang, L., Jin, S., Chen, X., Lu, Z.: Mmrole: A comprehensive framework for developing and evaluating multimodal role-playing agents. arXiv preprint arXiv:2408.04203 (2024)

  14. [21]

    arXiv preprint arXiv:2406.20094 (2024)

    Ge, T., Chan, X., Wang, X., Yu, D., Mi, H., Yu, D.: Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094 (2024)

  15. [22]

    In: Proceedings of the 17th ACM Conference on Recommender Systems, pp

    Dai, S., Shao, N., Zhao, H., Yu, W., Si, Z., Xu, C., Sun, Z., Zhang, X., Xu, J.: Uncovering chatgpt’s capabilities in recommender systems. In: Proceedings of the 17th ACM Conference on Recommender Systems, pp. 1126–1132 (2023)

  16. [23]

    arXiv preprint arXiv:2310.17512 (2023)

    Zhao, Q., Wang, J., Zhang, Y., Jin, Y., Zhu, K., Chen, H., Xie, X.: Competeai: Understanding the competition behaviors in large language model-based agents. arXiv preprint arXiv:2310.17512 (2023)

  17. [24]

    arXiv preprint arXiv:2405.07960 (2024)

    Schmidgall, S., Ziaei, R., Harris, C., Reis, E., Jopling, J., Moor, M.: Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. arXiv preprint arXiv:2405.07960 (2024)

  18. [25]

    arXiv preprint arXiv:2406.17675 (2024)

    Li, Y., Huang, Y., Wang, H., Zhang, X., Zou, J., Sun, L.: Quantifying ai psy- chology: A psychometrics benchmark for large language models. arXiv preprint arXiv:2406.17675 (2024)

  19. [26]

    arXiv preprint arXiv:2402.01622 (2024)

    Xie, J., Zhang, K., Chen, J., Zhu, T., Lou, R., Tian, Y., Xiao, Y., Su, Y.: Trav- elplanner: A benchmark for real-world planning with language agents. arXiv preprint arXiv:2402.01622 (2024)

  20. [27]

    arXiv preprint arXiv:2402.06044 (2024)

    Xu, H., Zhao, R., Zhu, L., Du, J., He, Y.: Opentom: A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models. arXiv preprint arXiv:2402.06044 (2024)

  21. [28]

    In: Proceedings of the 47th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval, pp

    Lin, X., Wang, W., Li, Y., Yang, S., Feng, F., Wei, Y., Chua, T.-S.: Data-efficient 18 fine-tuning for llm-based recommendation. In: Proceedings of the 47th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 365–374 (2024)

  22. [29]

    In: Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, pp

    Coignion, T., Quinton, C., Rouvoy, R.: A performance study of llm-generated code on leetcode. In: Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, pp. 79–89 (2024)

  23. [30]

    In: International Conference on Computers Helping People with Special Needs, pp

    Mucha, W., Cuconasu, F., Etori, N.A., Kalokyri, V., Trappolini, G.: Text2taste: a versatile egocentric vision system for intelligent reading assistance using large language model. In: International Conference on Computers Helping People with Special Needs, pp. 285–291 (2024). Springer

  24. [31]

    Accident Analysis & Prevention 125, 257–266 (2019)

    Dixit, V., Xiong, Z., Jian, S., Saxena, N.: Risk of automated driving: Implications on safety acceptability and productivity. Accident Analysis & Prevention 125, 257–266 (2019)

  25. [32]

    Transportation Research Part C: Emerging Technologies 162, 104603 (2024)

    Liu, J., Jian, S., Wu, C., Dixit, V.: Risky choice and diminishing sensitivity in maas context: A nonlinear logit analysis of traveler behavior. Transportation Research Part C: Emerging Technologies 162, 104603 (2024)

  26. [33]

    Transport Policy82, 36–45 (2019)

    Dixit, V., Jian, S., Hassan, A., Robson, E.: Eliciting perceptions of travel time risk and exploring its impact on value of time. Transport Policy82, 36–45 (2019)

  27. [34]

    Guo, M., Liu, J., Jian, S., Ren, G., Wu, C.: Exploring Subscription and Travel Choice Changes Under Mobility as a Service Bundles: Evidence from Experi- mental Economics Presented at the 104th Annual Meeting of the Transportation Research Board, Washington, DC, United States (2025)

  28. [35]

    Nature 623(7987), 493–498 (2023)

    Shanahan, M., McDonell, K., Reynolds, L.: Role play with large language models. Nature 623(7987), 493–498 (2023)

  29. [36]

    arXiv preprint arXiv:2404.18231 (2024)

    Chen, J., Wang, X., Xu, R., Yuan, S., Zhang, Y., Shi, W., Xie, J., Li, S., Yang, R., Zhu, T., et al.: From persona to personalization: A survey on role-playing language agents. arXiv preprint arXiv:2404.18231 (2024)

  30. [37]

    arXiv e-prints, 2403 (2024)

    Chen, H., Chen, H., Yan, M., Xu, W., Gao, X., Shen, W., Quan, X., Li, C., Zhang, J., Huang, F., et al.: Roleinteract: Evaluating the social interaction of role-playing agents. arXiv e-prints, 2403 (2024)

  31. [38]

    In: Proceedings of the 30th International Conference on Intelligent User Interfaces, pp

    Takagi, H., Moriya, S., Sato, T., Nagao, M., Higuchi, K.: A framework for efficient development and debugging of role-playing agents with large language models. In: Proceedings of the 30th International Conference on Intelligent User Interfaces, pp. 70–88 (2025)

  32. [39]

    Xu, R., Wang, X., Chen, J., Yuan, S., Yuan, X., Liang, J., Chen, Z., Dong, X., 19 Xiao, Y.: Character is destiny: Can role-playing language agents make persona- driven decisions? arXiv preprint arXiv:2404.12138 (2024)

  33. [40]

    In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp

    Alizadeh, K., Mirzadeh, S.I., Belenko, D., Khatamifard, S., Cho, M., Del Mundo, C.C., Rastegari, M., Farajtabar, M.: Llm in a flash: Efficient large language model inference with limited memory. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Li...

  34. [41]

    arXiv preprint arXiv:2412.03563 (2024)

    Mou, X., Ding, X., He, Q., Wang, L., Liang, J., Zhang, X., Sun, L., Lin, J., Zhou, J., Huang, X., et al.: From individual to society: A survey on social simulation driven by large language model-based agents. arXiv preprint arXiv:2412.03563 (2024)

  35. [42]

    arXiv preprint arXiv:2204.05239 (2022)

    Xu, L., Chen, Y., Cui, G., Gao, H., Liu, Z.: Exploring the universal vulnerability of prompt-based learning paradigm. arXiv preprint arXiv:2204.05239 (2022)

  36. [43]

    In: Uncertainty in Economics, pp

    Pratt, J.W.: Risk aversion in the small and in the large. In: Uncertainty in Economics, pp. 59–79. Elsevier, ??? (1978)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.