{"id":"92b861c9-98cf-434f-8b10-4bebb74a8dae","arxiv_id":"2506.23107","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"ChatGPT 4o and o1-mini chose more risk-averse lottery options than real respondents in Sydney, Hong Kong, Dhaka, and Nanjing; o1-mini was closer to humans, and Chinese prompts widened the gap.","lead":"This paper asked two large language models to play lottery games as survey respondents from four cities, and compared their choices with real human answers. It found the models were systematically more risk-averse than people, and that prompting them in Chinese moved their answers further from real human choices than English prompts did.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CRRA estimates are not shown to be identified; without a stated choice model per lottery, Tables 6–9 may compare artifacts rather than risk preferences.","rationale":"The reader's weakest_assumption is exactly the unstated probabilistic choice model and comparability of r across the four lottery designs. Reading the full text confirms that Section 4.2 gives only the CRRA utility function and the linear r equation with no likelihood or mapping from discrete choices to r. The tables give means and star annotations only, so the reader's CONDITIONAL verdict is appropriate. I did not find a stronger internal inconsistency that would change the verdict; the Dhaka/Nanjing table format is confusing but the 'Prob. payoff 2' labels and EV columns appear interpretable. The paper has real strengths: it uses four real datasets, a clear LLM-vs-human comparison design, multiple simulations per choice, and reports a counterintuitive language effect. But those strengths do not fix the identification gap. One could argue the concern is 'outside current consensus' if CRRA is standard, but the issue is not the choice of CRRA per se; it is that no estimation model is specified at all, so no error bar, t-statistic, or robustness claim can be checked. The concrete test I propose—re-estimating with an explicit logit CRRA model—would settle whether the reported means are robust. If the authors already have this code (likely, given the t-tests), providing it plus a sensitivity analysis would resolve the concern, so the verdict should remain CONDITIONAL rather than REJECT or UNVERDICTED.","tokens_in":12984,"tokens_out":1526,"duration_ms":15345,"concrete_test":"Re-estimate all means in Tables 6–9 with an explicit choice model, e.g., Luce/logit choice probability P(choose risky) = 1/(1+exp(-lambda*(EV_risky^gamma - EV_safe^gamma))) or CRRA logit per individual, and report the distribution of estimated r per city/model. Then check whether the qualitative ordering (4o > o1-mini > human; Chinese > English) survives. If the ordering flips or vanishes under a defensible estimation model, the headline claim is an artifact of the unspecified estimator.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.2 defines U(x) = x^(1-r)/(1-r) and invokes r = r0 + alpha1*X1 from [44], but it never states the probabilistic choice rule that maps binary lottery answers to individual r estimates, nor whether each individual gets one r or one per task. Tables 6–9 then treat the resulting means as comparable across four surveys that differ in payoff magnitudes, task counts, and risky-safe vs. risky-risky framing (Tables 2–5). The headline claims that LLMs are 'more risk averse than humans' and that Chinese prompts 'increase overestimation' depend entirely on this estimation and comparability step. The risk-averse ordering within a CRRA model is scale-invariant for singleton gambles in payoff units, but mixture behavior across the four lottery designs is not invariant, so nominal r differences across cities and languages can reflect design artifacts or LLM choice-pattern deviations (e.g., probability weighting, random noise, option-order effects) rather than curvature differences. The paper presents no likelihood, no parameter standard errors, no per-city variance, and no robustness check with an alternative model, so the central comparison is not identifiable from the manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether ChatGPT 4o and ChatGPT o1-mini can mimic human risky choices by simulating respondents in four transport stated-preference surveys (Sydney, Hong Kong, Dhaka, Nanjing). The authors feed each LLM a profile built from age, gender, education, and income, have the model answer repeated binary lottery tasks in English (and, for Hong Kong and Nanjing, also in Chinese), take a majority vote over three runs, and estimate CRRA risk-aversion parameters r using the specification r = r0 + α1X1 from reference [44]. They conclude that LLM-simulated respondents are generally more risk-averse than humans, that o1-mini is closer to human estimates than 4o, and that Chinese prompts produce larger deviations. The paper positions this as evidence of an intrinsic conservative bias and warns against using LLMs for risk-sensitive policymaking without calibration.","tokens_in":13177,"tokens_out":8942,"duration_ms":87991,"significance":"If the estimation were fully specified and reproducible, this would be a useful cross-cultural, cross-lingual benchmark for LLM-based behavioral simulation: four diverse datasets, two models, controlled demographic conditioning, and a formal utility framework. The paper also makes concrete, falsifiable claims about systematic overestimation and language-dependent bias, and it draws policy implications that are directly testable. However, the manuscript does not provide the choice model, likelihood, standard errors, or code/data needed to verify the point estimates, and one headline claim is contradicted by the paper's own Table 7. The contribution is potentially valuable but not yet verifiable in this form.","major_comments":[{"comment":"The paper never specifies the probabilistic choice model that maps each binary lottery answer to an individual r estimate. Eq. (1) defines CRRA utility and the text invokes r = r0 + α1X1 from [44], but there is no likelihood, no statement of whether r is estimated per person or from all choices pooled, and no parameter uncertainty. Tables 6–9 therefore report means of an unstated estimator, and the paired t-tests in Section 5.1 are presented only through significance stars, with no standard errors, p-values, or test statistics. The central quantitative claims that LLMs are more risk-averse than humans and that o1-mini is closer cannot be assessed without this identification step. Please state the choice model (e.g., a logit over CRRA expected utilities with a specific noise term), the estimation method, and report standard errors or confidence intervals for each mean.","section":"§4.2, Eq. (1)"},{"comment":"The headline claim that “both ChatGPT 4o and ChatGPT o1-mini systematically exhibited higher levels of risk aversion compared to actual human respondents” is contradicted by Table 7. For Hong Kong, the o1-mini mean is 0.509 versus the real mean of 0.765, so the model is less risk-averse than the human sample. The text in Section 5.1 carefully avoids saying “higher” for that cell, but the abstract and Section 6.1 overstate the result. Please qualify the claim to the datasets/cells where the direction holds, or add a statistical test that supports a general “systematic” direction.","section":"Abstract; §6.1"},{"comment":"The Dhaka lottery description conflicts with Table 4. The text states that in the left lottery the lower payoff 6 occurs with probability p and the higher payoff 8 with probability 1−p, and symmetrically for the right lottery (1 with p and 20 with 1−p). However, the table's expected values (e.g., 6.2 = 0.1×8 + 0.9×6 and 2.9 = 0.1×20 + 0.9×1) show that p is instead the probability of the first-listed payoff, i.e., the higher payoff in both lotteries. Because every Dhaka CRRA estimate, and hence the Dhaka rows of Tables 6–8 and Figure 1(c), depends on the correct probability mapping, this discrepancy must be resolved and corrected.","section":"§3.2, Table 4"},{"comment":"Table 9 does not test the English-versus-Chinese difference that Section 5.2 claims. The reported significance stars are for each simulated mean relative to the real mean (the column header “Ha: mean(diff) ≠ 0”), not for the difference between the English and Chinese simulated means. To support the statement that switching to Chinese prompts consistently amplifies the bias, the authors need a paired test of English versus Chinese simulated risk attitudes, with standard errors, and ideally a test on absolute deviations from the real values. As it stands, the linguistic effect is not statistically established.","section":"§5.2, Table 9"},{"comment":"The four lottery designs are not directly comparable for the cross-city claim. Sydney and Hong Kong pair a risky option against a fixed payoff; Dhaka and Nanjing pair two risky options; and the number of tasks (9 versus 10), payoff magnitudes, and probability grids differ. Under CRRA expected utility, nominal r comparisons are only meaningful if the same choice rule is assumed across designs. The manuscript should include a robustness analysis using a single structural model, or an alternative model, to show that the cross-city and cross-language comparisons are not artifacts of design differences, option-order effects, or probability-weighting deviations in LLM choices.","section":"§3.2, Tables 2–5"}],"minor_comments":[{"comment":"The heading “Planing” is a typo and should read “Planning”.","section":"§4.1"},{"comment":"The header row “Prob. payoff 1 2” is unclear; define the columns (e.g., Prob(payoff 1), Prob(payoff 2)) and label which payoff is lower and which is higher, consistent with the text.","section":"Table 4"},{"comment":"The sample prompt tells the agent that there will be “nine consecutive lottery questions,” but the Dhaka and Nanjing surveys have ten tasks; the prompt should be dataset-specific.","section":"§4.1"},{"comment":"The prompt text “a age-year-old” should be “an age-year-old”.","section":"§4.1"},{"comment":"Eq. (1) is undefined at r = 1; adding the standard limiting form would avoid a gap for estimates near risk neutrality.","section":"§4.2, Eq. (1)"},{"comment":"The significance-star notation is used without t-statistics or p-values; even after the estimation is clarified, the tables should include measures of uncertainty.","section":"Tables 6–9"},{"comment":"The kernel density plots should state the bandwidth and the sample sizes; with n = 64 in Sydney, density estimates are noisy and may be misleading.","section":"Figure 1"},{"comment":"For a simulation study with bespoke prompts and a proprietary LLM, releasing at least the prompts and the estimation code would be necessary for independent verification; the current statement limits access to aggregated data only.","section":"Data availability"},{"comment":"The text says “the primary focus should be on the differences in expected values between the left and right options,” but for CRRA estimation the full payoff distributions, not just EV differences, are needed; the text should not suggest otherwise.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The central problem is verifiability: the estimation procedure is absent, and the headline claim is internally contradicted by the Hong Kong o1-mini row in Table 7. I do not see these as fatal if the authors can provide the full specification, standard errors, corrected claims, and a proper test of the language effect; hence major revision rather than rejection. The paper's scope—transport SP data, LLM simulation, and risk preferences—fits the journal's interests, but the authors should also address the dependence on their own prior work [44] for the central estimating equation, since a reader cannot check the likelihood from the current citation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful empirical benchmark, not a methodological breakthrough. The headline result — two ChatGPT models mostly come out more risk-averse than human respondents on real lottery data, and Chinese prompts widen the gap — is plausible and worth knowing. But the paper needs a major revision before the numbers can be trusted.\n\nWhat's new: applying role-playing language agents to four real stated-preference lottery datasets across four cities, with demographic profiles, is a legitimate new empirical setting. The prompt-language effect in Chinese is a genuinely new observation, and the majority-voting over three runs is a sensible noise-reduction step.\n\nThe biggest problem is the estimation step. Section 4.2 gives the CRRA utility function and cites [44] for r = r0 + α1X1, but never states how an individual's r is recovered from a series of binary choices — no likelihood, no choice rule, no standard errors. Tables 6-9 report means and significance stars only. Without that, I cannot tell whether the differences reflect risk-preference curvature or artifacts of the choice model, the lottery mix, or the majority-voting rule. The comparability of r across four survey designs with different payoff magnitudes and risky-safe vs risky-risky frames is asserted, not defended. The stress-test concern about identification is fair.\n\nThere is also a factual overstatement in the abstract and Section 6.1: they say both models exhibit more risk-aversion than humans, but Table 7 shows o1-mini in Hong Kong at 0.509 vs real 0.765 — less risk-averse. That needs correction. Minor: the sample prompt says \"nine consecutive lottery questions\" but Dhaka and Nanjing have ten tasks. Data and code are not public, so a reader cannot verify any of the central estimates.\n\nWho this is for: researchers using LLM agents to simulate travelers or consumers; it gives a concrete caution. The paper deserves peer review after the estimation procedure is specified, the overclaim fixed, and ideally the data/code released. I'd send it with a request for major revision.","headline":"Useful benchmark result, but the CRRA estimation is under-specified and the abstract overstates the findings.","tokens_in":13719,"tokens_out":4727,"would_cite":true,"duration_ms":45999,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that LLM-simulated survey respondents, built from demographic profiles, choose more risk-averse lottery options than the human respondents they mimic, and that prompting in Chinese widens the divergence.","keywords":["large language models","risk preferences","CRRA","lottery choice","role-playing language agents","cross-cultural study","stated preference survey","prompt language"],"falsifier":"Re-estimate the risk coefficients from the same human and LLM choices with an explicit stochastic choice model, such as a logistic choice rule over CRRA utilities with an error term, and compare the LLM-human gap; if the systematic conservative bias disappears or reverses under that specification, the headline result is an artifact of the unstated mapping from choices to risk scores.","tokens_in":12767,"feed_emoji":"🎲","tokens_out":10866,"duration_ms":103623,"temperature":0.7,"pith_summary":"The paper asks whether large language models can stand in for human respondents in risky-choice surveys. Across four cities, simulated respondents built from age, gender, education, income, and city cues chose more cautiously in lottery games than the real participants, and their estimated risk-aversion coefficients were generally higher. ChatGPT o1-mini tracked human choices more closely than ChatGPT 4o did. When the prompt language for Hong Kong and Nanjing was switched from English to Chinese, the distance between simulated and real risk attitudes grew, even though those respondents natively speak Chinese. If these results hold, LLM-based stand-ins for human survey respondents will need calibration before they can be trusted in policy or market research about risk.","feed_headline":"LLM survey agents skew more risk-averse than humans","feed_subtitle":"The gap widens when prompts switch to Chinese—a caution for AI-based surveys and policy modeling.","key_machinery":"The central object is the CRRA risk-preference coefficient $r$, defined through the utility function $U(x)=x^{1-r}/(1-r)$, with $r>0$ meaning risk aversion, $r=0$ risk neutrality, and $r<0$ risk seeking. The paper adopts the linear specification $r = r_0 + \\alpha_1 X_1$ from reference [44] to connect socio-demographic variables $X_1$ to risk preferences, and it builds role-playing agents whose prompts combine an individual profile (age, gender, education, income, city) with chain-of-thought planning instructions and a closed-form action rule that outputs only the chosen lottery option. Each task is simulated three times and the majority choice is kept as the simulated response.","core_discovery":"Using the Constant Relative Risk Aversion (CRRA) framework, the paper estimates risk-preference coefficients $r$ from real and LLM-simulated lottery choices. It finds that ChatGPT 4o's simulated respondents are systematically more risk-averse than the human respondents in all four city samples, and that ChatGPT o1-mini is generally more risk-averse than humans as well, though its estimates are closer to the observed human values, with Hong Kong as the one reported exception. The paper further finds that for the Hong Kong and Nanjing samples, presenting the same tasks in Chinese rather than English moves the simulated risk attitudes further from the human estimates, a bias amplification the authors attribute to the models' predominantly English-language training.","pith_inferences":["A testable extension: vary the stakes or the gain-loss framing while holding the demographic profile fixed; if the conservative bias grows with stakes, the mechanism is closer to loss aversion or safety-oriented calibration than to CRRA curvature alone.","An inference not in the paper: because the mapping from lottery answers to individual $r$ estimates is never spelled out, an explicit stochastic choice model, such as a logit over CRRA utilities, could change the size or even the sign of the reported LLM-human gap.","A testable extension: run the same English-versus-Chinese comparison with models trained primarily on Chinese corpora; the paper's training-bias account predicts the gap would shrink or reverse.","An inference not in the paper: the four lotteries differ in payoff magnitudes and in risky-safe versus risky-risky framing, so a matched-design replication is needed before the cross-city differences in human risk aversion can be read as cultural rather than artifactual."],"forward_implications":["Uncalibrated LLM stand-ins will make populations look more risk-averse than they are, pushing policy recommendations toward overly conservative measures.","Model choice matters: ChatGPT o1-mini is a closer proxy for human risk preferences than ChatGPT 4o in these tasks, so conclusions drawn from one LLM may not transfer to another.","Prompt language is an active variable: even when respondents natively speak Chinese, Chinese prompts can degrade fidelity, so localization alone does not improve simulation accuracy.","Demographically narrow samples, such as the Dhaka taxi-driver data, can erase measurable differences between models, suggesting that LLMs fall back on generic rather than person-specific risk priors."],"supporting_citations":[{"why":"Supplies the Sydney dataset's nine lottery tasks and the human responses used in the real versus simulated comparison.","marker":"[31]"},{"why":"Supplies the Hong Kong dataset with scaled lottery payoffs and the human responses used for the Hong Kong comparison.","marker":"[32]"},{"why":"Supplies the Dhaka dataset of taxi-driver respondents and their risky-risky lottery choices.","marker":"[33]"},{"why":"Supplies the Nanjing dataset and its lottery design for the Nanjing comparison.","marker":"[34]"},{"why":"Provides the linear CRRA risk-preference specification $r = r_0 + \\alpha_1 X_1$ that the paper adopts to estimate risk attitudes from choices.","marker":"[44]"}],"fun_headline_variants":["LLMs more risk-averse than humans in four-city lotteries","Chinese prompts widen LLM-human risk preference gap","ChatGPT o1-mini beats 4o on human risk mimicry","LLM risk aversion: language matters, Chinese worsens","Survey AI skews risk-averse, more so in Chinese"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that a single CRRA risk-score equation taken from an earlier transportation study converts each respondent's lottery answers into a comparable risk-aversion number across four differently designed lotteries, but the paper does not state the choice model that performs this conversion.","fun_headline_variants_meta":{"raw":{"variants":["LLMs more risk-averse than humans in four-city lotteries","Chinese prompts widen LLM-human risk preference gap","ChatGPT o1-mini beats 4o on human risk mimicry","LLM risk aversion: language matters, Chinese worsens","Survey AI skews risk-averse, more so in Chinese"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1685,"prompt_tokens":907,"completion_tokens":778,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":691}},"tokens_in":523,"tokens_out":778,"duration_ms":8443,"temperature":1.0,"reasoning_tokens":691,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:49:19.902730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the risk coefficients from the same human and LLM choices with an explicit stochastic choice model, such as a logistic choice rule over CRRA utilities with an error term, and compare the LLM-human gap; if the systematic conservative bias disappears or reverses under that specification, the headline result is an artifact of the unstated mapping from choices to risk scores.","supporting_citations":[{"cited_title":"Accident Analysis & Prevention 125, 257–266 (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the Sydney dataset's nine lottery tasks and the human responses used in the real versus simulated comparison."},{"cited_title":"Transportation Research Part C: Emerging Technologies 162, 104603 (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the Hong Kong dataset with scaled lottery payoffs and the human responses used for the Hong Kong comparison."},{"cited_title":"Transport Policy82, 36–45 (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the Dhaka dataset of taxi-driver respondents and their risky-risky lottery choices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Nanjing dataset and its lottery design for the Nanjing comparison."},{"cited_title":"Travel behaviour and society 33, 100604 (2023) 20","cited_arxiv_id":null,"evidence_quote":"Provides the linear CRRA risk-preference specification $r = r_0 + \\alpha_1 X_1$ that the paper adopts to estimate risk attitudes from choices."}],"review_version":1}