{"id":"aa77596a-b323-40bb-8ce3-9d113db720d4","arxiv_id":"2601.11049","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs can partially predict individual and sample-level biased decisions in conversational settings, with GPT-4 aligning best, but reproduction of cognitive-load interactions is inconsistent.","lead":"This paper tests whether large language models can predict biased human decisions in chatbot conversations. Results are mixed, but GPT-4 comes closest to human bias patterns, while claims about reproducing cognitive-load effects are only weakly supported.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bias interaction claim rests on a marginal n=6 Spearman correlation (p=.07) in the explicit-bias-prompt condition, which also produces false positives; the abstract overstates this as reproduced.","rationale":"I read the paper in good faith. The human study is careful, pre-registered, and the cognitive-load manipulation is well validated through NASA-TLX and behavioral indicators. The LLM experiments include multiple models, ablations, and perturbation analyses, which is commendable. The most load-bearing issue I find is not the memorization confound (which the authors acknowledge and which is a known challenge for this literature), but the statistical basis for the abstract's headline claim that LLM predictions 'reproduced the same bias patterns and load-bias interactions observed in humans.' The evidence for load-bias interaction reproduction is a single marginal Spearman correlation (rho = 0.771, p = .07) computed over only six choice problems, and it appears only in the HL3 condition where the model was explicitly told to be biased — a condition that also produces false positives and lower accuracy. The neutral HL1/HL2 conditions, which are the more meaningful test of naturalistic simulation, show weaker and sometimes directionally wrong correlations. The paper itself tempers this finding as 'preliminary' and notes that LLMs struggled with load-bias interactions unless explicitly prompted. Therefore, the abstract overstates the strength and generality of the interaction result. This does not change my overall verdict: conditional acceptance with required revisions is appropriate, since the human results and individual-level prediction findings retain value. The concrete test I propose would settle whether the interaction claim survives proper inference.","tokens_in":38782,"tokens_out":5138,"duration_ms":56432,"concrete_test":"Using the released data or Table 6, recompute the Spearman correlation between human z-scores and GPT-4.1 HL3 z-scores with an exact permutation test over the six choice problems (all 720 permutations). If the exact two-sided p >= 0.05, the claim that LLMs reproduce human load-bias interactions is not supported at the conventional level. As a sensitivity check, also recompute the correlation after excluding the Status Quo problems where human z-scores are near zero; if the correlation collapses, the result is driven by two Framing problems and cannot support the general abstract claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims that LLM predictions 'reproduced the same bias patterns and load-bias interactions observed in humans.' The load-bias interaction part is not supported at the conventional significance level. In Section 4.3.3/Table 6, the paper computes z-scores for the change in Cohen's h between complex and simple dialogue across six choice problems. The only positive correlation between human and GPT-4.1 z-scores is in the HL3 condition: Spearman rho = 0.771, p = .07 (Table 7). With n=6, this is marginal at best; the exact permutation p is not reported. Crucially, HL3 is the prompt that explicitly instructs the model to 'Be highly susceptible to cognitive biases.' Table 5 shows that HL3 produces false positives in 5 of the 6 cases where humans showed no bias, and sample-level accuracy drops to 58%. In the neutral conditions HL1 and HL2 — which the paper highlights as reproducing bias patterns — the correlations are 0.600 and not significant, and Table 6 shows the direction reverses for Goal Framing (human +2.29, HL1 -1.61, HL2 -1.81). The paper itself concedes in Section 4.3.3 that 'LLMs struggled to reproduce load-bias interactions, such as the impact of cognitive load, unless explicitly prompted, like in HL3.' Thus the abstract's load-bias reproduction claim rests on a marginal correlation in a condition known to over-predict bias, and it does not generalize across models (GPT-5 correlations are negative; open-source models are weak). This is a statistical overstatement of the central novelty, independent of the memorization confound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pre-registered human experiment (N=1,648) in which participants completed six classic choice problems through a chatbot after either a simple or a cognitively demanding dialogue. The human results show framing effects (risky-choice and goal framing) that are stronger after complex dialogue, status quo effects in two of three scenarios, and NASA-TLX/behavioral validation that complex dialogue increased mental load. The paper then prompts several LLMs (GPT-4.1 family, GPT-5 family, and open-source models) with demographic information and dialogue transcripts at three human-likeness prompt levels, evaluating individual-level prediction accuracy, sample-level bias reproduction, and whether the LLMs reproduce the human load-bias interaction. The authors conclude that LLMs, especially GPT-4.1, reproduced the same bias patterns and load-bias interactions observed in humans.","tokens_in":39198,"tokens_out":5240,"duration_ms":62315,"significance":"The human experiment is a solid, well-controlled contribution: it is pre-registered, powered, and validated with NASA-TLX and recall checks, and the open datasets and code are valuable for future work. If the human findings stand, they provide a useful demonstration that classic biases persist in conversational interfaces and that dialogue complexity can be manipulated to induce cognitive load. The LLM component is more tentative. The central claim that LLM predictions reproduced the same load-bias interactions rests on a single marginal correlation in one explicitly biased prompt condition, and the use of canonical choice problems leaves a serious training-data-contamination concern that the paper acknowledges but does not resolve. The individual-level prediction claim is also weakened by the paper's own perturbation analysis, which shows the models are insensitive to the participant's actual utterances. These issues make the LLM conclusions substantially overstated relative to the evidence.","major_comments":[{"comment":"The abstract claims that LLM predictions 'reproduced the same bias patterns and load-bias interactions observed in humans.' The load-bias interaction is supported only by a Spearman correlation of rho=0.771 with p=.07 for GPT-4.1 in the HL3 condition (n=6). HL1 and HL2 correlations are rho=0.600 and not significant, GPT-5 correlations are negative, and open-source models are weak. Moreover, HL3 is the prompt that explicitly instructs the model to be highly susceptible to cognitive biases, and Table 5 shows it produces false positives in 5 of 6 cases where humans showed no bias, with sample-level accuracy of only 58%. This is not sufficient to claim reproduction of load-bias interactions. Please temper the claim to a marginal, prompt-dependent effect and report the exact permutation p-value.","section":"Abstract; §4.3.3; Table 7"},{"comment":"The load-bias interaction direction is not robust. For Goal Framing, humans show a positive z-score of 2.29, but HL1 and HL2 show negative z-scores (-1.61 and -1.81), meaning the models move in the opposite direction from humans under cognitive load. Only HL3 gives a positive z-score (2.66), and the overall HL3 correlation is driven by a single model-prompt combination. With only six z-score pairs, the Spearman test has very low power; the paper should provide a permutation test and should not present a p=.07 result as 'reproducing' the interaction.","section":"§4.3.3; Table 6"},{"comment":"The paper acknowledges in §5.1 that the LLMs may be 'matching patterns based on learned statistical associations, especially given the widespread use of these choice problems in existing datasets.' However, the Limitations section dismisses this concern by noting that the models have a September 2024 training cutoff while data were collected in 2025. That argument is irrelevant to the actual contamination risk: the six choice problems are canonical (Asian Disease Problem, Samuelson and Zeckhauser scenarios) and predate the cutoff by decades. The LLM results therefore cannot distinguish simulation from memory retrieval unless the authors add novel variants of the choice problems, a held-out control set, or some other contamination test. At minimum, the abstract and discussion must carry the caveat that the bias reproduction may be pattern matching rather than predictive simulation.","section":"§5.1; Limitations"},{"comment":"The individual-level prediction claim is weakened by the paper's own perturbation and ablation results. Replacing the participant's actual responses in the dialogue transcript with randomly generated text leaves prediction accuracy essentially unchanged (Figure 2), and removing demographic information also leaves accuracy roughly unchanged (Section 4.5). This indicates that the LLMs are predicting from the experimental condition and dialogue structure rather than from the individual participant's utterances or demographics. The paper should reframe RQ3 as condition-level or group-level prediction, not individual-level prediction, or provide evidence that the model uses individual-specific information.","section":"§4.3.1; §4.5; Figure 2"}],"minor_comments":[{"comment":"The 'Interaction With Dialogue Complexity' column uses 'Positive'/'Negative' without explanation. Also, the p-values for Simple and Complex dialogue are within-condition significance tests; the interaction claim should be supported by a formal interaction test (e.g., logistic regression with a dialogue-complexity × framing term), not only by visual comparison of confidence intervals.","section":"Table 2"},{"comment":"The confusion matrix totals are unclear: the row 'Not Biased' sums to 3 but there are six choice problems, each with two dialogue conditions. Please clarify the unit of analysis (e.g., 12 condition-level observations) and present the counts consistently.","section":"Table 5"},{"comment":"The phrase 'accuracy (distinct from individual-level prediction accuracy used in Section 4.3.1...)' is confusing. Consider using a different term, such as 'sample-level alignment rate,' to avoid ambiguity.","section":"§4.3.2"},{"comment":"The label 'gpt4_1_blrp' is not defined in the caption. Please spell out that it denotes GPT-4.1 baseline with human response perturbation.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The human study is careful and the authors should be commended for preregistration, power analysis, data/code availability, and multi-method cognitive-load validation. My concern is that the LLM conclusions go well beyond what the data support. The load-bias interaction claim rests on one marginal p=.07 correlation in the explicitly biased HL3 condition, and the training-data contamination issue is not addressed by the training-cutoff argument. I believe the paper can be made publishable by substantially tempering the abstract and claims, adding a formal interaction test and permutation test, and either providing a contamination control or clearly framing the LLM results as exploratory and pattern-matching."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read for you. The paper is better than the abstract suggests in one way and worse in another. The human experiment is a model of care: pre-registered, N=1,648, power analysis, NASA-TLX plus recall checks, response-time validation. That part gives you a solid, reusable demonstration that framing and status quo biases show up in chatbot dialogues, and that a deliberately complex prior dialogue raises mental demand and selectively strengthens framing effects. I'd trust those results.\n\nThe LLM half is the genuine new contribution: asking models to predict individual choices from demographics plus the actual dialogue transcript, and then checking whether the aggregate predictions reproduce the human bias pattern. The methodology is thorough -- seven model variants, three human-likeness prompts, ablations, and a perturbation check. The dialogue-conditioned improvements on Goal Framing and Investment are real.\n\nNow the soft spots. The abstract says the models 'reproduced the same bias patterns and load-bias interactions observed in humans.' That load-bias interaction claim is fragile. It comes from a Spearman correlation across six choice problems between human and LLM z-scores. The only marginally significant positive value is rho = 0.771, p = .07, for GPT-4.1 under HL3 -- the prompt that explicitly tells the model to be highly susceptible to biases. HL3 also produces false positives in five of the six cases where humans showed no bias, and its sample-level accuracy is 58%. HL1 and HL2, the conditions the paper otherwise highlights, give rho = 0.600, not significant, and they actually reverse the Goal Framing interaction. So the abstract is overstating what the data support. The paper does concede this in Section 4.3.3 -- 'LLMs struggled to reproduce load-bias interactions... unless explicitly prompted' -- but the abstract doesn't carry that caveat.\n\nThe second soft spot is the memorization confound, which the authors themselves flag in Section 5.1: the six choice problems are classics, so the models may be retrieving learned associations rather than simulating. That is not a fatal indictment -- the individual-level prediction with dialogue context goes beyond simple retrieval -- but it does undermine the 'simulation' language. I'd want to see at least one novel choice problem to show the effect generalizes.\n\nNet: this is a solid, honest paper with a new application and careful human data, and one clearly overstated summary claim. It deserves a serious referee, but I'd send it back with a request to temper the abstract and add a novel-task control or at least frame the results as preliminary. Worth a reading-group session on LLM-as-proxy methodology.","headline":"The human study is solid and the LLM application is genuinely new, but the abstract overclaims the load-bias interaction, which rides on one marginal n=6 correlation in an explicitly-biased prompt condition.","tokens_in":39603,"tokens_out":2791,"would_cite":true,"duration_ms":29047,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs can predict biased human decisions in conversational settings using limited dialogue and demographics, reproducing population-level bias patterns and their interaction with cognitive load.","keywords":["conversational AI","framing effect","status quo bias","cognitive load","LLM simulation","decision prediction","human-likeness prompting","dialogue complexity"],"falsifier":"Run the same prediction protocol on a set of newly constructed framing and status-quo choice problems that have established human effect sizes but are not present in any LLM training corpus; if the models' predictive accuracy and bias reproduction drop drastically relative to the classic problems, the central claim is driven by memorization rather than generalizable simulation.","tokens_in":38688,"feed_emoji":"🧠","tokens_out":2593,"duration_ms":30822,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models can predict individual human decisions in chatbot conversations when given a participant's demographics and prior dialogue transcript—and that, at the population level, their simulated choices reproduce two classic cognitive biases (the Framing Effect and Status Quo Bias) as well as the way those biases intensify under cognitive load. To test this, the authors ran a pre-registered human study (N = 1,648) in which participants faced six classic decision problems after either a simple or cognitively demanding dialogue. They confirm that framing and status-quo biases appear in conversational settings, that complex dialogue raises mental demand, and that this load selectively amplifies framing effects. They then show that LLMs, especially the GPT-4 family, can predict individual choices more accurately when given dialogue context, and can mirror human bias patterns and load-bias interactions—though the paper itself acknowledges this could be statistical pattern matching rather than genuine simulation.","feed_headline":"LLMs can predict biased choices from chatbot dialogue","feed_subtitle":"GPT-4.1 reproduces framing and status-quo biases, and how cognitive load shifts them, from chat transcripts alone.","key_machinery":"The argument rests on pairing classical choice problems (three framing tasks, three status-quo tasks) with a controlled chatbot dialogue at two complexity levels, and then prompting LLMs at three 'human-likeness' levels: minimal role, naturalistic instruction, and explicit bias instruction. The central quantities are effect sizes (Cohen's h) for bias presence, z-scores for the change in bias under complex dialogue, and Spearman correlations between human and LLM z-score patterns—these carry the claim that LLM behavior can align with human bias and load-bias interactions. Ablations isolating memory vs. arithmetic components of the dialogue indicate that memory cues, not arithmetic, are what a","core_discovery":"On its own terms, the paper claims that LLMs are capable of simulating biased human decision-making in conversational settings: given demographic information plus the transcript of a prior dialogue, GPT-4.1 predictions were significantly more accurate than chance in several choice problems (e.g., Goal Framing accuracy rose from 47% to 63% with dialogue; Investment Decisions from 62% to 76%). At the sample level, models under neutral prompts reproduced the presence or absence of bias across all six choice problems with 75% agreement with human findings, and when explicitly instructed to be biased (HL3) they reproduced the direction of load-bias interactions (Spearman ρ = 0.771, p = .07). The","pith_inferences":["A decisive test the paper leaves implicit: using novel choice problems that were not present in LLM training data would separate genuine human-like generalization from memorization of canonical psychology experiments.","Since removing demographics barely changed predictions, future simulation systems could rely primarily on dialogue structure, which may simplify deployment and reduce privacy concerns.","The ablation finding that memory cues drive load-bias reproduction suggests a testable design principle: chat systems that require users to hold referents in working memory are more likely to amplify framing effects—and could be deliberately calibrated to nudge or debias.","If LLM simulation is accepted as a proxy, the community should treat pre-registered human baselines as the ground truth for every new bias and interaction, since prompt-level overfitting (as seen in HL3) can otherwise produce confidently wrong simulations."],"forward_implications":["Conversational agents could infer a user's bias susceptibility from dialogue history alone and adapt how options are presented, without needing explicit personal data.","LLM-based simulations could serve as low-cost proxies for user studies, enabling rapid A/B testing of dialogue designs for unintended bias amplification.","The selective load-bias interaction implies that increasing conversational complexity is not neutral: it can systematically strengthen framing-type biases while leaving status-quo bias unchanged.","Explicitly instructing models to be biased produces false positives on tasks where humans show no bias, so practical bias-aware simulation requires calibration against real human data.","Model choice matters: GPT-4-family models outperformed GPT-5 and open-source models in both accuracy and bias fidelity, so simulation claims should not be assumed to transfer across model generations."],"fun_headline_variants":["LLMs predict biased choices from chat dialogue","GPT-4.1 mirrors human bias patterns from chat transcripts","Dialogue context boosts LLM prediction of biased decisions","LLMs reproduce framing and status-quo biases from dialogue","Chatbot talk helps AI foresee human cognitive biases"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the LLMs' reproduction of bias patterns reflects simulation of human behavior rather than memorization of these classic choice problems from training data—the paper itself flags this as an open possibility in its discussion.","fun_headline_variants_meta":{"raw":{"variants":["LLMs predict biased choices from chat dialogue","GPT-4.1 mirrors human bias patterns from chat transcripts","Dialogue context boosts LLM prediction of biased decisions","LLMs reproduce framing and status-quo biases from dialogue","Chatbot talk helps AI foresee human cognitive biases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1325,"prompt_tokens":770,"completion_tokens":555,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":514,"tokens_out":555,"duration_ms":6052,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:05:46.259747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same prediction protocol on a set of newly constructed framing and status-quo choice problems that have established human effect sizes but are not present in any LLM training corpus; if the models' predictive accuracy and bias reproduction drop drastically relative to the classic problems, the central claim is driven by memorization rather than generalizable simulation.","supporting_citations":[],"review_version":1}