{"id":"e4d718fa-6dda-4654-b7ad-f77643b4aa57","arxiv_id":"2412.16772","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Personality-prompted LLMs do not reliably behave in line with the ascribed Big Five traits in Ultimatum Game and Milgram-style tests, with trends sometimes reversing human patterns.","lead":"This paper tests whether giving large language models personality descriptions, like 'very agreeable', changes their decisions in two social experiments, the Ultimatum Game and a Milgram-style obedience test. It finds that prompted personalities often fail to produce human-like behavior, sometimes moving in the opposite direction from human studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central negative result depends on unverified single-study human benchmarks for Openness–UG and Agreeableness–ME; if those correlations do not replicate, the 'opposite direction' finding is uninterpretable.","rationale":"The reader's weakest assumption identifies exactly the load-bearing point: the human benchmark is thin and unvalidated. I agree. The paper's most striking finding—LLMs move opposite to humans for Openness in UG and Agreeableness in ME—is only interpretable if the human correlations are real and in the stated direction. Neither Mehta (2007, a dissertation) nor Bègue et al. (2015, a single study) is accompanied by effect sizes or replications. The paper's robustness checks (alternative prompts, model snapshots) all hold the human benchmark fixed, so they cannot detect a failure of that benchmark. A meta-analysis or direct human replication would settle whether the 'opposite' label is meaningful. The secondary LLM-as-a-judge issue in ME is real but affects only one of the four case studies and is somewhat mitigated by retries and filtering; the human benchmark issue affects both experiments and the central RQ1 answer. Therefore, the reader's CONDITIONAL verdict is appropriate: the paper is a coherent, multi-model negative result, but conditionality is warranted until the human benchmark is verified. My stress-test does not change the verdict.","tokens_in":13522,"tokens_out":6571,"duration_ms":56568,"concrete_test":"Perform a preregistered meta-analysis (or at minimum a systematic review with effect-size extraction) of studies linking Big Five traits to Ultimatum Game responder acceptance, specifically Openness and Agreeableness. If the pooled Openness–acceptance effect is not significantly positive (p<0.05, random effects), the paper's claim that LLMs show an 'opposite' Openness trend is unsupported and RQ1 needs to be revisited for UG. Independently, locate or collect a second human sample for the Milgram paradigm to confirm that Agreeableness positively predicts shock intensity; if the effect does not replicate, the ME opposite-direction conclusion should be withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim—that personality-prompted LLMs fail to behave like humans—rests on the direction of two human correlations: (i) Mehta (2007) reporting that Openness and Agreeableness positively predict accepting unfair UG offers, and (ii) Bègue et al. (2015) reporting that Agreeableness and Conscientiousness positively predict shock intensity in a Milgram paradigm. Both are single studies, and the paper provides no effect sizes, confidence intervals, or independent replication. In the UG case, the LLM Openness trend is downward across all seven models; this is labeled 'opposite to humans' solely on the authority of one dissertation. If a meta-analysis or replication shows the Openness–acceptance correlation is null or negative, the 'opposite direction' becomes a false signal. Similarly, the ME Agreeableness finding is interpreted as human-opposing, but the human baseline is one journal article whose direction may not be robust. The paper's robustness checks perturb the LLM prompts, never the human ground truth, leaving the foundation of the RQ1 negative answer untested. A secondary measurement concern is the LLM-as-a-judge in the ME Stop/Obey steps, which could introduce personality-dependent bias, but the primary load-bearing assumption is the validity and direction of the human correlations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates whether personality-prompted large language models behave consistently with the personality they are prompted to embody, using two classic social decision-making settings: the Ultimatum Game (UG) and the Milgram Experiment (ME). The authors prompt seven models from four vendors with nine levels of trait intensity for Agreeableness, Openness, and Conscientiousness, then measure acceptance rates in UG and withdrawal levels and disobediences in ME. They report two main findings: (i) in two of four case studies (Openness in UG, Agreeableness in ME) the direction of behavior change is opposite to the direction reported in human studies, and in a third case (Conscientiousness in ME) no significant change is found; (ii) the effect of trait intensity on behavior is generally non-monotonic, so behavior cannot be finely steered by prompt intensity. The authors conclude that personality-prompted LLMs should not be assumed to exhibit human-aligned behavior even when questionnaire-based assessments suggest the personality has been induced.","tokens_in":13745,"tokens_out":3945,"duration_ms":36533,"significance":"If the results hold, the paper makes a useful contribution by moving personality evaluation away from self-report questionnaires and toward behavioral benchmarks, and by providing cross-vendor evidence of shared failure modes. The robustness checks (prompt perturbations, model-update comparison, filtering of invalid responses) are valuable, and the negative answer to RQ2 regarding monotonic steering is plausible and interesting. However, the central negative answers to RQ1 rest on the direction of two specific human correlations, neither of which is established with effect sizes or replication in the manuscript, and the main regression trends are reported without uncertainty quantification. These are load-bearing gaps that currently make the headline claim stronger than the evidence supports.","major_comments":[{"comment":"The claim that the Openness–UG trend is 'opposite to human data' rests entirely on Mehta (2007), a single doctoral dissertation, and the paper reports no effect size, confidence interval, or independent replication for the Openness–acceptance correlation. The secondary citation Zhao and Smillie (2015) supports a general link between prosocial traits and bargaining behavior, not specifically the Openness–unfair-offer correlation that the paper needs. As written, the 'opposite direction' finding for Openness is uninterpretable if the human benchmark does not replicate. Please add meta-analytic or multiple-study evidence for the sign of the human correlation, or alternatively reframe the conclusion as 'opposite to the direction reported in Mehta (2007)' and weaken RQ1 accordingly.","section":"§Introduction; §Results and Discussion (Ultimatum Game)"},{"comment":"The regression coefficients Θi in Eq. (1) are the quantitative basis for both RQ1 and RQ2, but Fig. 6 plots point estimates with no confidence intervals, standard errors, or significance tests. Since each Θi is estimated from many trials (e.g., 50 runs per offer level), standard errors are readily available. Without them, the paper cannot support claims that the Openness trend is significantly downward, that the Agreeableness trend is significantly upward, or that non-monotonicity is a robust phenomenon rather than sampling noise. Please report uncertainty for all Θi and provide inferential tests for trend direction and monotonicity.","section":"§Results and Discussion, Eq. (1) and Fig. 6"},{"comment":"The ME measurement of withdrawal levels and disobediences depends on LLM-as-a-judge classifications in the Stop? and Obey? steps, which the paper itself acknowledges as imperfect (citing Zheng et al. 2023). Because the Agreeableness-opposite result is one of the two headline 'opposite direction' findings, a judge that misclassifies hesitation or stopping in a personality-dependent way could produce the observed pattern without reflecting genuine behavioral differences. Please validate the judge against a labeled sample of completions or against the log-probability-based method of Aher et al. (2023), and report agreement rates broken down by personality condition.","section":"§Methodology (Milgram Experiment); Fig. 7, Fig. 8, Table 3"},{"comment":"The paper reports that low- and high-Conscientiousness withdrawal levels are 'not significantly different' from baseline based on Welch's t-test, but no test statistics, p-values, effect sizes, or multiple-comparison corrections are provided. Given that the null Conscientiousness result is one of the four case-study outcomes used to answer RQ1, the absence of these statistics makes the null claim unverifiable. Please report the full test results or replace the claim with descriptive evidence.","section":"§Results and Discussion (Milgram Experiment); Table 3"}],"minor_comments":[{"comment":"The Conclusion refers to varying 'Agreeableness and Consciousness'; this should be 'Conscientiousness'.","section":"Conclusion"},{"comment":"The caption contains a duplicated word: 'results results'.","section":"Fig. 7 caption"},{"comment":"The caption contains a typo: 'Least Agreeablee' should be 'Least Agreeable'.","section":"Fig. 8 caption"},{"comment":"The caption contains a typo: 'firs-order response' should be 'first-order response'.","section":"Table 3 caption"},{"comment":"The non-monotonicity claim is currently supported by visual inspection of Fig. 6; a formal test, such as comparing adjacent trait levels or testing a quadratic term, would make the RQ2 conclusion more rigorous.","section":"§Results and Discussion (Ultimatum Game)"},{"comment":"The citation to Mehta (2007) gives a page number for the UG correlation, but the paper does not state whether the reported correlation is partial, zero-order, or controlled for other Big Five traits; please clarify the statistical basis for the human benchmark.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a computational social science audience and addresses a timely question. My main concern is that the headline negative result depends on the sign of two human correlations, one from a dissertation and one from a single journal article, and the manuscript does not quantify uncertainty in its own central regression trends. These are fixable with additional analysis and more cautious framing; I do not see a need for rejection. I would also suggest the authors make the judge-validation data for the Milgram Experiment available to reviewers, since that mechanism is central to one of the two opposite-direction findings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version: this paper reports something that, if it holds, matters. Prompt-based personality engineering is used in commercial companions, and this paper finds that across seven LLMs, raising prompted Openness flips Ultimatum Game acceptance in the opposite direction from the human correlation in Mehta (2007), and raising Agreeableness flips Milgram obedience opposite to Bègue et al. (2015). Those reversals are consistent across models and survive prompt perturbations and a model-update check. That is a real addition. The paper also does a simple but important control: it verifies with IPIP questionnaires that the prompts actually move the measured trait in the intended direction (Fig. 12), so the failure is not \"the prompt didn't take\" – it's that a successfully induced trait does not produce the corresponding behavior.\n\nThe soft spots are real but not evenly distributed. The biggest one is the human ground truth. The \"opposite direction\" claim rests on two single studies: a dissertation for Openness–UG acceptance, and one journal article for Agreeableness–shock intensity. The paper gives no effect sizes or confidence intervals for those correlations, and the stress-test is right that no perturbation of the human benchmark is explored. That said, the paper does cite Zhao and Smillie (2015) as independent support for the prosocial-bargaining link, which dilutes the single-study worry somewhat. Still, the interpretation would be stronger with a second human source or a meta-analytic effect size.\n\nThe UG regression is presented without error bars on the Θ_i coefficients, and RQ2's \"non-monotonic\" claim is asserted from the plotted pattern rather than tested. A formal monotonicity test and confidence intervals would close that gap. The LLM-as-judge in the Milgram Stop/Obey steps is a known weak point; the authors filter incoherent runs, but judge bias could be personality-dependent and they don't measure it. No code or data is released, which limits the reproducibility of the exact numbers.\n\nThe paper's own Limitations section is honest: \"summoned\" agents in a brief conversation may not capture persistent personality. That is a genuine boundary condition, not a flaw the paper hides.\n\nTake it as a conditional negative result, not a final one. Who benefits: anyone working on persona prompting, AI alignment, or using LLMs for behavioral simulation. It deserves a serious referee – the experimental design and consistency checks are above the threshold. My own verdict would be major-revision: add error bars and a monotonicity test, strengthen the human-benchmark justification with at least one replication or effect size, and release code/data. That is fixable rather than fatal.","headline":"Cross-model negative result on personality prompting is worth taking seriously, but the human-benchmark grounding is thinner than the headline claims.","tokens_in":14309,"tokens_out":1933,"would_cite":true,"duration_ms":16935,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Personality-prompted LLMs do not behave like humans with the same personality in social decision-making tasks.","keywords":["personality prompting","large language models","Ultimatum Game","Milgram experiment","Big Five personality traits","social alignment","LLM-as-a-judge","behavioral benchmarks"],"falsifier":"A replication in which at least one LLM shows acceptance rates for unfair Ultimatum offers increasing monotonically with prompted Openness from 1 to 9, or a Milgram run in which the high-Agreeableness condition reaches a significantly higher mean final shock level than the low-Agreeableness condition under this paper's prompting scheme, would contradict the reported opposite-direction findings.","tokens_in":13290,"feed_emoji":"🤖","tokens_out":8581,"duration_ms":69444,"temperature":0.7,"pith_summary":"Personality prompting is the cheapest and most common way to give a chatbot a character, yet this paper asks whether a prompted trait actually changes what the model does when a social decision must be made. It tests models from four vendors in two classic behavioral experiments, the Ultimatum Game and the Milgram obedience paradigm, with Big Five traits prompted at nine intensity levels. The answer to both research questions is no: in two of four trait–task combinations behavior moved in the opposite direction from the human trend, and in no case did stronger trait prompts produce a monotonic increase in the corresponding behavior. The authors conclude that a model can pass a personality questionnaire while failing to act in a human-aligned way in a consequential social interaction, so personality benchmarks should evaluate behavior in realistic tasks rather than self-report style answers.","feed_headline":"Personality prompts steer LLMs against human social trends","feed_subtitle":"In Ultimatum and Milgram tests, higher trait scores moved models the wrong way and never monotonically.","key_machinery":"The machinery is a pair of standardized social-interaction testbeds borrowed from behavioral economics and social psychology, re-run as iterative prompt loops. In the Ultimatum Game, a responder prompted with a personality description sees offers from $0 to $10 and must accept or reject; acceptance is recorded over 50 runs per offer and trait level. In the Milgram paradigm, a prompted teacher is progressively ordered to increase shock voltage while an LLM-as-a-judge classifies whether the teacher stopped, hesitated, or obeyed. The quantitative core is a linear regression of acceptance on one-hot trait levels plus normalized offer, $y(trait, o) = \\sum_{i=1}^{9} \\Theta_i x_i + \\Theta_o o + c$, whose $\\Theta_i$ coefficients encode whether trait intensity moves behavior and whether that movement is monotonic. Personality induction uses a nine-point adjective-plus-qualifier prompting scale that previous work showed to shift questionnaire scores, and the authors verify that prompted traits register on the IPIP questionnaire in a supplementary check.","core_discovery":"The paper's central claim is that prompt-based personality induction does not reliably transfer to social decision behavior. In the Ultimatum Game, where human studies show Openness and Agreeableness positively correlated with accepting unfair offers, all seven tested models showed increasing Agreeableness increasing acceptance but increasing Openness decreasing acceptance, reversing the human correlation. In the Milgram setup, where human data show Agreeableness and Conscientiousness positively correlated with administering higher shocks, high-Agreeableness models withdrew earlier than low-Agreeableness models in every model that could complete the task, while Conscientiousness produced no significant difference from baseline. Across the four case studies, two results went against the human trend, one was null, and the one aligned trend was not monotonic across the nine trait levels. The paper presents this as evidence that personality-prompted LLMs cannot be expected to exhibit human-aligned behavior by default, even when the model correctly reports the prompted trait on psychological questionnaires.","pith_inferences":["Beyond the paper, a plausible mechanism is that adjective prompts evoke text about a trait rather than a stable decision policy; comparing adjective prompts with prompts containing concrete trait-consistent examples would test this.","The null Conscientiousness result hints that safety and refusal behavior can override the trait signal; one could test whether the opposite-direction effects shrink when the experimental frame is explicitly separated from real harm.","The paper's limitation about brief, consequence-free agents points to a testable extension: run the same protocols with memory and long-term stakes to see whether persistence restores human-like trait effects.","If the Openness reversal is robust, it implies openness-prompted agents may react more strongly to perceived unfairness in other allocation tasks, a prediction that can be checked against existing negotiation benchmarks."],"forward_implications":["If a model answers a personality questionnaire in line with its prompt, that tells us little about whether it will act accordingly when bargaining, obeying orders, or otherwise making a social decision.","Fine control of personality by prompt intensity is not achievable with current adjective-based prompting: levels 1 through 9 do not produce a monotonic behavioral gradient.","The negative results hold across open- and closed-source models from four vendors and persist under several prompt phrasings, so the failure is not isolated to one model or one wording.","Behavioral benchmarks grounded in classic experiments, rather than self-report or text-style evaluation, are needed to validate personality-prompted agents before they are deployed as conversational companions or advisors.","A model that behaves in a human-like way in the unprompted baseline can become less human-like when a personality is injected, so prompting can actively degrade social alignment rather than only fail to improve it."],"supporting_citations":[{"why":"Supplies the Ultimatum Game and Milgram simulation scaffolds and the baseline silicon-population results that this paper modifies and compares against.","marker":"Aher, Arriaga, and Kalai 2023"},{"why":"Provides the nine-level adjective-plus-qualifier personality prompting method used for all trait inductions.","marker":"Serapio-García et al. 2023"},{"why":"Source of the human Ultimatum Game result that Openness and Agreeableness correlate with accepting unfair offers, the target trend for RQ1.","marker":"Mehta 2007"},{"why":"Source of the human Milgram result that Agreeableness and Conscientiousness correlate with higher shock intensity, the target trend for the ME case studies.","marker":"Bègue et al. 2015"},{"why":"Cited for the known imperfection of LLM-as-a-judge classifications, which the paper manages in the Milgram Stop/Obey steps with retries and filtering.","marker":"Zheng et al. 2023"},{"why":"Provides the classic human obedience curve used as the baseline for withdrawal levels in the Milgram setup.","marker":"Milgram 1963"}],"fun_headline_variants":["LLM personality prompts backfire in social tests","Prompted personality fails to humanize LLMs","LLMs' prompted traits diverge from human behavior","Personality prompting misfires in LLM social choices","LLM personality prompts reverse human social patterns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions treat the two human-study correlations used as ground truth, Openness and Agreeableness predicting acceptance of unfair offers, and Agreeableness and Conscientiousness predicting higher shocks, as correct and transferable to a text-based LLM testbed; if either correlation is weak, non-replicable, or does not transfer, the observed reversals can no longer be read as a failure of personality prompting.","fun_headline_variants_meta":{"raw":{"variants":["LLM personality prompts backfire in social tests","Prompted personality fails to humanize LLMs","LLMs' prompted traits diverge from human behavior","Personality prompting misfires in LLM social choices","LLM personality prompts reverse human social patterns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1386,"prompt_tokens":932,"completion_tokens":454,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":548,"tokens_out":454,"duration_ms":4021,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:33.599216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication in which at least one LLM shows acceptance rates for unfair Ultimatum offers increasing monotonically with prompted Openness from 1 to 9, or a Milgram run in which the high-Agreeableness condition reaches a significantly higher mean final shock level than the low-Agreeableness condition under this paper's prompting scheme, would contradict the reported opposite-direction findings.","supporting_citations":[],"review_version":1}