{"id":"bc3c93e1-4df7-4bb4-9974-35e9771f1d75","arxiv_id":"2502.08265","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs prompted with Big Five trait scores can respond consistently to personality questionnaires but generate free text that often fails to express the prompted trait, especially Neuroticism.","lead":"The paper tests whether four large language models can imitate Big Five personality traits in questionnaire answers and in generated texts. It finds the models answer questionnaires fairly consistently but fail to convey traits like Neuroticism in open text, and it releases a dataset and analysis framework.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4 Omni classifier's weak and unmeasured systematic error on Neuroticism (Level-2 F1=0.50, MAE=1.77) could produce the reported model biases, since all confusion matrices come from this classifier.","rationale":"The reader's weakest assumption correctly identifies the classifier as the load-bearing component. The paper's strongest empirical claims about model-level biases, especially the Neuroticism result, are not directly supported by human evaluation of the full corpus; they flow through a classifier that is weak on that exact trait and of unknown systematic bias. This is a correctness risk, not merely a matter of external consensus. At the same time, the concern is addressable: the paper provides a released framework and a documented annotation protocol, so a targeted human rescoring of a stratified sample would settle whether the classifier artifacts explain the observed patterns. The central claim that simple Big Five prompting is insufficient is plausible and partially supported by the questionnaire-stage results and by the low human agreement on Neuroticism, so I would not move the verdict to REJECT. The reader's CONDITIONAL verdict remains appropriate; no verdict adjustment is needed beyond the conditions already stated.","tokens_in":13552,"tokens_out":4323,"duration_ms":44816,"concrete_test":"Take a stratified sample of the generated corpus (e.g., 50 texts per model per trait, covering all prompted scores and temperatures) and have independent human annotators score them with the same -2 to +2 rubric used in Section 5.2. Then compute per-model and per-trait confusion matrices from the human scores and compare them to the GPT-4 Omni classifier matrices in Figure 2. In particular, report the mean signed difference (classifier minus human) per trait and per score group; if the Neuroticism low-score bias replicates only in the classifier but not in human labels, the central claim is unsupported. As a lighter check, recompute the Neuroticism analysis using an independent classifier (e.g., fine-tuned BERT or a different LLM) and see whether the direction of the bias persists.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that text generation fails to simulate prompted Big Five scores, with Neuroticism the hardest trait and several models biased toward low Neuroticism, is derived almost entirely from the GPT-4 Omni classifier outputs in Section 5.3 and Figure 2. The classifier is validated on only 288 human-annotated texts and is weakest precisely on the trait carrying the headline conclusion: Neuroticism has weighted Level-2 F1=0.50, Level-1 F1=0.63, and MAE=1.77 (Tables 2-3), while human agreement on Neuroticism is also lowest (kappa 0.57-0.59, Table 1). The paper reports only aggregate agreement/MAE, not the direction of classifier errors relative to human scores. If GPT-4 Omni's default assistant persona is high Agreeableness and low Neuroticism, as the paper itself suggests in Section 8, then the classifier may systematically label generated texts as low Neuroticism regardless of their actual content. Because the same model family is both a tested generator and the classifier, and because the classifier prompt shares the trait definitions and 'Nondistinguishable' instruction with the generation task, the reported per-model confusion matrices and trait-level biases could be artifacts of classifier bias rather than genuine simulation failures. The claim that 'some models demonstrated strong bias towards certain score groups' is therefore not securely established without knowing the signed bias of the classifier on the full corpus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLMs can simulate Big Five personality traits in two tasks: answering the BFI-44 questionnaire and generating free-text responses under prompted trait scores. The questionnaire stage reports reliability metrics (Cronbach's alpha, Guttman's lambda) and score distributions per model. The text generation stage prompts GPT-3.5 Turbo, GPT-4 Omni, Claude 3 Haiku, and Mixtral 8x22B with single-trait scores from 1 to 5, then evaluates the outputs via a small human annotation study (288 texts, 8 annotators) and a GPT-4 Omni based automatic classifier. The paper concludes that generating personality-consistent text is still challenging, with Neuroticism the hardest trait and several models showing strong score biases, and it releases a dataset and an analytical framework.","tokens_in":13797,"tokens_out":3359,"duration_ms":35459,"significance":"If the central claims hold, the paper provides a useful negative result for personality prompting: simple Big Five score prompts do not reliably transfer to generated text, and questionnaire behavior can diverge from text-generation behavior. The released dataset and open-source framework are practical contributions, and the questionnaire stage is methodologically more solid, with explicit reliability metrics. However, the text-generation conclusions rest almost entirely on an automatic classifier that is validated on only 288 texts and is weakest exactly on Neuroticism, the trait carrying the headline result; until the classifier's systematic biases are characterized, the per-model conclusions in Section 5.3 and Figure 2 are not secure.","major_comments":[{"comment":"The automated classifier's Neuroticism scores are the least reliable (Level-2 weighted F1 = 0.50, MAE = 1.77), yet Neuroticism is the trait for which the paper makes its strongest claim ('the most challenging trait to simulate'). The paper reports only aggregate precision/recall/F1 and overall MAE, not the direction of classifier errors relative to human judgments. Without knowing whether the classifier systematically maps generated texts to low Neuroticism regardless of content, the confusion matrices in Figure 2 and the related per-model conclusions could be artifacts of classifier bias. Please report per-class and per-direction error analysis (e.g., confusion matrices of classifier vs. human scores for low/middle/high and Nondistinguishable), and if the bias is confirmed, re-analyze the generated-text corpus with a more reliable classifier or a substantially larger human-annotated sample.","section":"Section 5.3, Tables 2-3"},{"comment":"The classifier is built on GPT-4 Omni, which is also one of the models being evaluated for generation, and both generation and classification prompts use the same trait definitions and the same 'Nondistinguishable' instruction. Since Section 8 itself suggests that GPT-4 Omni may have a default persona of high Agreeableness and low Neuroticism, the classifier may share the very bias the paper attributes to the generators. This circularity is load-bearing: the per-model comparisons in Figure 2 and the bullet conclusions about model biases would collapse if the classifier's errors are systematically aligned with its own default persona. Please demonstrate robustness by using an independent classifier (e.g., a fine-tuned model not from the GPT family, or a different model family) or by providing evidence that the classifier's decisions are unbiased with respect to the prompted trait score and the generating model.","section":"Section 5.3 and Section 8"},{"comment":"The manuscript states that Claude responses 'were edited by masking direct references to personality characteristics or by removing content that did not align with the task objectives.' This editing is not quantified, and no criteria or inter-annotator reliability for the editing process are provided. If editing removes trait-relevant cues, it could directly lower the detectability of the prompted trait for Claude and thereby change the comparative conclusions in Section 5.3. Please specify how many responses were edited, what kinds of edits were made, whether the editors were blind to the prompted trait score, and ideally report results on both raw and edited texts to show that editing does not drive the findings.","section":"Section 5.1"},{"comment":"The human evaluation that serves as ground truth for the classifier is based on only 288 texts, with three annotators per text, and the inter-annotator agreement is moderate at best (Level-2 Fleiss kappa ranges from 0.57 for Neuroticism to 0.70 for Agreeableness). Given that Neuroticism also has the lowest human agreement, the conclusion that Neuroticism is especially hard to simulate is partly an artifact of noisier measurement for that trait. The paper should explicitly discuss how the measurement error in the human labels bounds the conclusions, and consider reporting analyses that pool scores into coarser categories or that use the full distribution of annotator scores rather than majority vote.","section":"Section 5.2 and Table 1"}],"minor_comments":[{"comment":"The sentence 'The Claude model exhibited p performance in deciding whether to deliver a direct answer' appears to contain a typo; the word 'p' should likely be 'poor' or another adjective.","section":"Section 5.1"},{"comment":"The phrase 'Inner-annotation agreement metrics' should be 'inter-annotation agreement metrics', since the metrics measure agreement between annotators, not within a single annotation.","section":"Section 5.2"},{"comment":"The trait name is spelled inconsistently as both 'Extraversion' and 'Extroversion'; please choose one spelling (the standard Big Five term is 'Extraversion') and use it consistently.","section":"Tables 1-3 and throughout"},{"comment":"The text refers to 'Figure 4 presents the results of the classifier's detection...' but the figure showing trait detection appears in Appendix B; please renumber or cross-reference the appendix figures clearly so that the reader can locate them.","section":"Section 5.3 and Appendix B"},{"comment":"The mapping between the 1-5 score used in generation prompts and the -2 to +2 scale used by human annotators and the classifier is not stated explicitly; please clarify how prompt scores were mapped to annotation and classification labels, in particular how score 3 (middle) is treated.","section":"Section 5.1 and C.2"},{"comment":"The Limitations section mentions 'partial human dataset annotation', but the main body does not state that the classifier, rather than human judgment, is the primary instrument for all reported text-generation results; please move or echo this caveat in Section 5.3 where the classifier is first used.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper's main empirical contribution depends on the GPT-4 Omni based classifier, and the current validation leaves open the possibility that the headline Neuroticism finding is a classifier artifact. This is fixable with additional validation or re-analysis, so I do not recommend rejection, but the revision must address the circularity and the weak per-trait classifier performance before the claims can be accepted. The released framework and dataset are valuable assets, and the questionnaire-stage results are a solid part of the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper is worth engaging with, but the headline claim about text-generation failure is not as secure as the authors present it. The stress-test note is on target. The GPT-4 Omni classifier produces most of the trait scores, and it is weakest exactly on Neuroticism (Level-2 F1 0.50, MAE 1.77). It shares a model family and trait definitions with the generators, so the specific finding that models bias toward low Neuroticism could be partly a classifier artifact. That said, the paper earns real credit. The two-stage design (questionnaire plus text generation) is sensible, the questionnaire stage includes reliability metrics, and the human evaluation uses Fleiss kappa. The released framework and dataset are a genuine contribution, and the linguistic analysis adds texture. The questionnaire results do not depend on the classifier, and they already show trait differentiation difficulties, so the general conclusion that simple Big Five prompting is insufficient for text generation is plausible.\n\nSoft spots, in proportion: first, the classifier validation is small (288 texts) and the paper reports no signed error or confusion matrix for the classifier itself. Without knowing whether the classifier systematically under-detects or mislabels traits, the per-model confusion matrices in Figure 2 are hard to interpret. Second, some generated texts were edited before analysis; the editing is understandable, but it threatens reproducibility and blurs the measured behavior of the models. Third, human agreement on Neuroticism is the lowest (kappa 0.57–0.59), so even the ground truth for the headline trait is shaky. These are addressable, not fatal. The citation pattern looks fine; prior work on LLM personality is cited appropriately.\n\nWho this is for: people building personalized dialogue agents or game characters, and anyone studying how to elicit personality from LLMs. It is not a theoretical breakthrough, but it is a solid benchmark-style dataset plus a cautionary tale about using LLM-based evaluators. I would send this to peer review with the expectation that the authors add classifier error analysis including signed bias, report results without edited texts or clearly justify the edits, and ideally validate the classifier on a larger human-annotated set. The central argument holds up in spirit, but the evidence needs to be tightened before the trait-level conclusions are accepted.","headline":"Useful dataset and framework for LLM personality simulation, but the key claim about trait bias rests on a classifier that is weakest on the very trait it implicates.","tokens_in":14350,"tokens_out":2296,"would_cite":true,"duration_ms":25468,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Big Five score prompts fail to reliably drive LLM personality in generated text.","keywords":["Big Five personality","large language models","personality simulation","text generation","prompting","LLM evaluation","BFI-44","Neuroticism"],"falsifier":"Have three independent human annotators re-score a random sample of texts that the classifier labeled as low or non-detectable Neuroticism; if they frequently assign high Neuroticism scores, the paper's finding that models cannot simulate that trait would be an artifact of the classifier rather than a property of the models.","tokens_in":13328,"feed_emoji":"🎭","tokens_out":4977,"duration_ms":48325,"temperature":0.7,"pith_summary":"This paper asks whether telling an LLM to write as someone with a given Big Five personality score actually changes the personality expressed in the text it produces. The authors find that simple score-based prompting works for some traits and some models, but not reliably: several models drift toward a default 'helpful assistant' persona—low Neuroticism and high Agreeableness—even when told to simulate the opposite. The hardest trait is Neuroticism, which is often absent from generated texts or stuck at low scores. The paper releases the generated-text dataset and an analytical framework so other researchers can test personality simulation on their own models. The practical stakes are direct: chatbots and game agents are only as personalized as the language they actually produce.","feed_headline":"Big Five score prompts fail to drive LLM personality in text","feed_subtitle":"Across four chatbots, Neuroticism is the hardest trait to simulate, and middle scores collapse toward default biases.","key_machinery":"The argument runs on a two-stage evaluation protocol. First, each model answers BFI-44 questionnaire items under a prompt to adopt a high or low trait score, which tests whether the model knows the trait-behavior link. Second, each model answers open questions about preferences, perspectives, and life goals under a prompt specifying a trait score from 1 to 5, and those generated texts are scored by human annotators and by a GPT-4 Omni classifier built on the CARP clue-extraction method. Confusion matrices comparing prompted versus detected score groups are what turn observed text differences into a claim about personality simulation skill, supplemented by linguistic analyses of vocabulary across score levels.","core_discovery":"The central discovery is that LLMs do not consistently translate an explicit Big Five trait score into a matching personality in free-form text, even though the same models answer personality questionnaires coherently. Using four models (GPT-3.5 Turbo, GPT-4 Omni, Mixtral 8x22B, and Claude 3 Haiku) prompted with scores from 1 to 5 and judged by human annotators plus a GPT-4 Omni classifier, the study finds that Openness to Experience is the most reliably simulated trait, while Agreeableness often comes out high and Neuroticism is frequently undetectable or biased low. Middle scores are particularly unstable: when a model has a default bias, a prompted middle score tends to fall into that biased group rather than sounding neutral. The authors interpret this pattern as evidence that models fall back on an agreeable, emotionally stable assistant persona when generating personality-related text.","pith_inferences":["If the default assistant persona is the cause, then overriding bias may require behavioral examples or system-level instruction rather than more extreme trait scores.","Because the classifier used for most of the corpus is the same model family as one of the generators, the measured trait gaps might be understated; a classifier from a different model family could rank the models differently.","The linguistic-lexicon results suggest a cheap extension: test whether prompts that specify vocabulary style improve trait simulation more than score descriptions alone.","A natural next experiment is to prompt full Big Five profiles rather than single traits and see whether trait interactions reduce or amplify the per-trait biases found here."],"forward_implications":["Defining a persona by a single numeric Big Five score is not enough; generated text should be checked for score bias before deployment in chatbots or game characters.","Middle scores (around 3) are poor prompt targets; binary high/low trait definitions are more likely to yield distinguishable text.","For Neuroticism, the paper recommends setting the trait neutral or omitting it, because high-Neuroticism prompts fail to produce detectable emotional reactivity.","Questionnaire-based personality tests overstate a model's ability to simulate personality, since generation-task performance does not follow questionnaire performance.","Researchers can use the released framework to replicate the protocol on any model and compare trait-level bias directly."],"supporting_citations":[{"why":"Supplies the questionnaire-answering protocol and reliability/validity measurement that the paper adapts to test LLM knowledge of trait-behavior links.","marker":"Serapio-García et al., 2023"},{"why":"Provides the CARP clue-extraction method used to build the GPT-4 Omni classifier that scores the generated texts.","marker":"Sun et al., 2023"},{"why":"Prior finding that questionnaire answers do not predict LLM behavior, motivating the paper's move to text generation.","marker":"Ai et al., 2024"},{"why":"Defines the Big Five taxonomy and the BFI questionnaire that grounds the trait definitions and scoring used throughout.","marker":"John et al., 2008"},{"why":"Supplies the real-user trait distribution used to shape the prompt sampling setup.","marker":"Big Five Personality Test Dataset, 2024"},{"why":"Prior work evaluating and inducing personality in pretrained language models through text, which this study extends to current LLMs.","marker":"Jiang et al., 2024"}],"fun_headline_variants":["LLMs fail to turn Big Five scores into written personality","Big Five prompts don't shape LLM personality in text","Neuroticism hardest trait for LLMs to simulate in text","Middle scores collapse into LLM default personas in text","Openness easy to simulate, neuroticism hard for LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All comparisons of which model best simulates which trait assume the GPT-4 Omni classifier, which humans agree with only moderately and least on Neuroticism, is accurate enough on the full corpus to stand in for human judgment.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fail to turn Big Five scores into written personality","Big Five prompts don't shape LLM personality in text","Neuroticism hardest trait for LLMs to simulate in text","Middle scores collapse into LLM default personas in text","Openness easy to simulate, neuroticism hard for LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000521,"raw_usage":{"total_tokens":2466,"prompt_tokens":837,"completion_tokens":1629,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":1546}},"tokens_in":453,"tokens_out":1629,"duration_ms":11731,"temperature":1.0,"reasoning_tokens":1546,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T05:46:06.780571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have three independent human annotators re-score a random sample of texts that the classifier labeled as low or non-detectable Neuroticism; if they frequently assign high Neuroticism scores, the paper's finding that models cannot simulate that trait would be an artifact of the classifier rather than a property of the models.","supporting_citations":[{"cited_title":"Naumann, and Christopher J","cited_arxiv_id":null,"evidence_quote":"Defines the Big Five taxonomy and the BFI questionnaire that grounds the trait definitions and scoring used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the real-user trait distribution used to shape the prompt sampling setup."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior work evaluating and inducing personality in pretrained language models through text, which this study extends to current LLMs."}],"review_version":1}