{"id":"5c055a3e-59cc-44ca-ba54-bb99cc596136","arxiv_id":"2412.00207","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Chatbot self-report personality scores correlate only weakly with human-perceived personality and interaction quality across 500 GPT-4o chatbots, undermining the validity of self-report scales in this context.","lead":"This paper tested whether asking chatbots to fill out human personality questionnaires actually measures the personality that users experience. Across 500 designed chatbots and 500 human interactions, chatbot self-ratings matched user perceptions poorly and barely predicted interaction quality, so the authors argue evaluation should move to task-based interaction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-report prompt embeds the personality description ('respond in a way that matches this description'), so weak human correlations may reflect prompt-compliance artifact rather than invalidity of self-report scales.","rationale":"Reading the paper in good faith: the authors ran a large, transparent study (500 designs, 500 participants, released data), and their descriptive finding that this particular description-anchored self-report protocol agrees poorly with human perception and user experience is credible and useful. The convergent and discriminant analyses, ANOVA tables, and FDR adjustments show careful intent. My concern is not with the data but with the scope of the conclusion. Appendix B shows that the self-report prompt is 'respond in a way that matches this description,' where the description is the manipulated personality prompt. Under that instruction, the model is not being asked to introspect about behavior; it is being asked to comply with a label. The near-ceiling/floor self-report scores in Table 1 (e.g., SD 0.02-0.31, means 4.5-5.0 high / 1.0-1.8 low) are exactly what compliance with an explicit profile would produce. Human perception, by contrast, is constrained by task behavior and by the fact that many traits (e.g., neuroticism) are suppressed in service of the task; the Limitations section acknowledges this. Thus the gap between self-report and human perception may be an artifact of comparing a label-following task with a behavior-based rating task. The reader's conditional verdict already captures this, and I agree with it. If the paper were read only as 'this protocol has limited validity,' it would be fine. But the abstract and Section 4.3 claim a 'substantive disjunction between the traits reflected in questionnaires and the chatbot's actual conversational behavior.' That requires ruling out the alternative explanation that the questionnaire responses reflect the embedded description rather than any personality the chatbot 'has.' The proposed check, removing the description from the self-report prompt and asking for ratings grounded in the preceding conversation, would settle whether the low correlations generalize. Until then, the verdict should remain conditional, with the condition being a cleaner self-report protocol.","tokens_in":21643,"tokens_out":6797,"duration_ms":67733,"concrete_test":"Rerun the self-report collection for the same 500 designed chatbots in two conditions: (A) the original Appendix B prompt with the personality description; (B) a neutral condition in which the model completes the task interaction first and is then asked, without any personality description, to fill out BFI-2-XS using a prompt such as 'Based on your actual responses in the conversation, indicate how well each statement describes you.' Compare criterion correlations against the existing human-perceived and UEQ ratings, and report confidence intervals for the correlations. If condition (B) yields materially higher correlations (e.g., non-AGR traits above 0.4, UEQ above 0.2), the original low correlations are a prompt-compliance artifact and the paper's conclusion must be restricted to description-anchored protocols; if condition (B) remains as weak, the central claim survives this critique.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the scores elicited by Section 3.2.1 and Appendix B are self-reports of the chatbot's personality. In the actual protocol, the model is told, 'For the following task, respond in a way that matches this description: {personality description},' where the description is exactly the prompt used to design the chatbot (Section 3.1.1). This makes each scale response a measure of instruction-following and of whether the model can endorse items consistent with a supplied label. The high convergent correlations across BFI-2-XS/BFI-2/IPIP-NEO-120 (Table 2, mean approx. 0.85) are expected under this protocol: all three scales are administered with the same explicit description, so inter-scale agreement reflects a common input, not independent construct validity. The weak criterion correlations with human perception and UEQ (Tables 4 and 6) could therefore mean that the description-anchored questionnaire protocol produces artificial, context-free responses, or that the personality prompt is weakly realized in interactive behavior. Both interpretations are consistent with the data, but only the second supports the paper's conclusion that 'self-report methods' have limited validity. The conclusion overgeneralizes from a protocol that conflates self-report with direct prompt compliance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale empirical evaluation (500 chatbot configurations, 500 human participants) of whether self-report personality scales administered to GPT-4o-based chatbots have criterion and predictive validity. Chatbots were assigned high or low levels of one Big Five domain using adjective-based prompts across five tasks; their self-report scores were collected with BFI-2-XS, BFI-2, and IPIP-NEO-120, and human participants interacted with the chatbots, rated perceived personality (BFI-2-XS), and rated user experience (UEQ). The main findings are that self-report scores correlate strongly across the three inventories (mean rho ≈ 0.85), correlate only weakly with human-perceived personality except for Agreeableness (rho ≈ 0.58), and correlate weakly with UEQ, whereas human-perceived Agreeableness and Conscientiousness correlate substantially with UEQ in some tasks. The authors conclude that self-report scales have limited criterion and predictive validity for chatbot personality design and advocate task-based, interactive evaluation.","tokens_in":21802,"tokens_out":2881,"duration_ms":27915,"significance":"If the conclusions hold, the paper addresses an important methodological question in LLM-based chatbot evaluation: whether borrowed human personality inventories can serve as valid measures of designed personality. The study's strengths include a controlled design with 500 distinct personality configurations, human interaction data, public release of prompts and data, reproducibility details, and a constructive comparison with a fine-tuned transcript-based evaluator in Section 5. The empirical observation that description-anchored self-reports correlate weakly with human perception is a useful data point. However, the central claim is weakened by the specific self-report protocol used, as detailed in the major comments, so the significance of the paper depends on whether the authors can separate the validity of self-report as a method from the validity of their particular prompt-anchored administration.","major_comments":[{"comment":"The self-report protocol instructs the chatbot: 'For the following task, respond in a way that matches this description: \"{personality description}\"', where the description is exactly the personality prompt used to design the chatbot (Section 3.1.1). This makes each scale response a measure of instruction-following and item-endorsement consistency under an explicit anchoring manipulation, not an independent self-report of the chatbot's personality. The high convergent correlations in Table 2 (mean rho = 0.85) are therefore expected: all three scales are administered with the same explicit description, so inter-scale agreement reflects a common input rather than independent construct validity. The weak criterion correlations with human perception (Table 4) could mean that the description-anchored protocol produces artificial, context-free responses, or that the personality prompt is weakly realized in interactive behavior. Both interpretations are consistent with the data, but only the second supports the paper's conclusion that 'self-report methods' have limited criterion validity. The authors should either add a condition in which the chatbot responds without the embedded description, or substantially narrow the claim to the specific protocol studied.","section":"Appendix B and Section 3.2.1"},{"comment":"The text states that F values 'exceed the conventional threshold of 1', but an F ratio of 1 is not a conventional significance threshold; it merely indicates that between-group variance equals within-group variance. Many F values in Table 9 are close to 1 (e.g., 1.076 for EXT in Social Support, 1.030 for EXT in Job Interview), and these would not be statistically significant under standard F-distribution critical values. The claim that 'personality settings work for both human and chatbot' is therefore not supported by the F values alone. The authors should report formal significance tests or effect sizes with confidence intervals for the human-perceived scores, not just F > 1.","section":"Section 4.1, Table 9"},{"comment":"The comparison between self-report and human-perceived personality as predictors of UEQ is confounded by method variance. Human-perceived personality and UEQ are both rated by the same participants after the same interaction, so shared method variance can inflate their correlations. Self-report scores are generated separately by the model from the prompt and are not subject to that shared context. The conclusion that self-report traits are 'generally poor predictors of interaction quality' relative to perceived traits is therefore not a clean test of predictive validity. The authors should either acknowledge this confound explicitly and qualify the comparison, or use a design where self-report and human perception are obtained from independent sources with comparable measurement conditions.","section":"Section 4.3, Table 6"}],"minor_comments":[{"comment":"The text says 'Table J details the specific instructions used', but the table is numbered Table 15; please correct the cross-reference.","section":"Appendix J"},{"comment":"There is a typo in the note: 'social support task,,' has a double comma.","section":"Appendix I, Table 14"},{"comment":"The abstract and conclusion use the phrase 'self-repord' in the conclusion section (Section 6); this should be 'self-report'.","section":"Conclusion"},{"comment":"For BFI-2-XS, Table 9 shows many 'NA' entries due to no within-group variance, yet Table 10 reports significant p-values for those same cells; please clarify how p-values were computed when variance is zero, or restrict the significance claims to estimable cells.","section":"Appendix G, Table 9 and Appendix H, Table 10"},{"comment":"The limitations section acknowledges potential bias in test choice and the single prompt-based control method, but it does not mention that the self-report prompt in Appendix B embeds the personality description, which is a key protocol-specific limitation that should be discussed.","section":"Limitations (Appendix L)"}],"recommendation":"major_revision","confidential_remarks":"The paper's central empirical finding (weak correlation between the tested self-report protocol and human perception) is plausible and useful, but the protocol in Appendix B conflates self-report with direct prompt compliance. I would not reject the paper, but the authors need to either add an unanchored condition or substantially soften the generalization to 'self-report methods.' The F-value issue and the shared-method-variance confound in the predictive-validity comparison also need attention. The dataset and reproducibility efforts are commendable and should be preserved in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth reading. It is the first large-scale empirical check of whether chatbot self-reports on human personality inventories track what humans perceive in real task interactions. The setup is solid: 500 GPT-4o chatbots with distinct single-domain personality designs, 500 human participants, five task settings, three inventories, and a released interaction dataset. The headline finding—self-report scores correlate weakly (mostly rho < 0.4) with human-perceived personality and with user experience—is credible and useful. The Agreeableness exception (around 0.58) is worth noting.\n\nThe main soft spot is exactly what the stress-test note flags. The self-report prompt in Appendix B tells the model to “respond in a way that matches this description,” where the description is the same personality profile used to design the chatbot. The high convergent correlations across inventories (mean ~0.85) are therefore inflated by a shared input, and the weak criterion correlations may reflect the artificiality of that instruction rather than a general invalidity of self-report methods. The paper’s conclusion overgeneralizes from this protocol. A cleaner test would collect self-reports without embedding the design description, or ask the model to describe its own observed behavior. That said, the negative result is not vacuous: this is the protocol the field has been using, so showing it doesn’t track human perception is a real contribution.\n\nOther soft spots are minor. Correlations in Table 4 lack confidence intervals and significance tests. The fine-tuned evaluator in Section 5 is trained on the human labels it then predicts, so its better correlation is partly circular; it is presented as preliminary, but could be labeled more clearly.\n\nThis paper is for anyone working on LLM evaluation, especially personality and role-playing. The data and dataset are valuable, and the central question is important. It deserves a serious referee. Expect major revision on the interpretation: either run a cleaner self-report protocol or explicitly scope the claim to description-anchored self-reports. I would send it to review on that basis.","headline":"Solid empirical study showing that description-anchored self-reports don't track perceived chatbot personality, but the paper overgeneralizes from a protocol that conflates self-report with prompt compliance.","tokens_in":22387,"tokens_out":2841,"would_cite":true,"duration_ms":25531,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM-based chatbots' answers on human personality inventories correlate weakly with the personalities humans perceive during interaction and with interaction quality, so self-report scales have limited validity for…","keywords":["LLM chatbot personality","self-report validity","Big Five inventory","criterion validity","predictive validity","human perception","task-based evaluation","user experience"],"falsifier":"Run the same 500-chatbot design with a neutral self-report protocol that does not instruct the chatbot to match the profile—for example, asking it to rate whether each item describes its typical responses in the assigned task—and compare the scores with human-perceived personality and interaction quality; strong correlations in that condition would show the weak validity found here is driven by the matching instruction rather than by self-report measurement itself.","tokens_in":21370,"feed_emoji":"🤖","tokens_out":7209,"duration_ms":58699,"temperature":0.7,"pith_summary":"The paper asks whether LLM-based chatbots can meaningfully fill out human personality questionnaires about themselves. It builds 500 chatbots with distinct Big Five designs, has each one complete three standardized personality inventories, and then has 500 human participants interact with those chatbots, rate the chatbot's personality, and rate interaction quality. The self-report scores were internally consistent across inventories but correlated weakly with human-perceived traits (mostly below 0.4, with Agreeableness at 0.58) and even more weakly with user-experience ratings. The authors conclude that self-report scales have limited criterion and predictive validity for evaluating chatbot personality design, and argue that evaluation should be task-based and interaction-grounded.","feed_headline":"Self-report scales misjudge chatbot personality in 500-bot test","feed_subtitle":"In a 500-chatbot study, questionnaire scores track designed traits but not user perception or interaction quality.","key_machinery":"The evaluation pipeline is the central object: each chatbot is given a prompt that pairs a Big Five domain (one of five), a level (high or low), five adjective descriptors, and a task role; the chatbot then completes three self-report inventories, while a separate human participant interacts with the same chatbot in that task and rates its personality with the BFI-2-XS and the interaction with the User Experience Questionnaire. The load-bearing comparison is the Spearman correlation matrix between self-report and human-perceived scores, with the multitrait-multimethod matrix (a correlation matrix comparing multiple traits measured by multiple methods) supplying the convergent and discriminant analysis.","core_discovery":"The paper's central claim is that the 'self-report' personality scores of LLM-based chatbots, obtained by asking the chatbot to rate items from BFI-2-XS, BFI-2, and IPIP-NEO-120, do not track how the chatbot's personality is actually perceived by humans in task-based conversations, nor do they predict the quality of the interaction. While the scales showed moderate convergent and discriminant validity when compared with each other, their correlations with human-perceived personality were weak and unstable across tasks, and their correlations with user-experience ratings were near zero or null in most conditions. The authors interpret this as a substantive disjunction between questionnaire-elicited traits and the chatbot's observable conversational behavior, and therefore as evidence that self-report personality scales alone are insufficient for validating personality design in LLM-based chatbots.","pith_inferences":["Because the self-report prompt tells the chatbot to 'respond in a way that matches' its assigned profile, the weak correlations with human perception may partly reflect an unnatural instruction rather than a general incapacity of chatbots to self-report; testing a neutral phrasing would separate these explanations.","Agreeableness was the one trait with moderately strong self-report-to-perception correlations, which suggests self-report may remain useful for traits that are consistently visible in polite, cooperative dialogue, while failing for context-dependent traits like Conscientiousness or Extraversion.","The same logic likely extends beyond personality: any questionnaire administered to an LLM without grounding in interaction, such as empathy, values, or attitude scales, may show the same gap between what the model says and how its behavior is perceived.","The released transcripts and human ratings could support a different line of work: identifying which conversational cues drive human trait judgments and using them to build evaluation metrics that do not require a separate questionnaire round."],"forward_implications":["Designers who rely on self-report questionnaires to confirm a chatbot's personality may be misled, because scores can reflect the prompt instruction rather than the personality users actually experience.","Evaluation methods for chatbot personality should include human perception, since human ratings of traits such as Agreeableness and Conscientiousness were more strongly tied to interaction quality than self-report scores were.","Task context changes how personality traits show up, so the same trait can be expressed strongly in one task and weakly or even inversely in another, which static questionnaires cannot capture.","A model fine-tuned on human-chatbot transcripts rated personality in closer agreement with human perception than the self-report scales did, suggesting that interaction-based automated evaluation is a viable direction."],"supporting_citations":[{"why":"Supplies the adjective-shape method used to build the high/low Big Five personality profiles assigned to each chatbot.","marker":"Serapio-García et al. (2023)"},{"why":"Provides the BFI-2-XS, the 15-item scale used for both the chatbot self-report and the human-perceived personality ratings.","marker":"Soto & John (2017b)"},{"why":"Provides the full BFI-2 inventory, one of the three self-report instruments.","marker":"Soto & John (2017a)"},{"why":"Provides the IPIP-NEO-120 inventory, the third self-report instrument.","marker":"Johnson (2014b)"},{"why":"Supplies the multitrait-multimethod matrix approach used to evaluate convergent and discriminant validity.","marker":"Campbell & Fiske (1959)"},{"why":"Supplies the meta-analytic baseline that human self-reports and informant reports correlate only about 0.36, the comparison point for the weak self-report-to-perception correlations.","marker":"Connolly et al. (2007)"},{"why":"Provides the User Experience Questionnaire used to measure interaction quality for the predictive-validity analysis.","marker":"Laugwitz et al. (2008)"},{"why":"Supplies the definition of predictive validity that anchors the analysis of whether personality scores predict interaction quality.","marker":"Funder (2006)"}],"fun_headline_variants":["Chatbot self-reports fail to match human perception","500 chatbots reveal self-report personality scores miss the mark","LLM chatbot self-assessments don't predict interaction quality","Self-rated chatbot traits diverge sharply from user experience"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that asking the chatbot to 'respond in a way that matches' its assigned personality profile yields a genuine self-report of that designed personality, rather than a direct prompt-compliant restatement of the profile; if that instruction forces the scale scores, the weak correlations with human perception may be an artifact of the instruction rather than a general property of chatbot self-reports.","fun_headline_variants_meta":{"raw":{"variants":["Chatbot self-reports fail to match human perception","500 chatbots reveal self-report personality scores miss the mark","LLM chatbot self-assessments don't predict interaction quality","Self-rated chatbot traits diverge sharply from user experience"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1457,"prompt_tokens":916,"completion_tokens":541,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":476}},"tokens_in":532,"tokens_out":541,"duration_ms":5770,"temperature":1.0,"reasoning_tokens":476,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:36:56.766648+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 500-chatbot design with a neutral self-report protocol that does not instruct the chatbot to match the profile—for example, asking it to rate whether each item describes its typical responses in the assigned task—and compare the scores with human-perceived personality and interaction quality; strong correlations in that condition would show the weak validity found here is driven by the matching instruction rather than by self-report measurement itself.","supporting_citations":[{"cited_title":"Construction and evaluation of a user experience questionnaire","cited_arxiv_id":null,"evidence_quote":"Provides the User Experience Questionnaire used to measure interaction quality for the predictive-validity analysis."}],"review_version":1}