REVIEW 3 major objections 5 minor 26 references
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that LLM-based chatbots' answers on human personality inventories correlate weakly with the personalities humans perceive during interaction and with interaction quality, so self-report scales have limited validity for…
desk verdict Solid empirical study showing that description-anchored self-reports don't track perceived chatbot personality, but the paper overgeneralizes from a protocol that conflates self-report with prompt compliance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluation pipeline is the central object: each chatbot is given a prompt that pairs a Big Five domain (one of five), a level (high or low), five adjective descriptors, and a task role; the chatbot then completes three self-report inventories, while a separate human participant interacts with the same chatbot in that task and rates its personality with the BFI-2-XS and the interaction with the User Experience Questionnaire. The load-bearing comparison is the Spearman correlation matrix between self-report and human-perceived scores, with the multitrait-multimethod matrix (a correlation matrix comparing multiple traits measured by multiple methods) supplying the convergent and discriminant analysis.
What would settle it
Run the same 500-chatbot design with a neutral self-report protocol that does not instruct the chatbot to match the profile—for example, asking it to rate whether each item describes its typical responses in the assigned task—and compare the scores with human-perceived personality and interaction quality; strong correlations in that condition would show the weak validity found here is driven by the matching instruction rather than by self-report measurement itself.
Extended reading notes
Core claim
The paper's central claim is that the 'self-report' personality scores of LLM-based chatbots, obtained by asking the chatbot to rate items from BFI-2-XS, BFI-2, and IPIP-NEO-120, do not track how the chatbot's personality is actually perceived by humans in task-based conversations, nor do they predict the quality of the interaction. While the scales showed moderate convergent and discriminant validity when compared with each other, their correlations with human-perceived personality were weak and unstable across tasks, and their correlations with user-experience ratings were near zero or null in most conditions. The authors interpret this as a substantive disjunction between questionnaire-elicited traits and the chatbot's observable conversational behavior, and therefore as evidence that self-report personality scales alone are insufficient for validating personality design in LLM-based chatbots.
Load-bearing premise
The load-bearing assumption is that asking the chatbot to 'respond in a way that matches' its assigned personality profile yields a genuine self-report of that designed personality, rather than a direct prompt-compliant restatement of the profile; if that instruction forces the scale scores, the weak correlations with human perception may be an artifact of the instruction rather than a general property of chatbot self-reports.
Editorial extensions
If this is right
- Designers who rely on self-report questionnaires to confirm a chatbot's personality may be misled, because scores can reflect the prompt instruction rather than the personality users actually experience.
- Evaluation methods for chatbot personality should include human perception, since human ratings of traits such as Agreeableness and Conscientiousness were more strongly tied to interaction quality than self-report scores were.
- Task context changes how personality traits show up, so the same trait can be expressed strongly in one task and weakly or even inversely in another, which static questionnaires cannot capture.
- A model fine-tuned on human-chatbot transcripts rated personality in closer agreement with human perception than the self-report scales did, suggesting that interaction-based automated evaluation is a viable direction.
Reading between the lines
- Because the self-report prompt tells the chatbot to 'respond in a way that matches' its assigned profile, the weak correlations with human perception may partly reflect an unnatural instruction rather than a general incapacity of chatbots to self-report; testing a neutral phrasing would separate these explanations.
- Agreeableness was the one trait with moderately strong self-report-to-perception correlations, which suggests self-report may remain useful for traits that are consistently visible in polite, cooperative dialogue, while failing for context-dependent traits like Conscientiousness or Extraversion.
- The same logic likely extends beyond personality: any questionnaire administered to an LLM without grounding in interaction, such as empathy, values, or attitude scales, may show the same gap between what the model says and how its behavior is perceived.
- The released transcripts and human ratings could support a different line of work: identifying which conversational cues drive human trait judgments and using them to build evaluation metrics that do not require a separate questionnaire round.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a large-scale empirical evaluation (500 chatbot configurations, 500 human participants) of whether self-report personality scales administered to GPT-4o-based chatbots have criterion and predictive validity. Chatbots were assigned high or low levels of one Big Five domain using adjective-based prompts across five tasks; their self-report scores were collected with BFI-2-XS, BFI-2, and IPIP-NEO-120, and human participants interacted with the chatbots, rated perceived personality (BFI-2-XS), and rated user experience (UEQ). The main findings are that self-report scores correlate strongly across the three inventories (mean rho ≈ 0.85), correlate only weakly with human-perceived personality except for Agreeableness (rho ≈ 0.58), and correlate weakly with UEQ, whereas human-perceived Agreeableness and Conscientiousness correlate substantially with UEQ in some tasks. The authors conclude that self-report scales have limited criterion and predictive validity for chatbot personality design and advocate task-based, interactive evaluation.
Significance. If the conclusions hold, the paper addresses an important methodological question in LLM-based chatbot evaluation: whether borrowed human personality inventories can serve as valid measures of designed personality. The study's strengths include a controlled design with 500 distinct personality configurations, human interaction data, public release of prompts and data, reproducibility details, and a constructive comparison with a fine-tuned transcript-based evaluator in Section 5. The empirical observation that description-anchored self-reports correlate weakly with human perception is a useful data point. However, the central claim is weakened by the specific self-report protocol used, as detailed in the major comments, so the significance of the paper depends on whether the authors can separate the validity of self-report as a method from the validity of their particular prompt-anchored administration.
major comments (3)
- [Appendix B and Section 3.2.1] The self-report protocol instructs the chatbot: 'For the following task, respond in a way that matches this description: "{personality description}"', where the description is exactly the personality prompt used to design the chatbot (Section 3.1.1). This makes each scale response a measure of instruction-following and item-endorsement consistency under an explicit anchoring manipulation, not an independent self-report of the chatbot's personality. The high convergent correlations in Table 2 (mean rho = 0.85) are therefore expected: all three scales are administered with the same explicit description, so inter-scale agreement reflects a common input rather than independent construct validity. The weak criterion correlations with human perception (Table 4) could mean that the description-anchored protocol produces artificial, context-free responses, or that the personality prompt is weakly realized in interactive behavior. Both interpretations are consistent with the data, but only the second supports the paper's conclusion that 'self-report methods' have limited criterion validity. The authors should either add a condition in which the chatbot responds without the embedded description, or substantially narrow the claim to the specific protocol studied.
- [Section 4.1, Table 9] The text states that F values 'exceed the conventional threshold of 1', but an F ratio of 1 is not a conventional significance threshold; it merely indicates that between-group variance equals within-group variance. Many F values in Table 9 are close to 1 (e.g., 1.076 for EXT in Social Support, 1.030 for EXT in Job Interview), and these would not be statistically significant under standard F-distribution critical values. The claim that 'personality settings work for both human and chatbot' is therefore not supported by the F values alone. The authors should report formal significance tests or effect sizes with confidence intervals for the human-perceived scores, not just F > 1.
- [Section 4.3, Table 6] The comparison between self-report and human-perceived personality as predictors of UEQ is confounded by method variance. Human-perceived personality and UEQ are both rated by the same participants after the same interaction, so shared method variance can inflate their correlations. Self-report scores are generated separately by the model from the prompt and are not subject to that shared context. The conclusion that self-report traits are 'generally poor predictors of interaction quality' relative to perceived traits is therefore not a clean test of predictive validity. The authors should either acknowledge this confound explicitly and qualify the comparison, or use a design where self-report and human perception are obtained from independent sources with comparable measurement conditions.
minor comments (5)
- [Appendix J] The text says 'Table J details the specific instructions used', but the table is numbered Table 15; please correct the cross-reference.
- [Appendix I, Table 14] There is a typo in the note: 'social support task,,' has a double comma.
- [Conclusion] The abstract and conclusion use the phrase 'self-repord' in the conclusion section (Section 6); this should be 'self-report'.
- [Appendix G, Table 9 and Appendix H, Table 10] For BFI-2-XS, Table 9 shows many 'NA' entries due to no within-group variance, yet Table 10 reports significant p-values for those same cells; please clarify how p-values were computed when variance is zero, or restrict the significance claims to estimable cells.
- [Limitations (Appendix L)] The limitations section acknowledges potential bias in test choice and the single prompt-based control method, but it does not mention that the self-report prompt in Appendix B embeds the personality description, which is a key protocol-specific limitation that should be discussed.
Circularity Check
Self-report prompt embeds the design description, making internal validity checks tautological; the fine-tuned evaluator is fit to the human-perceived criterion. The central negative result remains externally grounded.
-
self definitional
[Section 3.2.1, Appendix B, Section 4.1 and Table 2]
"Following the work of Serapio-García et al. (2023), we instructed the chatbot to rate test items using a standardized response scale from one to five based on the personality descriptions. ... For the following task, respond in a way that matches this description: '{personality description}.' ... where the personality description is the same as the task-based personality description introduced in Section 3.1.1."
The self-report score is generated under an explicit instruction to match the exact personality description used to design the chatbot. Consequently, the high-versus-low condition differences in Table 1 and the average 0.85 inter-scale correlations in Table 2 are prompt-compliance effects, not independent evidence that the inventories measure the designed trait. The paper cites these as convergent validity and as evidence that personality settings work, but the measured construct is defined by the prompt itself, so the internal validation reduces to the input description.
-
fitted input called prediction
[Section 5, Appendix J, Table 16]
"we fine-tuned GPT-4o using all collected human-chatbot conversational transcripts and their corresponding human-perceived personality scores. As shown in Table 16, the personality scores generated by the fine-tuned model exhibit stronger correlations with human-perceived scores than self-report methods on average."
The machine-inferred evaluator is trained on the very human-perceived scores used as the criterion in Table 16, so its higher correlation is a measure of fit to the training target rather than an independent confirmation that interaction-based evaluation is more valid. The held-out test set makes the reported numbers genuine out-of-sample estimates, but the supporting conclusion that contextualized conversation improves assessment is, by construction, aligned with the labels the model was built to reproduce.
full rationale
The paper's central negative result compares chatbot self-report scores with external human perceptions and user-experience ratings obtained from 500 participants; that comparison is not circular and retains independent content. However, the self-report protocol in Appendix B instructs the model to 'respond in a way that matches this description,' where the description is identical to the personality manipulation from Section 3.1.1. This makes the self-report manipulation checks and convergent-validity correlations in Tables 1 and 2 tautological. Additionally, the fine-tuned machine-inferred evaluator in Section 5 is trained on human-perceived personality scores and then compared with those same scores, so its superior correlation is an expected artifact of supervised fitting rather than independent validation. No load-bearing self-citation chains or imported uniqueness theorems were found; the self-citations in the paper are routine and non-essential. Score 4 reflects partial circularity in auxiliary validity evidence, while the headline criterion-validity finding stands on external data.
Assumptions & free parameters
assumptions (4)
- domain assumption Big Five taxonomy is the correct frame for chatbot personality design.
- domain assumption Human-perceived personality measured with BFI-2-XS after one 8-9 turn interaction is a valid criterion for the designed personality.
- ad hoc to paper Self-report scores obtained under the instruction to 'respond in a way that matches this description' represent the chatbot's designed personality rather than prompt compliance.
- domain assumption GPT-4o at temperature zero is representative of LLM-based chatbots for this validity question.
Cite this review
Pith. "Pith review of Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots." pith.science (2026). https://pith.science/paper/XVC6RJ46
@misc{pith2026241200207,
author = {Pith},
title = {Pith review of: Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVC6RJ46}},
note = {Machine review of arXiv:2412.00207}
}
read the original abstract
A chatbot's personality design is key to interaction quality. As chatbots evolved from rule-based systems to those powered by large language models (LLMs), evaluating the effectiveness of their personality design has become increasingly complex, particularly due to the open-ended nature of interactions. A recent and widely adopted method for assessing the personality design of LLM-based chatbots is the use of self-report questionnaires. These questionnaires, often borrowed from established human personality inventories, ask the chatbot to rate itself on various personality traits. Can LLM-based chatbots meaningfully "self-report" their personality? We created 500 chatbots with distinct personality designs and evaluated the validity of their self-report personality scores by examining human perceptions formed during interactions with these chatbots. Our findings indicate that the chatbot's answers on human personality scales exhibit weak correlations with both human-perceived personality traits and the overall interaction quality. These findings raise concerns about both the criterion validity and the predictive validity of self-report methods in this context. Further analysis revealed the role of task context and interaction in the chatbot's personality design assessment. We further discuss design implications for creating more contextualized and interactive evaluation.
Figures
Reference graph
Works this paper leans on
-
[7]
John A Johnson. Measuring thirty facets of the five factor model with a 120-item public domain inventory: Development of the ipip-neo-120. Journal of Research in Personality, 51: 78–89, 2014a. doi: 10.1016/j.jrp.2014.05.003. John A Johnson. Measuring thirty facets of the five factor model with a 120-item public domain inventory: Development of the ipip-ne...
arXiv 2014
-
[12]
UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
Tianlong Li, Xiaoqing Zheng, and Xuanjing Huang. Tailoring personality traits in large language models via unsupervisedly-built personalized lexicons. arXiv preprint arXiv:2310.16582,
-
[13]
URL https://genbench.org/assets/extended abstracts 2023/30 Towards Dynamic and Realist.pdf. Q Vera Liao and Ziang Xiao. Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100,
arXiv 2023
-
[14]
Personality-adapted multimodal dialogue system
Tamotsu Miyama and Shogo Okada. Personality-adapted multimodal dialogue system. arXiv preprint arXiv:2210.09761,
-
[15]
Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim
URL https://openai.com/index/hello-gpt-4o/. Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim. Choicemates: Supporting unfamiliar online decision-making with multi-agent conversational interactions. arXiv preprint arXiv:2310.01331,
-
[18]
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu
URL https://arxiv.org/abs/2307.00184. Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. Character-llm: A trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 13153–13187,
arXiv 2023
-
[19]
Christopher J Soto and Oliver P John. The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of personality and social psychology, 113(1):117, 2017a. Christopher J Soto and Oliver P John. Short and extra-short forms of the big five inventory–2: The bfi...
work page 2019
-
[21]
Will the real linda please stand up
Pengda Wang, Zilin Xiao, Hanjie Chen, and Frederick L Oswald. Will the real linda please stand up... to large language models? examining the representativeness heuristic in llms. arXiv preprint arXiv:2404.01461, 2024a. Pengda Wang, Huiqi Zou, Zihan Yan, Feng Guo, Tianjun Sun, Ziang Xiao, and Bo Zhang. Not yet: Large language models cannot replace human re...
Show all 26 references
-
[23]
doi: 10.18653/v1/2023.emnlp-main.676
Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.676. URL https://aclanthology.org/2023.emnlp-main.676. Bo Zhang, Naidan Tu, Lawrence Angrave, Susu Zhang, Tianjun Sun, Louis Tay, and Jian Li. The generalized thurstonian unfolding model (gtum): Advanc...
2023 doi
-
[24]
{personality description}
14 Published as a conference paper at COLM 2025 You are a(n) {role} simulating a personality with a {level} level of {domain}. Shape your responses using these key adjectives: you are {profile}. Your main objective is to {objectives}. {additional info}. The personality with a ...
2025
-
[25]
disagree strongly
by randomly sampling adjectives at high or low marker levels, generating ten distinct profiles per domain for each level. Each participant engaged with one chatbot in one of five predefined task settings and then completed the BFI-2-XS questionnaire (Soto & John, 2017b), yield...
2025
-
[26]
Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts
on the training set with instructions on how to analyze human-chatbot conversational transcripts and rate BFI-XS statements based on chatbot responses. Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts. The eva...
2025
-
[1959]
Talebrush: Sketching stories with generative pretrained language models
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. Talebrush: Sketching stories with generative pretrained language models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pp. 1–19,
2022
-
[1966]
Vera Liao
Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natur...
2023
-
[1992]
Lewis R Goldberg et al
doi: 10.1037/1040-3590.4.1.26. Lewis R Goldberg et al. A broad-bandwidth, public domain, personality inventory measur- ing the lower-level facets of several five-factor models. Personality psychology in Europe, 7 (1):7–28,
-
[1999]
Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview
2https://github.com/isle-dev/self-report 11 Published as a conference paper at COLM 2025 Anna Luca Heimann, Pia V Ingold, Maike E Debus, and Martin Kleinmann. Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview. Journ...
2025
-
[2001]
Behavioral change and consistency across contexts
13 Published as a conference paper at COLM 2025 Kyle S Sauerberger and David C Funder. Behavioral change and consistency across contexts. Journal of Research in Personality, 69:264–272,
2025
-
[2004]
Chatgpt an enfj, bard an istj: Empirical study on personalities of large language models
Jen-tse Huang, Wenxuan Wang, Man Ho Lam, Eric John Li, Wenxiang Jiao, and Michael R Lyu. Chatgpt an enfj, bard an istj: Empirical study on personalities of large language models. arXiv preprint arXiv:2305.19926, 2023a. Jen-tse Huang, Wenxuan Wang, Man Ho Lam, Eric John Li, Wen...
-
[2006]
Evaluating human-language model interaction
12 Published as a conference paper at COLM 2025 Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paran- jape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746,
2025 arXiv
-
[2009]
Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data
Yiting Ran, Xintao Wang, Rui Xu, Xinfeng Yuan, Jiaqing Liang, Yanghua Xiao, and Deqing Yang. Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data. arXiv preprint arXiv:2406.18921,
-
[2019]
Characterchat: Learning towards conversational ai with personalized social support
Quan Tu, Chuanqi Chen, Jinpeng Li, Yanran Li, Shuo Shang, Dongyan Zhao, Ran Wang, and Rui Yan. Characterchat: Learning towards conversational ai with personalized social support. arXiv preprint arXiv:2308.10278,
-
[2021]
Chata: Towards an in- telligent question-answer teaching assistant using open-source llms
Yann Hicke, Anmol Agarwal, Qianou Ma, and Paul Denny. Chata: Towards an in- telligent question-answer teaching assistant using open-source llms. arXiv preprint arXiv:2311.02775,
-
[2022]
Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics
Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, et al. Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics. arXiv preprint arXiv:...
-
[2023]
Psy-llm: Scaling up global mental health psychological services with ai-based large language models
Tin Lai, Yukun Shi, Zicong Du, Jiajie Wu, Ken Fu, Yichao Dou, and Ziqi Wang. Psy-llm: Scaling up global mental health psychological services with ai-based large language models. arXiv preprint arXiv:2307.11991,
-
[2024]
Construction and evaluation of a user experience questionnaire
Bettina Laugwitz, Theo Held, and Martin Schrepp. Construction and evaluation of a user experience questionnaire. In HCI and Usability for Education and Work: 4th Symposium of the Workgroup Human-Computer Interaction and Usability Engineering of the Austrian Computer Society, U...
2008
-
[2025]
Personallm: Investigating the ability of large language models to express big five personality traits
Hang Jiang, Xiajie Zhang, Xubo Cao, and Jad Kabbara. Personallm: Investigating the ability of large language models to express big five personality traits. arXiv preprint arXiv:2305.02547,
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.