REVIEW 4 major objections 5 minor 33 references
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Assigning an LLM a persona such as 'economist' produces consistent introversion and other stereotype-driven spillover, but consistency depends heavily on the task and context.
desk verdict A solid, well-scoped empirical framework for persona consistency that deserves peer review, but the category-level rankings rest on an unvalidated single-judge neutral rule that needs sensitivity analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework's central object is a normalized Shannon entropy score computed over the distribution of judge-assigned labels across runs, personas, and task dimensions, with characteristic-specific intensity scores as a complement. Personas are constructed by combining traits from validated instruments—the Happiness Survey, the Political Compass Test, Holland's occupational themes, and the Big Five Inventory—and evaluated across five task dimensions: survey answering, essay writing, social media post generation, single-turn chat, and multi-turn chat. Open-ended outputs are labeled by GPT-4o as an LLM-as-a-judge, with survey dimensions scored by their own keys. The off-diagonal entropy comparisons between assigned persona categories and evaluation categories are what expose spillover consistency and defaults, and the metric's normalization is what allows comparisons across models, personas, and tasks without arbitrary thresholds.
What would settle it
Have human annotators label a stratified, dimension-balanced sample of the same outputs and compare with GPT-4o's labels, then repeat with a judge model whose own occupational-stereotype profile has been measured; if 'introverted economist' and similar spillover patterns do not survive across judges and against human labels, the spillover claim would be settled against the paper.
Extended reading notes
Core claim
On this paper's own terms, the discovery is that assigning a persona to an LLM produces high intra-persona consistency overall, but consistency is a function of three factors rather than a fixed model property. The assigned persona category matters: happiness and occupation are consistently expressed, while personality and political stance are harder. Stereotypical associations generate spillover consistency in unassigned traits, so a happy persona also reads as extroverted and an economist persona as introverted, and when no stereotype applies the model reverts to a default persona with socially desirable traits and a slightly left-leaning, libertarian political stance. Task format and context length modulate all of this: more structured tasks and additional conversational context raise consistency. The paper offers this as a standardized, multi-dimensional framework for evaluating persona consistency, replacing ad hoc prompt-perturbation tests with a survey-grounded, entropy-based methodology.
Load-bearing premise
The open-ended findings rest on GPT-4o reliably reading the intended traits out of generated text, with human validation on only 100 examples heavily weighted toward chat; if the judge's own stereotypes align with the tested traits, the spillover patterns could be artifacts of the judge rather than of the model's persona behavior.
Editorial extensions
If this is right
- In open-ended conversational settings, adding a persona does not significantly raise consistency over leaving traits unspecified, so persona-assigned chatbots may adhere to their role less reliably than structured-task results suggest.
- Consistency measured on surveys or essays does not predict consistency in single-turn chat, so applications should evaluate persona adherence under the exact task format they deploy.
- Because spillover follows stereotypes, asking for one trait can silently activate other traits, such as an economist persona being introverted or a data scientist persona leaning economically right, which matters for bias-sensitive deployments.
- Within the Llama family, larger models are more consistent; across families, equal-size models differ, so model choice and size both matter for persona reliability.
- The framework extends to arbitrary persona descriptions beyond the four predefined categories, giving a general way to benchmark any system-prompt persona.
Reading between the lines
- A likely consequence the authors do not spell out: the entropy metric's decision to count a neutral judge label as a partial vote for every option means hedging in open-ended chat is automatically penalized as inconsistency; a missing-data treatment of neutrality could raise chat consistency scores and narrow the structured-vs-chat gap.
- The stereotype patterns are measured through a single judge model, so an immediate test is to relabel a balanced sample with other judge families or with humans and see whether 'introverted economist' survives when the judge's own associations are controlled.
- Since the personas and surveys are English-language and Western in origin, the same routine in other languages could reveal which spillovers are training-data defaults of these models and which are artifacts of the instruments, a comparison the paper's framework already makes possible.
- The task-format effect suggests a deployment rule the paper leaves implicit: persona consistency should be measured on the distribution of prompts actually served to users, not on a canonical evaluation set.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces a framework for measuring persona consistency in LLMs across four persona categories (happiness, occupation, personality, political stance) and five evaluation dimensions (survey, essay, social media, singlechat, multichat). Consistency is quantified with entropy and characteristic-specific intensity over five runs of five open-weight models; open-ended responses are labeled by GPT-4o as judge with confidence-based neutral handling, while survey responses use their own scoring keys. The main findings are that intra-persona consistency is high but category-dependent (happiness and occupation greater than personality and political stance), that spillover consistency reflects stereotypes and default personas (e.g., economist more introverted), and that consistency is higher in structured tasks and longer contexts (survey/essay/social media greater than singlechat; multichat greater than singlechat). A manual validation of 100 judge labels reports Cohen's kappa 0.68.
Significance. The framework addresses a real gap: most prior persona evaluations focus on a single dimension or writing style, whereas this paper systematically varies persona category, task type, and model. Strengths include the use of five models with five runs, the combination of entropy and characteristic-specific scores, the inclusion of human validation and statistical tests (Wilcoxon, Friedman/Nemenyi), and public code. If the validation concerns below are resolved, the framework would be a useful standardized benchmark for persona consistency and for studying stereotype spillover in persona-assigned LLMs.
major comments (4)
- [§2.3–§2.4, Table 1, Figure 2] The central ordering claim—that happiness and occupation are more consistent than personality and political stance—is computed from GPT-4o labels under a neutral rule that treats confidence scores 1 and 2 as a 'neutral' category and redistributes that mass uniformly over all options. The paper reports no per-category or per-dimension validation of this rule, and the 100-sample manual validation is not analyzed by category (Section 2.3). If GPT-4o's confidence calibration is weaker for political stance or for certain personality traits, entropy in those categories is inflated irrespective of the generated text, which could manufacture the ranking in Table 1 and the spillover patterns in Figure 2. Please provide (i) human–judge agreement per persona category and per evaluation dimension, (ii) the distribution of judge confidence scores by category, and (iii) a sensitivity analysis of the entropy results to the neutral-mass rule (e.g., exclude neutrals, or treat them as a separate outcome).
- [§3.3, Figure 3, RQ3] The conclusion that consistency is higher in more structured tasks (survey, essay, social media) than in singlechat/multichat is confounded by the measurement instrument: survey responses are scored with the survey's own scoring key, whereas all open-response dimensions are labeled by GPT-4o with confidence-based neutral handling (Sections 2.3–2.4). Lower entropy on the survey dimension could therefore reflect the deterministic scoring procedure rather than higher consistency in the generated text. This affects the headline claim that consistency increases with more structured tasks and additional context. Please re-estimate RQ3 using a single evaluation instrument across task types, or otherwise show that the entropy differences in Figure 3 persist when the survey dimension is labeled by the same LLM-as-a-judge.
- [§2.3] The manual validation supporting the LLM-as-a-judge is too skewed and too coarse to support the aggregate claims. Of 100 examples, 50 come from singlechat, 42 from multichat, and only 4 each from essay and social media; no per-category kappa values are reported. The essay and social-media dimensions, which carry substantial weight in the averaged entropy scores, are effectively unvalidated, and there is no evidence that the judge's accuracy is uniform across happiness, occupation, personality, and political stance. Please report Cohen's kappa per dimension and per persona category, and augment the sample where the current validation is nearly empty.
- [§3.3, Figure 2] The spillover claims, in particular the economist-as-introvert result, are made from GPT-4o labels and could reflect the judge's own stereotypes rather than the generated text. Because GPT-4o and the tested models are trained on overlapping public text, shared stereotype associations are a real risk. The 100-sample human validation does not report agreement for the specific spillover cells used in the stereotype claims. Please provide human agreement or an independent judge on those cells, or at minimum make the raw generated texts available for the relevant persona–evaluation pairs so readers can assess the basis for the claim.
minor comments (5)
- [§2.4] The sentence 'We added the neutral category as a random prediction to every option' is ambiguous; if the neutral mass is split uniformly across all options, state that explicitly and give the formula for P(x) after redistribution.
- [§7] The numbers '65.42%' and '65.49%' are described as Cohen's kappa, while Section 2.3 reports a kappa of 0.68; please clarify whether these are percentage agreement values or kappa values and make the notation consistent.
- [§2.2] The citation 'Zhao et al.' appears without a year or venue; please complete the reference.
- [Figure 3] There is an unreadable sequence of '/uni...' glyphs immediately after the Figure 3 caption; this appears to be a rendering or encoding error and should be fixed.
- [Appendix B, Appendix J] Appendix B refers to 'RQ5' before that research question is introduced in Appendix J, and the spelling 'Personahub' in Appendix J is inconsistent with 'PersonaHub' elsewhere; please harmonize.
Circularity Check
No significant circularity: the consistency scores are direct measurements from external judges and fixed scoring keys, with no fitted parameter or self-citation chain forcing the results.
full rationale
The paper's derivation chain is an evaluation pipeline, not a fitted model. Section 2.3 defines labels via fixed survey scoring keys or GPT-4o judgments, and Section 2.4 computes Shannon entropy directly from those observed label frequencies (Eq. 1 and Eq. 2); no parameter is fitted to the outcomes, and no equation is defined in terms of the conclusion. The neutral rule ('A neutral choice was identified when the model’s confidence score was 1 or 2') is a fixed threshold applied uniformly and is disclosed; even if judge confidence varies by category, that would be a measurement-validity threat, not a circular reduction. The paper explicitly acknowledges judge sensitivity in Section 7 and provides an independent human anchor (Cohen's κ = 0.68), albeit on a skewed sample. The two prior works co-authored by David Jurgens (Shu et al. 2024; Zheng et al. 2023) are cited only for background motivation and category selection; they are not used to force the consistency result or to forbid alternatives. The survey dimension uses the same instruments that defined the personas, but this is a manipulation check rather than a circular derivation: the model can fail the check, and the paper's main claims do not rest on the survey dimension alone. No self-definitional, fitted-as-prediction, or imported-uniqueness step is present.
Assumptions & free parameters
assumptions (5)
- domain assumption The selected survey instruments (Happiness, Political Compass, Holland, Big Five) are valid for constructing personas and for scoring LLM adherence.
- domain assumption GPT-4o labels on open-ended responses are a reliable proxy for human judgments across all task dimensions.
- domain assumption Shannon entropy of predicted trait labels is an appropriate and comparable measure of consistency across binary and multi-class persona categories.
- domain assumption The LLaMA-3.2-1B interlocutor in the multichat condition does not systematically distort the consistency of the persona-assigned model being evaluated.
- domain assumption LLM responses to the survey questions reflect the assigned persona rather than task demands or random responding.
Cite this review
Pith. "Pith review of Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs." pith.science (2026). https://pith.science/paper/E7EYQ6BD
@misc{pith2026250602659,
author = {Pith},
title = {Pith review of: Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/E7EYQ6BD}},
note = {Machine review of arXiv:2506.02659}
}
read the original abstract
Personalized Large Language Models (LLMs) are increasingly used in diverse applications, where they are assigned a specific persona - such as a happy high school teacher - to guide their responses. While prior research has examined how well LLMs adhere to predefined personas in writing style, a comprehensive analysis of consistency across different personas and task types is lacking. In this paper, we introduce a new standardized framework to analyze consistency in persona-assigned LLMs. We define consistency as the extent to which a model maintains coherent responses when assigned the same persona across different tasks and runs. Our framework evaluates personas across four different categories (happiness, occupation, personality, and political stance) spanning multiple task dimensions (survey writing, essay generation, social media post generation, single turn, and multi-turn conversations). Our findings reveal that consistency is influenced by multiple factors, including the assigned persona, stereotypes, and model design choices. Consistency also varies across tasks, increasing with more structured tasks and additional context. All code is available on GitHub.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Shayan Alipour, Indira Sen, Mattia Samory, and Tanushree Mitra. 2024. Robustness and confounders in the demographic alignment of llms with human perceptions of offensiveness. arXiv preprint arXiv:2411.08977
arXiv 2024
-
[4]
Maarten Buyl, Alexander Rogiers, Sander Noels, Guillaume Bied, Iris Dominguez-Catena, Edith Heiter, Iman Johary, Alexandru-Cristian Mara, Raphaël Romero, Jefrey Lijffijt, and Tijl De Bie. 2025. https://arxiv.org/abs/2410.18417 Large language models reflect the ideology of their creators . Preprint, arXiv:2410.18417
arXiv 2025
-
[5]
Scott Allen Cambo and Darren Gergle. 2022. Model positionality and computational reflexivity: Promoting reflexivity in data science. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1--19
work page 2022
-
[6]
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024. Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094
arXiv 2024
-
[7]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava S...
arXiv 2024
-
[8]
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023. Bias runs deep: Implicit reasoning biases in persona-assigned llms. In The Twelfth International Conference on Learning Representations
work page 2023
Show all 33 references
-
[9]
John L. Holland. 1997. Making vocational choices: a theory of vocational personalities and work environments., third edition edition. PAR, Lutz
1997
-
[10]
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.229 P ersona LLM : Investigating the ability of large language models to express personality traits . In Findings of the Association for Comput...
2024 doi
-
[11]
OP John. 1999. The big-five trait taxonomy: History, measurement, and theoretical perspectives. Handbook of Personality: Theory and Research/Guilford
1999
-
[12]
Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, et al. 2024. Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics. CoRR
2024
-
[13]
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025. https://arxiv.org/abs/2411.16594 From generation to judgment: Opportunities and challenges of...
2025
-
[14]
Junyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2024. https://doi.org/10.18653/v1/2024.naacl-long.405 The steerability of large language models toward data-driven personas . In Proceedings of the 2024 Confer...
2024 doi
-
[15]
Andy Liu, Mona Diab, and Daniel Fried. 2024. https://doi.org/10.18653/v1/2024.findings-acl.586 Evaluating large language model biases in persona-steered generation . In Findings of the Association for Computational Linguistics ACL 2024, pages 9832--9850. Association for Comput...
2024 doi
-
[16]
Sonja Lyubomirsky and Heidi S Lepper. 1999. A measure of subjective happiness: Preliminary reliability and construct validation. Social indicators research, 46:137--155
1999
-
[17]
Manuj Malik, Jing Jiang, and Kian Ming A. Chai. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1079 An empirical analysis of the writing styles of persona-assigned LLM s . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 19369...
2024 doi
-
[18]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[19]
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.146 The shifted and the overlooked: A task-oriented investigation of user- GPT interactions . In Proc...
2023 doi
-
[20]
Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109
2024 arXiv
-
[21]
Paul R \"o ttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024. https://doi.org/10.18653/v1/2024.acl-long.816 Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large lan...
2024 doi
-
[22]
Aadesh Salecha, Molly E Ireland, Shashanka Subrahmanya, João Sedoc, Lyle H Ungar, and Johannes C Eichstaedt. 2024. https://doi.org/10.1093/pnasnexus/pgae533 Large language models display human-like social desirability biases in big five personality surveys . PNAS Nexus, 3(12):pgae533
2024 doi
-
[23]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004. PMLR
2023
-
[24]
Anastasiia Sedova, Robert Litschko, Diego Frassinelli, Benjamin Roth, and Barbara Plank. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.1003 To know or not to know? analyzing self-consistency of large language models under ambiguity . In Findings of the Association for ...
2024 doi
-
[25]
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. 2023. https://arxiv.org/abs/2307.00184 Personality traits in large language models . Preprint, arXiv:2307.00184
2023 arXiv
-
[26]
Claude Elwood Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379--423
1948
-
[27]
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. https://aclanthology.org/2023.emnlp-main.814 Character- LLM : A trainable agent for role-playing . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153--13187. Associati...
2023
-
[28]
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens. 2024. https://doi.org/10.18653/v1/2024.naacl-long.295 You don ' t need a personality test to know these models are unreliable: Assessing the reliability ...
2024 doi
-
[29]
Qwen Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models
2024
-
[30]
Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.878 R ole ...
2024 doi
-
[31]
Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.102 I n C haracter: Evaluating personality fidelity in role-playing a...
2024 doi
-
[32]
Wildchat: 1m chatgpt interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatgpt interaction logs in the wild. In The Twelfth International Conference on Learning Representations
-
[33]
a helpful assistant
Mingqian Zheng, Jiaxin Pei, and David Jurgens. 2023. Is" a helpful assistant" the best role for large language models? a systematic evaluation of social roles in system prompts. arXiv preprint arXiv:2311.10054, 8
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.