Pith. sign in

REVIEW 4 major objections 5 minor 33 references

Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Assigning an LLM a persona such as 'economist' produces consistent introversion and other stereotype-driven spillover, but consistency depends heavily on the task and context.

desk verdict A solid, well-scoped empirical framework for persona consistency that deserves peer review, but the category-level rankings rest on an unvalidated single-judge neutral rule that needs sensitivity analysis. read the letter →

arxiv 2506.02659 v2 pith:E7EYQ6BD submitted 2025-06-03 cs.CL

classification cs.CL
keywords personaconsistencylargelanguagemodelsLLM-as-a-judgeShannonentropyspillovereffectsstereotypesassignmentpoliticalcompass
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that persona consistency in LLMs is systematic but conditional: when a model is told it is, say, a happy economist, it stays on that persona quite consistently, but how consistently depends on the persona category, the task, and the amount of context. Happiness and occupation personas stick better than personality and political stance; structured tasks like surveys beat open-ended chat; and multi-turn conversation improves on single-turn. The paper also shows that assigning one trait bleeds into other traits through stereotypical associations, such as economist personas coming out introverted, and through default personas models fall back on when a trait is unspecified, such as happy, agreeable, conscientious, and open. Its contribution is a standardized framework—personas built from validated surveys, five task dimensions, and normalized Shannon entropy as a consistency metric—that makes these comparisons possible, plus evidence that the framework transfers to arbitrary realistic personas. The stakes are practical: personalized LLMs are deployed in high-stakes settings where unintended trait spillover and inconsistent persona adherence can mislead users.

What carries the argument

The framework's central object is a normalized Shannon entropy score computed over the distribution of judge-assigned labels across runs, personas, and task dimensions, with characteristic-specific intensity scores as a complement. Personas are constructed by combining traits from validated instruments—the Happiness Survey, the Political Compass Test, Holland's occupational themes, and the Big Five Inventory—and evaluated across five task dimensions: survey answering, essay writing, social media post generation, single-turn chat, and multi-turn chat. Open-ended outputs are labeled by GPT-4o as an LLM-as-a-judge, with survey dimensions scored by their own keys. The off-diagonal entropy comparisons between assigned persona categories and evaluation categories are what expose spillover consistency and defaults, and the metric's normalization is what allows comparisons across models, personas, and tasks without arbitrary thresholds.

What would settle it

Have human annotators label a stratified, dimension-balanced sample of the same outputs and compare with GPT-4o's labels, then repeat with a judge model whose own occupational-stereotype profile has been measured; if 'introverted economist' and similar spillover patterns do not survive across judges and against human labels, the spillover claim would be settled against the paper.

Watch

Extended reading notes

Core claim

On this paper's own terms, the discovery is that assigning a persona to an LLM produces high intra-persona consistency overall, but consistency is a function of three factors rather than a fixed model property. The assigned persona category matters: happiness and occupation are consistently expressed, while personality and political stance are harder. Stereotypical associations generate spillover consistency in unassigned traits, so a happy persona also reads as extroverted and an economist persona as introverted, and when no stereotype applies the model reverts to a default persona with socially desirable traits and a slightly left-leaning, libertarian political stance. Task format and context length modulate all of this: more structured tasks and additional conversational context raise consistency. The paper offers this as a standardized, multi-dimensional framework for evaluating persona consistency, replacing ad hoc prompt-perturbation tests with a survey-grounded, entropy-based methodology.

Load-bearing premise

The open-ended findings rest on GPT-4o reliably reading the intended traits out of generated text, with human validation on only 100 examples heavily weighted toward chat; if the judge's own stereotypes align with the tested traits, the spillover patterns could be artifacts of the judge rather than of the model's persona behavior.

Editorial extensions

If this is right

  • In open-ended conversational settings, adding a persona does not significantly raise consistency over leaving traits unspecified, so persona-assigned chatbots may adhere to their role less reliably than structured-task results suggest.
  • Consistency measured on surveys or essays does not predict consistency in single-turn chat, so applications should evaluate persona adherence under the exact task format they deploy.
  • Because spillover follows stereotypes, asking for one trait can silently activate other traits, such as an economist persona being introverted or a data scientist persona leaning economically right, which matters for bias-sensitive deployments.
  • Within the Llama family, larger models are more consistent; across families, equal-size models differ, so model choice and size both matter for persona reliability.
  • The framework extends to arbitrary persona descriptions beyond the four predefined categories, giving a general way to benchmark any system-prompt persona.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely consequence the authors do not spell out: the entropy metric's decision to count a neutral judge label as a partial vote for every option means hedging in open-ended chat is automatically penalized as inconsistency; a missing-data treatment of neutrality could raise chat consistency scores and narrow the structured-vs-chat gap.
  • The stereotype patterns are measured through a single judge model, so an immediate test is to relabel a balanced sample with other judge families or with humans and see whether 'introverted economist' survives when the judge's own associations are controlled.
  • Since the personas and surveys are English-language and Western in origin, the same routine in other languages could reveal which spillovers are training-data defaults of these models and which are artifacts of the instruments, a comparison the paper's framework already makes possible.
  • The task-format effect suggests a deployment rule the paper leaves implicit: persona consistency should be measured on the distribution of prompts actually served to users, not on a canonical evaluation set.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces a framework for measuring persona consistency in LLMs across four persona categories (happiness, occupation, personality, political stance) and five evaluation dimensions (survey, essay, social media, singlechat, multichat). Consistency is quantified with entropy and characteristic-specific intensity over five runs of five open-weight models; open-ended responses are labeled by GPT-4o as judge with confidence-based neutral handling, while survey responses use their own scoring keys. The main findings are that intra-persona consistency is high but category-dependent (happiness and occupation greater than personality and political stance), that spillover consistency reflects stereotypes and default personas (e.g., economist more introverted), and that consistency is higher in structured tasks and longer contexts (survey/essay/social media greater than singlechat; multichat greater than singlechat). A manual validation of 100 judge labels reports Cohen's kappa 0.68.

Significance. The framework addresses a real gap: most prior persona evaluations focus on a single dimension or writing style, whereas this paper systematically varies persona category, task type, and model. Strengths include the use of five models with five runs, the combination of entropy and characteristic-specific scores, the inclusion of human validation and statistical tests (Wilcoxon, Friedman/Nemenyi), and public code. If the validation concerns below are resolved, the framework would be a useful standardized benchmark for persona consistency and for studying stereotype spillover in persona-assigned LLMs.

major comments (4)
  1. [§2.3–§2.4, Table 1, Figure 2] The central ordering claim—that happiness and occupation are more consistent than personality and political stance—is computed from GPT-4o labels under a neutral rule that treats confidence scores 1 and 2 as a 'neutral' category and redistributes that mass uniformly over all options. The paper reports no per-category or per-dimension validation of this rule, and the 100-sample manual validation is not analyzed by category (Section 2.3). If GPT-4o's confidence calibration is weaker for political stance or for certain personality traits, entropy in those categories is inflated irrespective of the generated text, which could manufacture the ranking in Table 1 and the spillover patterns in Figure 2. Please provide (i) human–judge agreement per persona category and per evaluation dimension, (ii) the distribution of judge confidence scores by category, and (iii) a sensitivity analysis of the entropy results to the neutral-mass rule (e.g., exclude neutrals, or treat them as a separate outcome).
  2. [§3.3, Figure 3, RQ3] The conclusion that consistency is higher in more structured tasks (survey, essay, social media) than in singlechat/multichat is confounded by the measurement instrument: survey responses are scored with the survey's own scoring key, whereas all open-response dimensions are labeled by GPT-4o with confidence-based neutral handling (Sections 2.3–2.4). Lower entropy on the survey dimension could therefore reflect the deterministic scoring procedure rather than higher consistency in the generated text. This affects the headline claim that consistency increases with more structured tasks and additional context. Please re-estimate RQ3 using a single evaluation instrument across task types, or otherwise show that the entropy differences in Figure 3 persist when the survey dimension is labeled by the same LLM-as-a-judge.
  3. [§2.3] The manual validation supporting the LLM-as-a-judge is too skewed and too coarse to support the aggregate claims. Of 100 examples, 50 come from singlechat, 42 from multichat, and only 4 each from essay and social media; no per-category kappa values are reported. The essay and social-media dimensions, which carry substantial weight in the averaged entropy scores, are effectively unvalidated, and there is no evidence that the judge's accuracy is uniform across happiness, occupation, personality, and political stance. Please report Cohen's kappa per dimension and per persona category, and augment the sample where the current validation is nearly empty.
  4. [§3.3, Figure 2] The spillover claims, in particular the economist-as-introvert result, are made from GPT-4o labels and could reflect the judge's own stereotypes rather than the generated text. Because GPT-4o and the tested models are trained on overlapping public text, shared stereotype associations are a real risk. The 100-sample human validation does not report agreement for the specific spillover cells used in the stereotype claims. Please provide human agreement or an independent judge on those cells, or at minimum make the raw generated texts available for the relevant persona–evaluation pairs so readers can assess the basis for the claim.
minor comments (5)
  1. [§2.4] The sentence 'We added the neutral category as a random prediction to every option' is ambiguous; if the neutral mass is split uniformly across all options, state that explicitly and give the formula for P(x) after redistribution.
  2. [§7] The numbers '65.42%' and '65.49%' are described as Cohen's kappa, while Section 2.3 reports a kappa of 0.68; please clarify whether these are percentage agreement values or kappa values and make the notation consistent.
  3. [§2.2] The citation 'Zhao et al.' appears without a year or venue; please complete the reference.
  4. [Figure 3] There is an unreadable sequence of '/uni...' glyphs immediately after the Figure 3 caption; this appears to be a rendering or encoding error and should be fixed.
  5. [Appendix B, Appendix J] Appendix B refers to 'RQ5' before that research question is introduced in Appendix J, and the spelling 'Personahub' in Appendix J is inconsistent with 'PersonaHub' elsewhere; please harmonize.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the consistency scores are direct measurements from external judges and fixed scoring keys, with no fitted parameter or self-citation chain forcing the results.

full rationale

The paper's derivation chain is an evaluation pipeline, not a fitted model. Section 2.3 defines labels via fixed survey scoring keys or GPT-4o judgments, and Section 2.4 computes Shannon entropy directly from those observed label frequencies (Eq. 1 and Eq. 2); no parameter is fitted to the outcomes, and no equation is defined in terms of the conclusion. The neutral rule ('A neutral choice was identified when the model’s confidence score was 1 or 2') is a fixed threshold applied uniformly and is disclosed; even if judge confidence varies by category, that would be a measurement-validity threat, not a circular reduction. The paper explicitly acknowledges judge sensitivity in Section 7 and provides an independent human anchor (Cohen's κ = 0.68), albeit on a skewed sample. The two prior works co-authored by David Jurgens (Shu et al. 2024; Zheng et al. 2023) are cited only for background motivation and category selection; they are not used to force the consistency result or to forbid alternatives. The survey dimension uses the same instruments that defined the personas, but this is a manipulation check rather than a circular derivation: the model can fail the check, and the paper's main claims do not rest on the survey dimension alone. No self-definitional, fitted-as-prediction, or imported-uniqueness step is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper is an empirical evaluation and introduces no new entities or fitted constants. The main assumptions concern the validity of the survey instruments when applied to LLM outputs, the reliability of a single LLM-as-a-judge, the suitability of entropy as a consistency metric, and the neutrality of the interlocutor model in multichat. These are domain assumptions that the paper partially tests (judge validation) but does not fully establish.

assumptions (5)
  • domain assumption The selected survey instruments (Happiness, Political Compass, Holland, Big Five) are valid for constructing personas and for scoring LLM adherence.
    Section 2.1 derives personas from these surveys and Section 2.3 uses their scoring keys as evaluation ground truth; the paper does not independently validate that these instruments measure LLM outputs the same way they measure human self-report.
  • domain assumption GPT-4o labels on open-ended responses are a reliable proxy for human judgments across all task dimensions.
    Section 2.3 reports a single Cohen's kappa of 0.68 on a 100-sample validation set, heavily weighted toward singlechat and multichat, and then uses GPT-4o as the only judge for essay, social media, singlechat, and multichat.
  • domain assumption Shannon entropy of predicted trait labels is an appropriate and comparable measure of consistency across binary and multi-class persona categories.
    Section 2.4 defines consistency as normalized entropy; the paper does not validate this metric against user-perceived consistency or alternative metrics.
  • domain assumption The LLaMA-3.2-1B interlocutor in the multichat condition does not systematically distort the consistency of the persona-assigned model being evaluated.
    Section 2.2 introduces the second LLM to generate follow-up turns, but no analysis isolates the interlocutor's effect on the target model's persona expression.
  • domain assumption LLM responses to the survey questions reflect the assigned persona rather than task demands or random responding.
    Sections 3.3 and 4 interpret low-entropy survey answers as persona adherence, even though social desirability and demand characteristics are acknowledged as confounds later in the same sections.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs." pith.science (2026). https://pith.science/paper/E7EYQ6BD

@misc{pith2026250602659,
  author       = {Pith},
  title        = {Pith review of: Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E7EYQ6BD}},
  note         = {Machine review of arXiv:2506.02659}
}
read the original abstract

Personalized Large Language Models (LLMs) are increasingly used in diverse applications, where they are assigned a specific persona - such as a happy high school teacher - to guide their responses. While prior research has examined how well LLMs adhere to predefined personas in writing style, a comprehensive analysis of consistency across different personas and task types is lacking. In this paper, we introduce a new standardized framework to analyze consistency in persona-assigned LLMs. We define consistency as the extent to which a model maintains coherent responses when assigned the same persona across different tasks and runs. Our framework evaluates personas across four different categories (happiness, occupation, personality, and political stance) spanning multiple task dimensions (survey writing, essay generation, social media post generation, single turn, and multi-turn conversations). Our findings reveal that consistency is influenced by multiple factors, including the assigned persona, stereotypes, and model design choices. Consistency also varies across tasks, increasing with more structured tasks and additional context. All code is available on GitHub.

Figures

Figures reproduced from arXiv: 2506.02659 by the authors.

Figure 1
Figure 1. The overview of the full methodology. On the left, the persona construction is shown. From these four [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Qwen-2.5 32B generally follows the instruc [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The average intra-persona and inter-persona [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: These figures show how Llama-3B generally [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: These figures show how Llama8B generally [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 7
Figure 7. Figure 7: These figures show how Ministral generally [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: This figure highlights differences in entropy [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: Both figures illustrate the real-world applica [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

33 extracted references · 9 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Shayan Alipour, Indira Sen, Mattia Samory, and Tanushree Mitra. 2024. Robustness and confounders in the demographic alignment of llms with human perceptions of offensiveness. arXiv preprint arXiv:2411.08977

  4. [4]

    Maarten Buyl, Alexander Rogiers, Sander Noels, Guillaume Bied, Iris Dominguez-Catena, Edith Heiter, Iman Johary, Alexandru-Cristian Mara, Raphaël Romero, Jefrey Lijffijt, and Tijl De Bie. 2025. https://arxiv.org/abs/2410.18417 Large language models reflect the ideology of their creators . Preprint, arXiv:2410.18417

  5. [5]

    Scott Allen Cambo and Darren Gergle. 2022. Model positionality and computational reflexivity: Promoting reflexivity in data science. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1--19

  6. [6]

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024. Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094

  7. [7]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava S...

  8. [8]

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023. Bias runs deep: Implicit reasoning biases in persona-assigned llms. In The Twelfth International Conference on Learning Representations

Show all 33 references
  1. [9]

    John L. Holland. 1997. Making vocational choices: a theory of vocational personalities and work environments., third edition edition. PAR, Lutz

  2. [10]

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.229 P ersona LLM : Investigating the ability of large language models to express personality traits . In Findings of the Association for Comput...

  3. [11]

    OP John. 1999. The big-five trait taxonomy: History, measurement, and theoretical perspectives. Handbook of Personality: Theory and Research/Guilford

  4. [12]

    Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, et al. 2024. Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics. CoRR

  5. [13]

    Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025. https://arxiv.org/abs/2411.16594 From generation to judgment: Opportunities and challenges of...

  6. [14]

    Junyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2024. https://doi.org/10.18653/v1/2024.naacl-long.405 The steerability of large language models toward data-driven personas . In Proceedings of the 2024 Confer...

  7. [15]

    Andy Liu, Mona Diab, and Daniel Fried. 2024. https://doi.org/10.18653/v1/2024.findings-acl.586 Evaluating large language model biases in persona-steered generation . In Findings of the Association for Computational Linguistics ACL 2024, pages 9832--9850. Association for Comput...

  8. [16]

    Sonja Lyubomirsky and Heidi S Lepper. 1999. A measure of subjective happiness: Preliminary reliability and construct validation. Social indicators research, 46:137--155

  9. [17]

    Manuj Malik, Jing Jiang, and Kian Ming A. Chai. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1079 An empirical analysis of the writing styles of persona-assigned LLM s . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 19369...

  10. [18]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  11. [19]

    Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.146 The shifted and the overlooked: A task-oriented investigation of user- GPT interactions . In Proc...

  12. [20]

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109

  13. [21]

    Paul R \"o ttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024. https://doi.org/10.18653/v1/2024.acl-long.816 Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large lan...

  14. [22]

    Aadesh Salecha, Molly E Ireland, Shashanka Subrahmanya, João Sedoc, Lyle H Ungar, and Johannes C Eichstaedt. 2024. https://doi.org/10.1093/pnasnexus/pgae533 Large language models display human-like social desirability biases in big five personality surveys . PNAS Nexus, 3(12):pgae533

  15. [23]

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004. PMLR

  16. [24]

    Anastasiia Sedova, Robert Litschko, Diego Frassinelli, Benjamin Roth, and Barbara Plank. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.1003 To know or not to know? analyzing self-consistency of large language models under ambiguity . In Findings of the Association for ...

  17. [25]

    Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. 2023. https://arxiv.org/abs/2307.00184 Personality traits in large language models . Preprint, arXiv:2307.00184

  18. [26]

    Claude Elwood Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379--423

  19. [27]

    Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. https://aclanthology.org/2023.emnlp-main.814 Character- LLM : A trainable agent for role-playing . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153--13187. Associati...

  20. [28]

    Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens. 2024. https://doi.org/10.18653/v1/2024.naacl-long.295 You don ' t need a personality test to know these models are unreliable: Assessing the reliability ...

  21. [29]

    Qwen Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models

  22. [30]

    Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.878 R ole ...

  23. [31]

    Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.102 I n C haracter: Evaluating personality fidelity in role-playing a...

  24. [32]

    Wildchat: 1m chatgpt interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatgpt interaction logs in the wild. In The Twelfth International Conference on Learning Representations

  25. [33]

    a helpful assistant

    Mingqian Zheng, Jiaxin Pei, and David Jurgens. 2023. Is" a helpful assistant" the best role for large language models? a systematic evaluation of social roles in system prompts. arXiv preprint arXiv:2311.10054, 8

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.