Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues

T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LLM agents prompted with Big Five personality traits produce negotiation behavior that matches established personality-psychology predictions, the paper argues.

desk verdict Useful framework, solid directional findings, but the causal and evidential claims outrun the evaluation design. read the letter →

arxiv 2506.15928 v3 pith:2ODSS3JV submitted 2025-06-19 cs.AI cs.CLcs.HC

classification cs.AIcs.CLcs.HC
keywords LLMsimulationBigFivepersonalitynegotiationcausaldiscoverySotopiahuman-AIteaminglexicalanalysisagenticAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that prompting large language model agents with Big Five personality descriptions produces negotiation behavior that tracks established personality psychology: higher Agreeableness and Extraversion reliably shift believability, goal achievement, and knowledge acquisition, while Neuroticism moves outcomes in the opposite direction. The claim is established through two sets of Sotopia simulations—buyer-seller price bargaining and a human-AI job negotiation—analyzed with causal discovery methods that estimate the effect of trait levels on outcomes. The authors also show that lexical markers of empathy, moral foundations, and connotative language shift with trait levels in ways consistent with the human literature. The point of the exercise is practical: if prompt-based personality can steer agent behavior this reliably, then mission-critical AI systems can be stress-tested against diverse operator personalities before deployment.

What carries the argument

The machinery is the Sotopia simulation testbed (an LLM-based framework in which two agents converse in pursuit of private social goals) combined with prompt-based trait manipulations drawn from Big Five Inventory items. Outcomes are scored by Sotopia-Eval dimensions such as believability, goal achievement, and knowledge acquisition, and by a suite of lexical analytics for empathy, moral foundations, sentiment, toxicity, connotation frames, and subjectivity. Causal discovery—structural learning via CausalNex followed by average treatment effect estimation with Causal Forests—is what converts the simulated dialogue corpus into the paper's causal claims about trait levels. The same pipeline is applied to Experiment 2 with the addition of AI-agent trait prompts (transparency, competence, adaptability) and post-interaction questionnaires.

What would settle it

Take a random sample of the generated negotiation transcripts and have independent human raters score them on believability, goal achievement, and empathy without knowing which trait prompt produced them. If the human ratings do not show the same Agreeableness and Extraversion differences reported by the LLM-based Sotopia-Eval and lexical scores, the central claim that prompt-based traits produce theory-consistent simulated behavior would fail. A quicker check is to re-score the same transcripts with a different LLM evaluator and see whether the trait effects survive the change.

Watch

Extended reading notes

Core claim

The central claim is that Big Five personality traits, implemented as prompt text derived from BFI questionnaire items, are a valid and controllable lever on LLM-simulated social behavior. Experiment 1 uses 8,686 price-bargaining transcripts between gpt-4o-mini agents and causal discovery (CausalNex DAGs plus Causal Forest average treatment effects) to show that Agreeableness and Extraversion significantly influence Sotopia-Eval scores for believability, goal achievement, and knowledge acquisition, with Neuroticism associated with worse outcomes. Experiment 2 extends the same logic to human-AI negotiation, where simulated human candidates' Agreeableness and Extraversion dominate AI system transparency, competence, and adaptability in shaping questionnaire and lexical measures; AI traits mainly affect conversational balance (transactivity, verbal equity). The authors conclude that LLM-driven social simulation can serve as a valid platform for studying personality-driven negotiation dynamics and for pre-deployment testing of agentic AI.

Load-bearing premise

The scores that measure whether a simulated negotiation was believable, successful, or empathic come from LLM-based tools reading LLM-generated dialogue, with no human raters checking that those scores match real human judgments.

Editorial extensions

If this is right

  • If prompt-based Big Five manipulations work as claimed, LLM simulations become a scalable substitute for human-subject experiments when probing personality effects on negotiation.
  • AI agents deployed in high-stakes settings should be evaluated per operator personality profile, since the Agreeableness and Extraversion of the human side dominate the interaction measures.
  • The dominance of personality over AI characteristics implies that training and design should prioritize personality-aware communication strategies over purely technical AI enhancements.
  • The framework offers a repeatable pre-deployment test: run the agent across a spectrum of simulated personalities and inspect the causal effects before fielding it.
  • The lexical measures (empathy, moral foundations, connotation framing) provide an actionable diagnostic for when an agent's dialogue drifts from expected trait-consistent behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not test whether the LLM evaluator's scores agree with human judgments; a natural next step is a human-rater study on the same transcripts, which could confirm or overturn the validity of the Sotopia-Eval and lexical measures.
  • Because both the actors and the evaluators are LLMs, part of the observed trait 'effects' could come from prompt phrase matching rather than genuine behavioral simulation; comparing across different evaluator models would probe this.
  • The causal-discovery framing implies intervention on trait levels in the prompt, but the prompt is the only intervention; the paper's ATEs should be read as effects of prompt text on output text, not of human personality on negotiation outcomes.
  • A direct extension would be to run the same two scenarios with a different base LLM (e.g., a smaller open-weight model) and see whether the trait-effect patterns replicate; if they do not, the framework's generality is limited.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper presents an evaluation framework for LLM-simulated negotiations. Experiment 1 manipulates Big Five personality prompts in a price-bargaining dialogue between two gpt-4o-mini agents and reports scenario-based (Sotopia-Eval), lexical, and causal analyses of trait effects. Experiment 2 manipulates AI hiring-manager transparency, competence, and adaptability alongside simulated human Agreeableness and Extraversion in a job-negotiation scenario, adding questionnaire measures and transactivity/verbal-equity outcomes. The authors conclude that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition, that personality effects dominate AI-characteristic effects, and that the simulations reproduce established personality-negotiation findings.

Significance. If the causal and validity claims survive scrutiny, the proposed framework would be a scalable, controlled testbed for personality-aware evaluation of agentic AI, addressing a real gap beyond task-completion metrics. The paper is strong in its scope (thousands of simulated episodes across two negotiation settings), its ambition to move from correlation to causal analysis, and its candid acknowledgment of limitations, including prompt-based personality manipulations and missing non-verbal cues. However, two load-bearing issues currently prevent the advertised conclusions from being supported: the reported analyses are SEM weights that are never connected to the described causal-discovery pipeline, and the LLM-based evaluation pipeline is vulnerable to circularity because the same model family generates the dialogues and scores them with access to the personality labels.

major comments (4)
  1. [Section 4.1.4 and Figures 2–10] The Methods describe causal discovery with CausalNex and causal-forest average treatment effect estimation, but the Results report only "SEM Weights" in all figures, with no description of the structural equation model, its specification, estimation method, or fit statistics. The paper never explains how these weights relate to the promised ATEs or to the "when X increases, we see a decrease in Y" statements in Section 4.1.4. Because the causal language in the Abstract and General Discussion depends on this analysis, the authors must either report the actual causal-forest ATEs with confidence intervals and intervention definitions or specify the SEM and justify that its coefficients estimate causal, rather than associational, effects.
  2. [Section 3.2.1 and Appendix A] Sotopia-Eval scores are produced by an LLM evaluator that appears to be applied to episodes whose input includes the full character profile, including explicit personality labels such as "Personality Trait: Introversion" (Appendix A). Since the same model family (gpt-4o-mini) generates the dialogues and scores them, Believability and Goal scores can reflect the prompt label rather than the actual dialogue behavior. This makes the General Discussion's "strong evidence" claim (Section 6) vulnerable to circularity. The paper needs to show that the evaluator is blind to the trait labels (for example, by ablating the profile from the evaluation input), or corroborate the scenario-based measures with human judgments, or substantially temper the causal interpretation.
  3. [Section 4.1.2 and Table 2] The lexical measures (DistilBERT-based empathy, emotion, morality, and subjectivity classifiers) are applied to LLM-generated negotiation dialogues without any validation against human-annotated texts of this type. The observed trait effects on lexical outcome variables could therefore be artifacts of classifier sensitivity to prompt phrasing or genre-specific language rather than genuine personality-consistent communication differences. Please provide validation evidence, at least on a held-out sample of the generated dialogues, or explicitly reframe the lexical results as exploratory rather than confirmatory evidence.
  4. [Section 6 and Section 7] The General Discussion describes the results as "strong evidence" for theory-consistent personality effects, but Section 7 itself acknowledges that prompt-based personality manipulations may not capture the complexity of real human personality, that lexical measures miss non-verbal cues, and that only two scenarios were considered. Combined with the evaluator-circularity and analysis-reporting issues above, the strength of the claim goes beyond what the evidence supports. The conclusions should be recast as proof-of-concept results that require human validation and an evaluator-blindness check.
minor comments (7)
  1. [Section 4.1.3] The sentence "running a total of 4343 episodes for each treatment combination setting" is inconsistent with the reported total of 8686 transcripts; please clarify whether 4343 is the total number of episodes or the number per condition and reconcile the transcript count.
  2. [Section 5.1.1] The three AI dimensions are named "Transparency, Adaptability, and Reliability" in Section 5.1.1, but Appendix C, the Abstract, and all Results figures refer to "Competence" rather than "Reliability." Please use one consistent terminology throughout.
  3. [Section 4.1.2] The text says "seven original dimensions" and then lists eight measures: Believability, Financial and Material Benefits, Goal, Knowledge, Overall Score, Relationship, Secret, and Social Rules. Clarify which dimensions are the seven original Sotopia-Eval dimensions and how Overall Score is defined.
  4. [Section 4.2.5] The phrase "Smaller, positive trait level differences were found for Conscientiousness and Conscientiousness and Openness" appears to be a typo; it should likely read "for Conscientiousness and Openness."
  5. [Section 5.2.3] The text references "Figure 8" for the empathy measures, but the empathy measures are displayed in Figure 7; the figure cross-reference needs correction.
  6. [Figure 3a caption] The caption misspells "Anticipating" as "Ancitipating."
  7. [Section 4.2.4] The clause "which were reversed for for Love and Joy indicators" contains a duplicated "for."

Circularity Check

1 steps flagged · score 4.0 of 10

Believability dimension is self-referential because the character profile contains the personality prompt; Goal and lexical measures retain independent content.

  1. self definitional [Table 1 / Sec. 3.2.1 and Appendix A; interpreted in Sec. 6]
    "Believability: How natural, realistic, and consistent the agent’s behavior is with its character profile [0, 10]. ... "personality_and_values": Personality Model: Big 5 Personality Personality Trait: Introversion Task Assignment: Prefers independent tasks and may struggle with collaboration."

    The manipulated independent variable is the Big Five trait written into Sotopia's character profile (Appendix A). Sotopia-Eval's Believability dimension is explicitly defined as consistency with that same character profile. Thus, when the trait prompt changes the profile, an agent that follows its own profile will be rated more believable by the definition of the scale; the trait-level-to-Believability association is partially secured by the measurement definition rather than by an independent behavioral outcome. The General Discussion nonetheless cites Believability as part of the 'strong evidence' that trait prompts produce theory-consistent behavior.

full rationale

The paper's strongest claim—that Experiment 1 provides 'strong evidence' that personality prompts produce behavior consistent with personality and negotiation theory—depends on multiple measures. One of those measures, Believability, is self-referential: Sotopia-Eval defines it as consistency with the character profile, and the character profile is exactly where the personality manipulation is inserted. This makes high Believability under a high-trait prompt partly an artifact of the evaluation definition. However, the same conclusion is also supported by Goal, Knowledge, and a suite of external lexical classifiers (empathy, moral foundations, sentiment, connotation frames), which do not reduce to the trait label. The Sotopia-Eval framework is a self-citation ([1] shares an author with this paper), but it is a published, code-released framework and therefore counts as independent support under the code-reproduced criterion. Section 7 honestly acknowledges the absence of human validation and the limitation that prompt-based manipulations may not capture human complexity; those are validity concerns, not circular reductions. Because one headline outcome reduces by construction while the central claim retains independent content, a moderate partial-circularity score is appropriate.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The central claims rest on two families of assumptions: (1) LLM-generated dialogues and LLM-based scores are valid stand-ins for human interaction and human evaluation, and (2) the causal analysis identifies true effects. The paper states these as design choices rather than validating them, and its own limitations section acknowledges the first family.

assumptions (4)
  • domain assumption Sotopia-Eval scores (believability, goal achievement, knowledge) are valid measures of negotiation quality and social outcomes.
    Invoked in Section 3.2.1 as the primary scenario-based outcome measures; no human-validation evidence is provided in this paper.
  • domain assumption Prompt-based Big Five descriptions induce the intended personality traits in LLM agents.
    Used throughout Experiments 1 and 2; Section 2.1 itself cites mixed evidence on BFI prompt efficacy, and Section 7 acknowledges the limitation.
  • standard math Causal discovery and causal forest estimates identify true causal effects, requiring correct DAG specification and no unobserved confounding.
    Assumed in Sections 4.1.4 and 5.1.4; treatment assignment is randomized by design, but the DAG structure and forest hyperparameters are not reported.
  • domain assumption LLM-generated dialogue is a faithful proxy for human negotiation communication.
    All lexical and scenario measures are computed on simulated transcripts; Section 7 notes this reliance and the absence of non-verbal cues.
invented entities (1)
  • Human digital twin (HDT) job candidate
    purpose: A simulated human job candidate used in Experiment 2 to stand in for human participants.
    The paper introduces the HDT concept in Section 5.1.1 and treats it as a simulated human, but no evidence is provided that HDT behavior matches real human candidates. This is a modeling construct, not a new physical entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues." pith.science (2026). https://pith.science/paper/2ODSS3JV

@misc{pith2026250615928,
  author       = {Pith},
  title        = {Pith review of: Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2ODSS3JV}},
  note         = {Machine review of arXiv:2506.15928}
}
read the original abstract

This paper presents an evaluation framework for agentic AI systems in mission-critical negotiation contexts, addressing the need for AI agents that can adapt to diverse human operators and stakeholders. Using Sotopia as a simulation testbed, we present two experiments that systematically evaluated how personality traits and AI agent characteristics influence LLM-simulated social negotiation outcomes--a capability essential for a variety of applications involving cross-team coordination and civil-military interactions. Experiment 1 employs causal discovery methods to measure how personality traits impact price bargaining negotiations, through which we found that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition outcomes. Sociocognitive lexical measures extracted from team communications detected fine-grained differences in agents' empathic communication, moral foundations, and opinion patterns, providing actionable insights for agentic AI systems that must operate reliably in high-stakes operational scenarios. Experiment 2 evaluates human-AI job negotiations by manipulating both simulated human personality and AI system characteristics, specifically transparency, competence, adaptability, demonstrating how AI agent trustworthiness impact mission effectiveness. These findings establish a repeatable evaluation methodology for experimenting with AI agent reliability across diverse operator personalities and human-agent team dynamics, directly supporting operational requirements for reliable AI systems. Our work advances the evaluation of agentic AI workflows by moving beyond standard performance metrics to incorporate social dynamics essential for mission success in complex operations.

Figures

Figures reproduced from arXiv: 2506.15928 by the authors.

Figure 1
Figure 1. Sotopia simulation framework [50]. take these various lexical measures in concert as we consider personality-linked impacts on our simulated negotiation scenarios. Alongside outcome-based and lexical measures, questionnaire measures have also been used to measure humans’ perceptions of their experiences during social negotiations. Research shows that subjective measures of a negotiation partner’s trustworthiness, fa… view at source ↗
Figure 2
Figure 2. Trait level–Sotopia-Eval SEM Weights Agreeableness Conscientiousness Extraversion Neuroticism Openness -0.25 0.00 0.25 0.50 Prepared Hopeful Anxious Annoyed Ancitipating (a) Empathy Emotion Measures Agreeableness Conscientiousness Extraversion Neuroticism Openness -0.1 0.0 0.1 0.2 0.3 Sympathizing Suggesting Questioning Neutral Encouraging Agreeing Acknowledging (b) Empathy Intent Measures [PITH_FULL_IMAGE:figures/… view at source ↗
Figure 3
Figure 3. Trait level–Empathy lexical measure SEM Weights [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Trait level–Socio-cognitive-emotion SEM Weights [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Scenario-based measure SEM Weights Scenario Agreeableness Extraversion AI_Adaptability AI_Competence AI_Transparency -0.8 -0.4 0.0 0.4 0.8 Self Performance Reliable Score Overall Experience Other Performance Frustration Score [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Questionnaire measure SEM Weights 5.2 Results 5.2.1 Scenario-based Measures Our results suggest that both HDT personality traits and AI bot characteristics play a crucial role in shaping several qualities of simulated job negotiation interactions (Fig￾ure 5). AI transp…
Figure 7
Figure 7. Figure 7: Empathy measure SEM Weights: “Hopeful” and “Apprehensive” are emotion [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Moral Foundation, Sentiment, and Emotion measure SEM Weights [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Lexical Connotation Frame Measures. Suffix indicates whose perspective is [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Subjectivity measure SEM weights 5.2.6 Lexical measures: Subjectivity Measures of subjective or evocative language ( [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Language of Bargaining: Linguistic Effects in LLM Negotiations

    cs.AI 2026-01 conditional novelty 5.0 of 10

    Language labels shift LLM negotiation outcomes in three games, but the claimed dominance over model choice is inconsistent with the reported model-level gaps.

Reference graph

Works this paper leans on

79 extracted references · 77 canonical work pages · cited by 1 Pith paper

  1. [1]

    SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, March 2024

    Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, March 2024

  2. [2]

    Cohen, Grant Engber- son, Laura Cassani, Trenton W

    Svitlana Volkova, Daniel Nguyen, Hsien-Te Kao, Myke C. Cohen, Grant Engber- son, Laura Cassani, Trenton W. Ford, Michael G. Yankoski, Mohammed Almu- tairi, Charles Chiang, Nandini Banerjee, Matthew Belcher, Tim Weninger, and Diego Gomez-Zara. VirTLab: Augmented Intelligence for Modeling and Evaluating Human-AI Teaming through Agent Interactions. in press

  3. [3]

    Cohen, Hsien-Te Kao, Grant Engberson, Louis Penafiel, Spencer Lynch, and Svitlana Volkova

    Daniel Nguyen, Myke C. Cohen, Hsien-Te Kao, Grant Engberson, Louis Penafiel, Spencer Lynch, and Svitlana Volkova. Exploratory Models of Human-AI Teams: Leveraging Human Digital Twins to Investigate Trust Development, November 2024

  4. [4]

    Human-Autonomy Teaming on Autonomous Vehicles with Large Language Model-Enabled Human Dig- ital Twins

    Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang. Human-Autonomy Teaming on Autonomous Vehicles with Large Language Model-Enabled Human Dig- ital Twins. In 2023 IEEE/ACM Symposium on Edge Computing (SEC) , pages 319– 324, December 2023

  5. [5]

    An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems, March 2024

    Hanqing Yang, Marie Siew, and Carlee Joe-Wong. An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems, March 2024

  6. [6]

    Fiore, and Florian Jentsch

    Rhyse Bendell, Jessica Williams, Stephen M. Fiore, and Florian Jentsch. Individual and team profiling to support theory of mind in artificial social intelligence.Scientific Reports, 14(1):12635, June 2024

  7. [7]

    Semmelink

    Adrian Furnham, Stephen Cuppello, and David S. Semmelink. Personality and Interpersonal Influence: Low Adjustment and Low Competitiveness is Associated With Low Assertiveness. Psychological Reports, page 00332941241246201, November 2024

  8. [8]

    How Personality Traits Influence Negotiation Out- comes? A Simulation based on Large Language Models

    Yin Jou Huang and Rafik Hadfi. How Personality Traits Influence Negotiation Out- comes? A Simulation based on Large Language Models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024 , pages 10336–10351, Miami, Florida, USA, November

Show all 79 references
  1. [9]

    Shadish, Thomas D

    William R. Shadish, Thomas D. Cook, and Donald T. Campbell. Experimental and Quasi-Experimental Designs for Generalized Causal Inference . Cengage Learning, Belmont, CA, 2nd edition edition, January 2001

  2. [10]

    CausalNex, October 2021

    Paul Beaumont, Ben Horsburgh, Philip Pilgerstorfer, Angel Droth, Richard Oen- taryo, Steven Ler, Hiep Nguyen, Gabriel Azevedo Ferreira, Zain Patel, and Wesley Leong. CausalNex, October 2021

  3. [11]

    Generalized random forests

    Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics , 47(2):1148–1178, April 2019. 19

  4. [12]

    McCrae and Oliver P

    Robert R. McCrae and Oliver P. John. An Introduction to the Five-Factor Model and Its Applications. Journal of Personality , 60(2):175–215, June 1992

  5. [13]

    Trait-names: A psycho-lexical study

    Gordon W Allport and Henry S Odbert. Trait-names: A psycho-lexical study. Psy- chological monographs, 47(1):i, 1936

  6. [14]

    Description and measurement of personality

    Raymond Bernard Cattell. Description and measurement of personality. 1946

  7. [15]

    Consistency of the factorial structures of personality ratings from different sources

    Donald W Fiske. Consistency of the factorial structures of personality ratings from different sources. The Journal of Abnormal and Social Psychology , 44(3):329, 1949

  8. [16]

    E. C. Tupes and R. E. Christal. Recurrent personality factors based on trait ratings. USAF ASD Tech. Rep. No. 61-97, US Air Force, Lackland Air Force Base, TX, 1961

  9. [17]

    Toward an adequate taxonomy of personality attributes: Repli- cated factor structure in peer nomination personality ratings

    Warren T Norman. Toward an adequate taxonomy of personality attributes: Repli- cated factor structure in peer nomination personality ratings. The journal of abnor- mal and social psychology , 66(6):574, 1963

  10. [18]

    The revised neo personality inventory (neo-pi- r)

    Paul T Costa and Robert R McCrae. The revised neo personality inventory (neo-pi- r). The SAGE handbook of personality theory and assessment , 2(2):179–198, 2008

  11. [19]

    LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models, February 2024

    Ivar Frisch and Mario Giulianelli. LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models, February 2024

  12. [20]

    Estimating the Personality of White-Box Language Models, May 2023

    Saketh Reddy Karra, Son The Nguyen, and Theja Tulabandhula. Estimating the Personality of White-Box Language Models, May 2023

  13. [21]

    Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. Who is ChatGPT? Bench- marking LLMs’ Psychological Portrayal Using PsychoBench, January 2024

  14. [22]

    PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, April 2024

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kab- bara. PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, April 2024

  15. [23]

    The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents, February 2025

    Yifan Duan, Yihong Tang, Xuefeng Bai, Kehai Chen, Juntao Li, and Min Zhang. The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents, February 2025

  16. [24]

    Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model’s Personality, December 2024

    Yiming Ai, Zhiwei He, Ziyin Zhang, Wenhong Zhu, Hongkun Hao, Kai Yu, Lingjun Chen, and Rui Wang. Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model’s Personality, December 2024

  17. [25]

    Petrov, Gregory Serapio-Garc ´ ıa, and Jason Rentfrow

    Nikolay B. Petrov, Gregory Serapio-Garc ´ ıa, and Jason Rentfrow. Limited Ability of LLMs to Simulate Human Psychological Behaviours: A Psychometric Analysis, May 2024

  18. [26]

    Exploring the Potential of Large Language Models to Simu- late Personality, February 2025

    Maria Molchanova, Anna Mikhailova, Anna Korzanova, Lidiia Ostyakova, and Alexandra Dolidze. Exploring the Potential of Large Language Models to Simu- late Personality, February 2025

  19. [27]

    The Art and Science of Negotiation

    Howard Raiffa. The Art and Science of Negotiation . Harvard University Press, 1982. 20

  20. [28]

    The role of personality in successful negotiating

    Roderick W Gilkey and Leonard Greenhalgh. The role of personality in successful negotiating. Negotiation Journal, 2(3):245–256, 1986

  21. [29]

    Personality and Negotiation Performance: The People Matter, January 2015

    Mary Sass and Matthew Liao-Troth. Personality and Negotiation Performance: The People Matter, January 2015

  22. [30]

    Gerui (Grace) Kang, Lin Xiu, and Alan C. Roline. How do interviewers respond to applicants’ initiation of salary negotiation? An exploratory study on the role of gender and personality. Evidence-based HRM: a Global Forum for Empirical Schol- arship, 3(2):145–158, August 2015

  23. [31]

    Bargainer characteristics in distributive and integrative negotiation

    Bruce Barry and Raymond A Friedman. Bargainer characteristics in distributive and integrative negotiation. Journal of personality and social psychology , 74(2):345, 1998

  24. [32]

    Amanatullah, Michael W

    Emily T. Amanatullah, Michael W. Morris, and Jared R. Curhan. Negotiators who give too much: Unmitigated communion, relational anxieties, and economic costs in distributive and integrative bargaining. Journal of Personality and Social Psychology, 95(3):723–738, 2008

  25. [33]

    Team coordination dynamics

    Jamie C Gorman, Polemnia G Amazeen, and Nancy J Cooke. Team coordination dynamics. Nonlinear dynamics, psychology, and life sciences , 14(3):265–289, July 2010

  26. [34]

    Amazeen, Nathan J

    Mustafa Demir, Polomnia G. Amazeen, Nathan J. McNeese, Aaron Likens, and Nancy J. Cooke. Team Coordination Dynamics in Human-Autonomy Teaming. Pro- ceedings of the Human Factors and Ergonomics Society Annual Meeting , 61(1):236– 236, September 2017

  27. [35]

    Big Five personality traits in simulated negotiation settings

    Pedro Fontes Falc˜ ao, Manuel Saraiva, Eduardo Santos, and Miguel Pina E Cunha. Big Five personality traits in simulated negotiation settings. EuroMed Journal of Business, 13(2):201–213, July 2018

  28. [36]

    Pennebaker and Laura A

    James W. Pennebaker and Laura A. King. Linguistic styles: Language use as an individual difference. Journal of Personality and Social Psychology, 77(6):1296–1312, 1999

  29. [37]

    Tausczik and James W

    Yla R. Tausczik and James W. Pennebaker. The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods. Journal of Language and Social Psychology, 29(1):24–54, March 2010

  30. [38]

    Towards emotion-aware agents for improved user satisfaction and partner perception in negotiation dialogues

    Kushal Chawla, Rene Clever, Jaysa Ramirez, Gale M Lucas, and Jonathan Gratch. Towards emotion-aware agents for improved user satisfaction and partner perception in negotiation dialogues. IEEE Transactions on Affective Computing , 2023

  31. [39]

    Graziano, Meara M

    William G. Graziano, Meara M. Habashi, Brad E. Sheese, and Ren´ ee M. Tobin. Agreeableness, empathy, and helping: A person × situation perspective. Journal of Personality and Social Psychology , 93(4):583–599, 2007

  32. [40]

    Does GPT-3 Generate Empa- thetic Dialogues? A Novel In-Context Example Selection Method and Automatic 21 Evaluation Metric for Empathetic Dialogue Generation

    Young-Jun Lee, Chae-Gyun Lim, and Ho-Jin Choi. Does GPT-3 Generate Empa- thetic Dialogues? A Novel In-Context Example Selection Method and Automatic 21 Evaluation Metric for Empathetic Dialogue Generation. In Nicoletta Calzolari, Chu- Ren Huang, Hansaem Kim, James Pustejovsky,...

  33. [41]

    Morality between the lines: Detecting moral sentiment in text

    Justin Garten, Reihane Boghrati, Joe Hoover, Kate M Johnson, and Morteza De- hghani. Morality between the lines: Detecting moral sentiment in text. In Proceed- ings of IJCAI 2016 Workshop on Computational Modeling of Attitudes , 2016

  34. [43]

    Jesse Graham, Jonathan Haidt, and Brian A. Nosek. Liberals and conservatives rely on different sets of moral foundations. Journal of Personality and Social Psychology , 96(5):1029–1046, May 2009

  35. [44]

    Detoxify, November 2020

    Laura Hanu and Unitary team. Detoxify, November 2020

  36. [45]

    Curhan, Hillary Anger Elfenbein, and Heng Xu

    Jared R. Curhan, Hillary Anger Elfenbein, and Heng Xu. What do people value when they negotiate? Mapping the domain of subjective value in negotiation. Journal of Personality and Social Psychology , 91(3):493–512, September 2006

  37. [46]

    Curhan, Noah Eisenkraft, Aiwa Shirako, and Lucio Baccaro

    Hillary Anger Elfenbein, Jared R. Curhan, Noah Eisenkraft, Aiwa Shirako, and Lucio Baccaro. Are Some Negotiators Better Than Others? Individual Differences in Bar- gaining Outcomes. Journal of research in personality , 42(6):1463–1475, December 2008

  38. [47]

    Conlon, and Remus Ilies

    Nikolaos Dimotakis, Donald E. Conlon, and Remus Ilies. The mind and heart (lit- erally) of the negotiator: Personality and contextual determinants of experiential reactions and economic outcomes in negotiation. Journal of Applied Psychology , 97(1):183–193, 2012

  39. [48]

    The Effect of Virtual Agent Warmth on Human-Agent Negotiation

    Pooja Prajod, Mohammed Al Owayyed, and Tim Rietveld. The Effect of Virtual Agent Warmth on Human-Agent Negotiation. 2019

  40. [49]

    Zhou, Gloria Mark, Jingyi Li, and Huahai Yang

    Michelle X. Zhou, Gloria Mark, Jingyi Li, and Huahai Yang. Trusting Virtual Agents: The Effect of Personality. ACM Trans. Interact. Intell. Syst. , 9(2-3):10:1–10:36, March 2019

  41. [50]

    Sotopia: Interactive evaluation for social intelligence in language agents

    Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Zhengyang Qi, Haofei Yu, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. Sotopia: Interactive evaluation for social intelligence in language agents. In- ternational Conference on Learning Rep...

  42. [51]

    What makes a good conversation? how controllable attributes affect human judgments

    Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. What makes a good conversation? how controllable attributes affect human judgments. arXiv preprint arXiv:1902.08654, 2019. 22

  43. [52]

    Connotation Frames: A Data- Driven Investigation, August 2016

    Hannah Rashkin, Sameer Singh, and Yejin Choi. Connotation Frames: A Data- Driven Investigation, August 2016

  44. [53]

    Moral foundations theory: The pragmatic validity of moral pluralism

    Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. Moral foundations theory: The pragmatic validity of moral pluralism. In Advances in experimental social psychology , volume 47, pages 55–130. Elsevier, 2013

  45. [54]

    Truth of Varying Shades: Analyzing Language in Fake News and Political Fact- Checking

    Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. Truth of Varying Shades: Analyzing Language in Fake News and Political Fact- Checking. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors, Proceed- ings of the 2017 Conference on Empirical M...

  46. [55]

    DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter, March 2020

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter, March 2020

  47. [56]

    Detoxify

    Laura Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020

  48. [57]

    DistilBERT for emotion recognition, May 2024

    Bhadresh Savani. DistilBERT for emotion recognition, May 2024

  49. [58]

    Machine intelligence to detect, characterise, and defend against in- fluence operations in the information environment

    M Glenski, E Ayton, E Saldanha, J Mendoza, D Arendt, Z Shaw, K Cronk, S Smith, and M Greaves. Machine intelligence to detect, characterise, and defend against in- fluence operations in the information environment. Journal of Information Warfare , 20(2):42–66, 2021

  50. [59]

    BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding, May 2019

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding, May 2019

  51. [60]

    A Deep Dive into Multilingual Hate Speech Classification

    Sai Saketh Aluru, Binny Mathew, Punyajoy Saha, and Animesh Mukherjee. A Deep Dive into Multilingual Hate Speech Classification. In Yuxiao Dong, Georgiana Ifrim, Dunja Mladeni´ c, Craig Saunders, and Sofie Van Hoecke, editors, Machine Learning and Knowledge Discovery in Databas...

  52. [61]

    Decoupling strategy and generation in negotiation dialogues, 2018

    He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. Decoupling strategy and generation in negotiation dialogues, 2018

  53. [62]

    The Book of Why: The New Science of Cause and Effect

    Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect. Basic Books, New York, 1st edition edition, May 2018

  54. [63]

    EconML: A python package for ml-based hetero- geneous treatment effects estimation

    Keith Battocchi, Eleanor Dillon, Maggie Hei, Greg Lewis, Paul Oka, Miruna Oprescu, and Vasilis Syrgkanis. EconML: A python package for ml-based hetero- geneous treatment effects estimation. https://github.com/py-why/EconML, 2019. Version 0.x

  55. [64]

    Bradley, John E

    Bret H. Bradley, John E. Baur, Christopher G. Banford, and Bennett E. Postleth- waite. Team Players and Collective Performance: How Agreeableness Affects Team Performance Over Time. Small Group Research, 44(6):680–711, December 2013. 23

  56. [65]

    Driskell, Gerald F

    James E. Driskell, Gerald F. Goodwin, Eduardo Salas, and Patrick Gavan O’Shea. What makes a good team player? Personality and team effectiveness. Group Dy- namics: Theory, Research, and Practice , 10(4):249–271, December 2006

  57. [66]

    Adam M. Grant. Rethinking the Extraverted Sales Ideal: The Ambivert Advantage. Psychological Science, 24(6):1024–1030, June 2013

  58. [67]

    Lepine, Jason A

    Jeffrey A. Lepine, Jason A. Colquitt, and Amir Erez. Adaptability to Changing Task Contexts: Effects of General Cognitive Ability, Conscientiousness, and Openness to Experience. Personnel Psychology, 53(3):563–593, 2000

  59. [68]

    Suzanne T. Bell. Deep-level composition variables as predictors of team performance: A meta-analysis. The Journal of Applied Psychology , 92(3):595–615, May 2007

  60. [69]

    K. J. Klein, J. L. Saltz, and D. M. Mayer. HOW DO THEY GET THERE? AN EXAMINATION OF THE ANTECEDENTS OF CENTRALITY IN TEAM NET- WORKS. Academy of Management Journal , 47(6):952–963, December 2004

  61. [70]

    Miranda A. G. Peeters, Harrie F. J. M. Van Tuijl, Christel G. Rutte, and Isabelle M. M. J. Reymen. Personality and team performance: A meta-analysis. European Journal of Personality , 20(5):377–396, August 2006

  62. [71]

    The PANAS-X: Manual for the positive and negative affect schedule-expanded form

    David Watson and Lee Anna Clark. The PANAS-X: Manual for the positive and negative affect schedule-expanded form. 1994

  63. [72]

    Robert R. McCrae. NEO-PI-R Data from 36 Cultures. In Robert R. McCrae and J¨ uri Allik, editors, The Five-Factor Model of Personality Across Cultures , pages 105–125. Springer US, Boston, MA, 2002

  64. [73]

    Habashi, William G

    Meara M. Habashi, William G. Graziano, and Ann E. Hoover. Searching for the Prosocial Personality: A Big Five Approach to Linking Personality and Prosocial Behavior. Personality and Social Psychology Bulletin , 42(9):1177–1192, September 2016

  65. [74]

    Hirsh, Colin G

    Jacob B. Hirsh, Colin G. DeYoung, Xiaowen Xu, and Jordan B. Peterson. Com- passionate Liberals and Polite Conservatives: Associations of Agreeableness With Political Ideology and Moral Values. Personality and Social Psychology Bulletin , 36(5):655–664, May 2010

  66. [75]

    Extraversion and Its Positive Emotional Core

    David Watson and Lee Anna Clark. Extraversion and Its Positive Emotional Core. In Handbook of Personality Psychology , pages 767–793. Elsevier, 1997

  67. [76]

    P. A. Hancock, Theresa T. Kessler, Alexandra D. Kaplan, John C. Brill, and James L. Szalma. Evolving Trust in Robots: Specification Through Sequential and Compar- ative Meta-Analyses. Human Factors, 63(7):1196–1229, November 2021

  68. [77]

    Schaefer, Jessie Y

    Kristin E. Schaefer, Jessie Y. C. Chen, James L. Szalma, and P. A. Hancock. A Meta- Analysis of Factors Influencing the Development of Trust in Automation: Implica- tions for Understanding Autonomy in Future Systems. Human Factors, 58(3):377– 400, May 2016. 24

  69. [78]

    Hancock, Deborah R

    Peter A. Hancock, Deborah R. Billings, Kristin E. Schaefer, Jessie Y. C. Chen, Ewart J. de Visser, and Raja Parasuraman. A Meta-Analysis of Factors Affecting Trust in Human-Robot Interaction. Human Factors: The Journal of the Human Factors and Ergonomics Society , 53(5):517–52...

  70. [79]

    first_name

    Sarah A. Jessup, Tamera R. Schneider, Gene M. Alarcon, Tyler J. Ryan, and August Capiola. The Measurement of the Propensity to Trust Automation. In Jessie Y.C. Chen and Gino Fragomeni, editors, Virtual, Augmented and Mixed Reality. Appli- cations and Case Studies , pages 476–4...

  71. [2024]

    Association for Computational Linguistics

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.