Pith. sign in

REVIEW 5 major objections 5 minor 1 cited by

LLMs can turn a personality trait into both a virtual agent's speech and its body language, so that viewers can reliably tell an extravert from an introvert.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

LLM prompting can steer both speech and nonverbal cues of virtual agents toward intended extraversion levels, with human observers detecting the difference.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Solid application paper with a clean user study, but the abstract overstates a scenario-dependent result and the nonverbal analysis lacks inferential support. the 5 major comments →

arxiv 2508.21087 v1 pith:LAET3ZGW submitted 2025-08-27 cs.HC

Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?

classification cs.HC
keywords large language modelsvirtual agentspersonalityextraversionnonverbal behaviorpersonality promptingLIWCuser perception
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a single LLM pipeline can generate both verbal and nonverbal behavior for a virtual agent from a personality prompt, and that people can read the intended trait from what the agent says and how it moves. Focusing on extraversion, the authors built introverted and extroverted agents, ran them through negotiation and ice-breaking conversations, and checked the outputs with automated text analysis and a human perception study. The automated checks show that the agents' word choices track known linguistic markers of extraversion, and that the selected gestures, facial expressions, and voice settings follow the intended trait. The user study shows that viewers can reliably distinguish the two agent types, with verbal content the strongest cue. If the claim holds, personality-aligned virtual characters become a prompt away rather than requiring hand-built behavior rules.

Core claim

The paper's central claim is that LLMs can generate coherent verbal and nonverbal behaviors for embodied virtual agents that consistently reflect an intended personality trait, specifically extraversion, and that human observers can recognize the trait from those behaviors. In agent-to-agent simulations, extroverted agents produced longer utterances, more social and power-related words, and fewer tentative and cognitive-process words than introverted agents, in directions mostly consistent with the human psychology literature. The nonverbal action choices also split cleanly: extraverts got broader gestures, more expressive faces, and louder/faster speech, while introverts got restrained gest

What carries the argument

The load-bearing component is the Nonverbal Action List, a predefined taxonomy of facial labels, body movements, and voice settings (such as 'Smile Broadly', 'Gesture Narrowly', 'Fast Pace') derived from empirical markers of extraversion. Around it sits a Nonverbal Description Generation Module that uses GPT to turn each animation clip into natural-language descriptions of the physical movement, plus a personality prompt that tells the LLM which trait to project. The LLM selects actions from the list as it writes each utterance, so verbal and nonverbal choices are generated together rather than separately; the list is what lets the system map an abstract trait into concrete, animatable behav

Load-bearing premise

The paper assumes that the extraversion markers established for human behavior and language—broad gestures, frequent gaze, fast loud speech, specific word categories—transfer directly to what an LLM writes and what a virtual agent animates, and that the text-analysis tools and the prompt-selected action list measure those markers faithfully.

What would settle it

Show viewers the same generated script with the opposite agent's animations and voice, swapping which body and voice accompany which words. If ratings track the words rather than the body/voice, the nonverbal-generation claim fails; if ratings track the body/voice only, the verbal claim is weakened. A more targeted version: run the BERT classifier on matched utterances from introvert and extravert prompts in the negotiation scenario—where the current classifier barely separates them—and see whether the trait signal disappears outside the ice-breaking scenario.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Personality-aligned virtual agents can be built on demand for negotiation, ice-breaking, and similar tasks without task-specific training data or hand-coded scripts.
  • Because verbal content dominates perceived personality, prompt designers can expect the strongest trait signal to come from wording, with gesture and expression as supporting cues.
  • Scenario context matters: the same personality prompts produced clearer trait differences in ice-breaking than in negotiation, so task constraints can mask or amplify personality expression.
  • The pipeline is reusable across conversational scenarios, since the same action list and prompt structure ran both experiments without redesign.
  • The trait is readable by untrained raters in short video clips, which makes the generated agents usable as experimental confederates in social-interaction research.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the negotiation scenario's automated classifier did not separate the traits while ice-breaking did, the practical limit may be the LLM's task model rather than its personality control; testing the same prompts across more tasks would map where the manipulation weakens.
  • Since only one voice identity was used per trait (Eric for the extravert, Brian for the introvert), some of the perceived difference could come from voice identity rather than generated behavior; a cross-voice replication with the same scripts would isolate the behavioral contribution.
  • The 'assumed similarity' effect the authors report suggests participant personality should be treated as a covariate in any agent-personality study; future experiments could deliberately recruit extreme introverts and extraverts to amplify the interaction.
  • The action-list architecture could transfer to other Big Five dimensions by replacing the marker set, but the trait-to-language mapping would first need to be validated with the same two-step pipeline.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a framework that uses LLM prompting to generate both verbal and nonverbal behaviors for embodied virtual agents, conditioned on the personality trait of extraversion. The method compiles a nonverbal action list from prior literature, embeds it in prompts together with personality definitions and animation descriptions, and renders selected actions in a Unity-based agent. Experiment 1 uses agent-agent simulations in negotiation and ice-breaking scenarios, evaluating generated utterances with LIWC and a BERT-based personality classifier, and reporting average probabilities of nonverbal actions. Experiment 2 is a user study in which participants watch videos of extroverted and introverted agents and rate perceived extraversion, cue influence, and behavior consistency. The paper claims that LLMs can generate verbal and nonverbal behaviors aligned with personality traits and that users can recognize these traits. The user study is the strongest component; the generation-side evidence has important gaps, including a null classifier result in one scenario and a lack of inferential statistics for nonverbal behavior.

Significance. If the central claim were fully supported, the framework would be a practically useful method for building controllable virtual confederates, potentially enabling systematic studies of personality effects in social interaction. The paper makes a good-faith effort to evaluate both generation and perception: the agent-agent simulation covers two distinct social contexts; the user study uses a mixed design, multiple video versions to control for idiosyncratic content, and collects cue-influence and consistency judgments; the BFI-10 measure of participant extraversion enables a perception-interaction analysis. These are real strengths. However, the evidence for the generation claim is currently incomplete: the BERT personality classifier fails to distinguish the two agent types in the negotiation scenario, and the nonverbal analyses are descriptive only. The paper also does not address the degree to which the observed behavioral differences are constructed by the prompt itself, since the evaluation reuses the same action taxonomy that was embedded in the prompts. The user study provides independent evidence for human perception, but the overall claim as stated in the abstract and concl

major comments (5)
  1. [§4.3.1, Figure 3] The BERT personality classifier results are internally inconsistent and do not support the unqualified claim in the abstract. The text states that in the Negotiation scenario both Extrovert and Introvert Agents had the same proportion (68%) of utterances classified as extroverted, yet then reports a 'trend' with χ²(1)=2.65, p=.10. If the proportions are identical, the chi-square must be zero. Please report the actual per-agent frequencies and the correct 2x2 table for each scenario. More substantively, this null result in one of two scenarios undercuts the claim that 'LLMs can generate verbal behaviors that align with personality traits'; at best the claim holds for the ice-breaking context. The discussion should either temper the claim or provide evidence for why the negotiation null is not problematic.
  2. [§4.3.2, Figures 4-5] The nonverbal behavior analysis is purely descriptive. The paper reports average probabilities of actions such as Gesture Widely and Smile Broadly, but no statistical tests, confidence intervals, or effect sizes are provided. Consequently, the reader cannot infer that the observed differences between extroverted and introverted agents are reliable rather than noise from the 10 simulation trials. This is load-bearing for RQ1 and for the Discussion statement that 'generated non-verbal behaviors also varied meaningfully based on personality' (§6). Please add inferential statistics (e.g., mixed-effects models or permutation tests on action probabilities), or reframe the section as an exploratory illustration and adjust the abstract and conclusions accordingly.
  3. [§3.1–§3.2, §4.2] There is a circularity concern in the evaluation of the generation pipeline. The Nonverbal Action List is compiled from the same personality literature that defines which behaviors should be extroverted vs. introverted (e.g., 'Subtle expressions ... are designed to reflect the restrained emotional display often associated with introverts,' while 'Extreme expressions ... capture the dynamic and exaggerated expressions typical of extroverts'). This list is then embedded in the LLM prompts, and the analysis in §4.2 measures the frequencies of exactly those action labels. The observed nonverbal differences may therefore be an artifact of the prompt instruction rather than evidence that the LLM can independently generate trait-consistent behaviors. At minimum, the paper should acknowledge this confound and, ideally, include an ablation where the action list is not annotated with trait associa
  4. [§4.1, Abstract, §7] The paper generalizes from a single LLM (gpt-4o-mini-2024-07-18, §4.1) to 'LLMs' in the abstract and conclusion. While this is a common first step, the central claim is about LLM capability, and a single model with default temperature provides no evidence of cross-model consistency. The abstract and discussion should either be qualified to 'GPT-4o-mini' or the authors should report results with at least one or two additional models (e.g., a Llama or Claude variant) to support the general claim. This is especially important because the personality-prompting literature cited in §2.2 shows considerable variation across model families.
  5. [§4.3.1, Tables 2-3] The LIWC analysis is selectively reported. The text says features were filtered to those with significant t-tests (p<.05) and Cohen's d>0.5, but no correction for multiple comparisons is applied despite testing dozens of LIWC categories. Moreover, the 'Aligned' column is a post-hoc, subjective judgment, and two features (nonfluencies, informal) are flagged as not aligned with prior literature yet still included in the 'success' narrative. Reporting all tested features with effect sizes and a clear multiple-comparison policy would strengthen the claim that the LIWC differences are systematic rather than cherry-picked.
minor comments (5)
  1. [§5.2.2] The reported statistic for Video Version, F=2.99, p=.622, appears implausible; the degrees of freedom are missing and the p-value is likely inconsistent with that F value. Please report the complete ANOVA table with df and corrected values.
  2. [§4.1] The text says 'We use GPT1'—this should be 'GPT-4o-mini' with the exact model identifier already given in the footnote. Similarly, 'approachs' in §3.1 is a typo.
  3. [Figure 3] The axis labels of Figure 3 are ambiguous: it is not clear whether the y-axis is the proportion of utterances or the proportion of trials, and whether the percentages are averaged over the ten simulation runs. Please clarify in the caption.
  4. [§5.2.3] The interaction between participant extraversion and agent type is interesting, but the paper does not report the regression coefficients or the cutoff used to define 'introverted' and 'extroverted' participants in Figure 7. Please provide these details.
  5. [§5.1] The sample-size description is ambiguous: '30 individuals assigned to one of the two scenarios' could mean 30 total or 30 per scenario. Given the between-subjects scenario factor, please state the number of participants per scenario explicitly.

Circularity Check

2 steps flagged

Nonverbal evaluation reuses the same action list inserted into the prompts, and voice stimuli were preselected via a perception pilot, so part of the validation is constructed.

specific steps
  1. self definitional [§3.1–3.2 and §4.2 (Nonverbal Action List; Behavior Analysis)]
    "Based on prior work, we compiled a set of observable behaviors—our Nonverbal Action List—used in prompts and animation descriptions to guide the LLM’s generation of appropriate nonverbal actions... LLM selects appropriate nonverbal behaviors from a predefined action list, ensuring that gestures, facial expressions, and speech characteristics align with the intended personality traits... For nonverbal behavior analysis, we evaluate the predicted probabilities of selected nonverbal actions from a predefined list to identify patterns specific to introverted and extroverted agents."

    The same hand-compiled, personality-labeled action list is fed into the prompt as the generation target and then used as the evaluation rubric. Nonverbal 'results'—e.g., extroverts use Gesture Widely and Smile Broadly, introverts use Gesture Narrowly and Avert Gaze—are therefore largely restatements of the prompt’s built-in mapping rather than independent evidence that the LLM inferred extraversion from personality alone. Some stochasticity remains, so this is partial circularity rather than a complete identity.

  2. fitted input called prediction [§3.2 (voice pilot) and §5.2.2 (perceived personality)]
    "To ensure that the selected voices matched the intended personality traits, we conducted a pilot study with 15 participants using four American English voice IDs provided by ElevenLabs. Based on participant feedback, the extrovert agent was assigned voice ID Eric, while the introvert agent used voice ID Brian. ... The analysis revealed a significant main effect of Agent Type (F = 44.57, p <.001), indicating that extroverted agents were perceived as significantly more extroverted than introverted agents."

    The two voice conditions were selected from pilot participants’ personality judgments, so the stimuli were deliberately built to differ in perceived extraversion before the main user study ran. The Experiment 2 main effect of Agent Type on perceived extraversion is therefore not an independent test of LLM-generated behavior: one of the manipulated cues (voice identity) was pre-fitted to produce exactly that perception. This inflates the conclusion that users can recognize the intended traits through the agents’ behaviors.

full rationale

The core derivation chain is not entirely circular: the LIWC analyses and the pre-trained BERT personality classifier are external, not fitted to the current data, and the user study is a real perception experiment with a significant Agent Type main effect. However, two load-bearing steps reduce partly to their own inputs. First, the Nonverbal Action List is compiled from the same personality literature that defines the expected differences, inserted into the LLM prompts as the generation target, and then reused as the nonverbal evaluation rubric—so the nonverbal 'alignment' is substantially constructed by the prompt’s built-in mapping. Second, the voices used in the user study were selected from a pilot in which participants judged which voice IDs matched the intended personality, meaning the perceived-extraversion result is partially forced by stimulus preselection. The paper’s abstract overstates the support, especially since the BERT classifier showed only a non-significant trend in the negotiation scenario; that is a support weakness rather than a circularity. Overall, the central claim retains independent content, so a moderate score of 4 is appropriate.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper imports psychological constructs and behavioral marker mappings from prior literature, then uses those same mappings to both prompt and evaluate the LLM. Hand-picked thresholds, a pilot-selected voice pair, and one LLM shape all conclusions. No new entities are postulated.

free parameters (3)
  • LIWC significance thresholds = p < .05, Cohen's d > 0.5
    Hand-chosen thresholds for reporting features; affects which features are presented as evidence.
  • Nonverbal action-to-trait mapping = e.g., Smile Broadly->extrovert, Coy Smile->introvert
    Hand-assigned based on literature; used in prompts and evaluation, partly determines the outcome.
  • Voice selection = Eric (extrovert), Brian (introvert)
    Selected via pilot study with 15 participants; affects nonverbal voice cue perception.
axioms (4)
  • domain assumption Big Five personality model is a valid description of personality.
    Used as the theoretical foundation throughout Sections 2 and 3.
  • domain assumption Prior findings on extraversion and behavior (Rutter, Buck, etc.) transfer to LLM-generated utterances and animations.
    This transfer is the core bridge between human psychology and the agent evaluation; invoked in Sections 3.1 and 4.2.
  • domain assumption The BERT-based personality classifier from Kazameini et al. is valid for classifying extraversion in short LLM-generated utterances.
    Used as a key evaluation metric in Section 4.2 without validation on LLM output.
  • domain assumption LIWC categories are reliable indicators of personality in this context.
    LIWC-based alignment is a central evidence source in Section 4.3.1.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?." pith.science (2026). https://pith.science/paper/LAET3ZGW

@misc{pith2026250821087,
  author       = {Pith},
  title        = {Pith review of: Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LAET3ZGW}},
  note         = {Machine review of arXiv:2508.21087}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This study proposes a framework that employs personality prompting with Large Language Models to generate verbal and nonverbal behaviors for virtual agents based on personality traits. Focusing on extraversion, we evaluated the system in two scenarios: negotiation and ice breaking, using both introverted and extroverted agents. In Experiment 1, we conducted agent to agent simulations and performed linguistic analysis and personality classification to assess whether the LLM generated language reflected the intended traits and whether the corresponding nonverbal behaviors varied by personality. In Experiment 2, we carried out a user study to evaluate whether these personality aligned behaviors were consistent with their intended traits and perceptible to human observers. Our results show that LLMs can generate verbal and nonverbal behaviors that align with personality traits, and that users are able to recognize these traits through the agents' behaviors. This work underscores the potential of LLMs in shaping personality aligned virtual agents.

Figures

Figures reproduced from arXiv: 2508.21087 by Bin Han, Deuksin Kwon, Jonathan Gratch, Kaleen Shrestha, Spencer Lin.

Figure 1
Figure 1. Figure 1: Conversational system overview behind the intro [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Embodied Virtual Agent System Overview pace, reflecting their respective personality traits in communication. To reflect these differences, we include four voice-related labels in our action list: Loud Volume, Small Volume, Fast Pace, and Slow Pace. 3.2 LLM-driven Behavior Realization We implement a behavior realization pipeline that maps LLM￾generated output to verbal and nonverbal behaviors in an embodie… view at source ↗
Figure 3
Figure 3. Figure 3: Personality Detection Results [60] LIWC Analysis: The LIWC analysis revealed significant lexical differences between extroverted and introverted agents in both scenarios (see [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Negotiation- Distribution of nonverbal behavior probabilities across modalities. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ice-breaking- Distribution of nonverbal behavior probabilities across modalities. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Perceived Extroversion Score [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Interaction between Participant Personality and [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training

    cs.CL 2026-06 unverdicted novelty 5.0

    Introduces a French OSCE dialogue dataset of 240 interactions and a modular LLM-based controllable virtual patient generation system with multi-level LLM-as-Judge evaluation for clinical skills training.

Reference graph

Works this paper leans on

66 extracted references · 61 canonical work pages · cited by 1 Pith paper · 3 internal anchors

  1. [1]

    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision 958–979

    Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, and others 2024 A survey on multimodal large language models for autonomous driving. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision 958–979

  2. [2]

    arXiv preprint arXiv:2402.12451

    Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara 2024 The revolution of multimodal large language models: a survey. arXiv preprint arXiv:2402.12451

  3. [3]

    arXiv preprint arXiv:2408.01319

    Jiaqi Wang, Hanqi Jiang, Yiheng Liu, Chong Ma, Xu Zhang, Yi Pan, Mengyuan Liu, Peiran Gu, Sichen Xia, Wenjun Li, and others 2024 A comprehensive review of multimodal large language models: Performance and challenges across different tasks. arXiv preprint arXiv:2408.01319

  4. [4]

    In Proceedings of the Twenty-fourth International Symposium on Theory, Algo- rithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing 545–550

    Eleanor Lin, James Hale, and Jonathan Gratch 2023 Toward a better understand- ing of the emotional dynamics of negotiation with large language models. In Proceedings of the Twenty-fourth International Symposium on Theory, Algo- rithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing 545–550

  5. [5]

    In Proceedings of the 24th ACM Interna- tional Conference on Intelligent Virtual Agents 1–3

    James Hale, Lindsey Schweitzer, and Jonathan Gratch 2024 Integration of LLMs with Virtual Character Embodiment. In Proceedings of the 24th ACM Interna- tional Conference on Intelligent Virtual Agents 1–3

  6. [6]

    In Proceedings of the 8th international conference on Multimodal interfaces 51–58

    Norbert Reithinger, Patrick Gebhard, Markus Löckelt, Alassane Ndiaye, Norbert Pfleger, and Martin Klesen 2006 Virtualhuman: dialogic and affective interaction with virtual characters. In Proceedings of the 8th international conference on Multimodal interfaces 51–58

  7. [7]

    In Proceedings of the 2014 international conference on Autonomous agents and multi-agent systems 1061–1068

    David DeVault, Ron Artstein, Grace Benn, Teresa Dey, Ed Fast, Alesia Gainer, Kallirroi Georgila, Jon Gratch, Arno Hartholt, Margaux Lhommet, and others 2014 SimSensei Kiosk: A virtual human interviewer for healthcare decision support. In Proceedings of the 2014 international conference on Autonomous agents and multi-agent systems 1061–1068

  8. [8]

    In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems 1–17

    Hanseob Kim, Bin Han, Jieun Kim, Muhammad Firdaus Syawaludin Lubis, Gerard Jounghyun Kim, and Jae-In Hwang 2024 Engaged and affective virtual agents: their impact on social presence, trustworthiness, and decision-making in the group discussion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems 1–17

  9. [9]

    Computational Linguistics 37(3) 455–488 MIT Press One Rogers Street, Cambridge, MA 02142- 1209, USA journals-info˜

    François Mairesse and Marilyn A Walker 2011 Controlling user perceptions of linguistic style: Trainable generation of personality traits. Computational Linguistics 37(3) 455–488 MIT Press One Rogers Street, Cambridge, MA 02142- 1209, USA journals-info˜

  10. [10]

    In Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization 399–405

    Hassan Soliman, Milos Kravcik, Nagasandeepa Basvoju, and Patrick Jaehnichen 2024 Using Large Language Models for Adaptive Dialogue Management in Digital Telephone Assistants. In Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization 399–405

  11. [11]

    In Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents 1–10

    Mehdi Arjmand, Farnaz Nouraei, Ian Steenstra, and Timothy Bickmore 2024 Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents. In Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents 1–10. Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?

  12. [12]

    In Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents 1–10

    Ian Steenstra, Farnaz Nouraei, Mehdi Arjmand, and Timothy Bickmore 2024 Virtual agents for alcohol use counseling: Exploring llm-powered motivational interviewing. In Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents 1–10

  13. [13]

    In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems 1–7

    Hongyu Wan, Jinda Zhang, Abdulaziz Arif Suria, Bingsheng Yao, Dakuo Wang, Yvonne Coady, and Mirjana Prpa 2024 Building llm-based ai agents in social virtual reality. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems 1–7

  14. [14]

    arXiv preprint arXiv:2307.00184

    Gregory Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić 2023 Personality traits in large language models. arXiv preprint arXiv:2307.00184

  15. [15]

    Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu 2023 Evaluating and Inducing Personality in Pre-trained Language Models

  16. [16]

    William Morrow & Co

    Jim Blascovich and Jeremy Bailenson 2011 Infinite reality: Avatars, eternal life, new worlds, and the dawn of the virtual revolution. William Morrow & Co

  17. [17]

    In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 3 937–944

    Celso M de Melo, Peter Carnevale, and Jonathan Gratch 2011 The effect of expression of anger and happiness in computer agents on negotiations with humans. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 3 937–944

  18. [18]

    Large Language Models for Virtual Human Gesture Selection

    Parisa Ghanad Torshizi, Laura B Hensel, Ari Shapiro, and Stacy C Marsella 2025 Large Language Models for Virtual Human Gesture Selection. arXiv preprint arXiv:2503.14408

  19. [19]

    In Proceedings of the 19th ACM International Conference on Intelligent Virtual Agents 111–118

    Rens Hoegen, Deepali Aneja, Daniel McDuff, and Mary Czerwinski 2019 An end-to-end conversational style matching agent. In Proceedings of the 19th ACM International Conference on Intelligent Virtual Agents 111–118

  20. [20]

    In the 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) 1–5

    Yoon Kyung Lee, Yoonwon Jung, Gyuyi Kang, and Sowon Hahn 2023 Devel- oping Social Robots with Empathetic Non-Verbal Cues Using Large Language Models.(2023). In the 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) 1–5

  21. [21]

    In Intelligent Virtual Agents: 10th International Conference, IVA 2010, Philadelphia, PA, USA, September 20-22, 2010

    Michael Neff, Yingying Wang, Rob Abbott, and Marilyn Walker 2010 Evaluating the effect of gesture and language on personality perception in conversational agents. In Intelligent Virtual Agents: 10th International Conference, IVA 2010, Philadelphia, PA, USA, September 20-22, 2010. Proceedings 10 222–235 Springer

  22. [22]

    Journal of Research in Personality 32(1) 80–107 Elsevier

    Richard Lippa 1998 The nonverbal display and judgment of extraversion, mas- culinity, femininity, and gender diagnosticity: A lens model analysis. Journal of Research in Personality 32(1) 80–107 Elsevier

  23. [23]

    Journal of personality and social psychology 30(4) 587 American Psychological Association

    Ross Buck, Robert E Miller, and William F Caul 1974 Sex, personality, and physi- ological variables in the communication of affect via facial expression. Journal of personality and social psychology 30(4) 587 American Psychological Association

  24. [24]

    Journal of Research in Interactive Marketing 17(3) 416–433 Emerald Publishing Limited

    Eunjoo Jin and Matthew S Eastin 2023 Birds of a feather flock together: matched personality effects of product recommendation chatbots and users. Journal of Research in Interactive Marketing 17(3) 416–433 Emerald Publishing Limited

  25. [25]

    Intelligent Service Robotics 1() 169–183 Springer

    Adriana Tapus, Cristian Ţăpuş, and Maja J Matarić 2008 User—robot personality matching and assistive robot behavior adaptation for post-stroke rehabilitation therapy. Intelligent Service Robotics 1() 169–183 Springer

  26. [26]

    International journal of human-computer studies 53(2) 251–267 Elsevier

    Katherine Isbister and Clifford Nass 2000 Consistency of personality in interactive characters: verbal cues, non-verbal cues, and user characteristics. International journal of human-computer studies 53(2) 251–267 Elsevier

  27. [27]

    description of personality

    Lewis R Goldberg 2013 An alternative “description of personality”: The Big-Five factor structure. In Personality and personality disorders 34–47 Routledge

  28. [28]

    Journal of personality and social psychology 52(6) 1258 American Psychological Association

    Robert R McCrae 1987 Creativity, divergent thinking, and openness to experience. Journal of personality and social psychology 52(6) 1258 American Psychological Association

  29. [29]

    Personnel psychology 44(1) 1–26 Wiley Online Library

    Murray R Barrick and Michael K Mount 1991 The big five personality dimensions and job performance: a meta-analysis. Personnel psychology 44(1) 1–26 Wiley Online Library

  30. [30]

    In Handbook of personality psychology 795–824 Elsevier

    William G Graziano and Nancy Eisenberg 1997 Agreeableness: A dimension of personality. In Handbook of personality psychology 795–824 Elsevier

  31. [31]

    European Journal of Social Psychology 2(4) 371–384 Wiley Online Library

    DR Rutter, Ian E Morley, and Jane C Graham 1972 Visual interaction in a group of introverts and extraverts. European Journal of Social Psychology 2(4) 371–384 Wiley Online Library

  32. [32]

    Journal of Research in Personality 40(1) 21–34 Elsevier

    David C Funder 2006 Towards a resolution of the personality triad: Persons, situations, and behaviors. Journal of Research in Personality 40(1) 21–34 Elsevier

  33. [33]

    Journal of research in personality 42(4) 914–932 Elsevier

    Tera D Letzring 2008 The good judge of personality: Characteristics, behaviors, and observer accuracy. Journal of research in personality 42(4) 914–932 Elsevier

  34. [34]

    Journal of Research in Personality 108() 104437 Elsevier

    Ryan L Klinger and Nathapon Siangchokyoo 2024 Finding better raters: The role of observer personality on the validity of observer-reported personality in predicting job performance. Journal of Research in Personality 108() 104437 Elsevier

  35. [35]

    arXiv preprint arXiv:2402.01765

    Aleksandra Sorokovikova, Natalia Fedorova, Sharwin Rezagholi, and Ivan P Yamshchikov 2024 LLMs Simulate Big Five Personality Traits: Further Evidence. arXiv preprint arXiv:2402.01765

  36. [36]

    Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić 2023 Personality Traits in Large Language Models

  37. [37]

    arXiv preprint arXiv:2406.14703

    Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, and others 2024 Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics. arXiv preprint arXiv:2406.14703

  38. [38]

    In CCF International Conference on Natural Language Processing and Chinese Computing 241–254 Springer

    Shengyu Mao, Xiaohan Wang, Mengru Wang, Yong Jiang, Pengjun Xie, Fei Huang, and Ningyu Zhang 2024 Editing Personality for Large Language Models. In CCF International Conference on Natural Language Processing and Chinese Computing 241–254 Springer

  39. [39]

    In Proceedings of the Eleventh International Symposium of Chinese CHI 127–138

    Liuping Wang, Hongxin Li, Tong Wu, Fan Yang, Yang Yang, Jin Huang, and Feng Tian 2023 Extrovert Increases Consensus? Exploring the Effects of Conversational Agent Personality for Group Decision Support. In Proceedings of the Eleventh International Symposium of Chinese CHI 127–138

  40. [40]

    Frontiers in psychology 12() 660895 Frontiers Media SA

    Maryam Saberi, Steve DiPaola, and Ulysses Bernardet 2021 Expressing per- sonality through non-verbal behaviour in real-time interaction. Frontiers in psychology 12() 660895 Frontiers Media SA

  41. [41]

    ACM Transactions on Graphics (TOG) 36(1) 1–16 ACM New York, NY, USA

    Funda Durupinar, Mubbasir Kapadia, Susan Deutsch, Michael Neff, and Norman I Badler 2016 Perform: Perceptual approach for adding ocean personality to human motion using laban movement analysis. ACM Transactions on Graphics (TOG) 36(1) 1–16 ACM New York, NY, USA

  42. [42]

    ACM Transactions on Graphics (TOG) 40(1) 1–16 ACM New York, NY, USA

    Sinan Sonlu, Uğur Güdükbay, and Funda Durupinar 2021 A conversational agent framework with multi-modal personality expression. ACM Transactions on Graphics (TOG) 40(1) 1–16 ACM New York, NY, USA

  43. [43]

    In 2025 IEEE Conference Virtual Reality and 3D User Interfaces (VR) 718–728 IEEE

    Sinan Sonlu, Bennie Bendiksen, Funda Durupinar, and Uğur Güdükbay 2025 Effects of Embodiment and Personality in LLM-Based Conversational Agents. In 2025 IEEE Conference Virtual Reality and 3D User Interfaces (VR) 718–728 IEEE

  44. [44]

    IEEE Transactions on Visualization and Computer Graphics IEEE

    Haoyang Li, Zan Wang, Wei Liang, and Yizhuo Wang 2025 X’s Day: Personality- Driven Virtual Human Behavior Generation. IEEE Transactions on Visualization and Computer Graphics IEEE

  45. [45]

    Journal of artificial intelligence research 30() 457–500

    François Mairesse, Marilyn A Walker, Matthias R Mehl, and Roger K Moore 2007 Using linguistic cues for the automatic recognition of personality in conversation and text. Journal of artificial intelligence research 30() 457–500

  46. [46]

    Journal of personality and social psychology 77(6) 1296 American Psychological Association

    James W Pennebaker and Laura A King 1999 Linguistic styles: language use as an individual difference. Journal of personality and social psychology 77(6) 1296 American Psychological Association

  47. [47]

    European Journal of Social Psychology 8(4) 467–487 Wiley Online Library

    Klaus R Scherer 1978 Personality inference from voice quality: The loud voice of extroversion. European Journal of Social Psychology 8(4) 467–487 Wiley Online Library

  48. [48]

    American Psychological Association

    David Ed Matsumoto, Hyisung C Hwang, and Mark G Frank 2016 APA handbook of nonverbal communication. American Psychological Association

  49. [49]

    British Journal of Social and Clinical Psychology 17(1) 31–36 Wiley Online Library

    Anne Campbell and J Philippe Rushton 1978 Bodily communication and person- ality. British Journal of Social and Clinical Psychology 17(1) 31–36 Wiley Online Library

  50. [50]

    Routledge

    Michael Argyle 2013 Bodily communication. Routledge

  51. [51]

    In Individual differences in movement 27–41 Springer

    John Brebner 1985 Personality theory and movement. In Individual differences in movement 27–41 Springer

  52. [52]

    Applied Sciences 11(18) 8776 MDPI

    Sinae Lee, Jangwoon Park, and Dugan Um 2021 Speech characteristics as indica- tors of personality traits. Applied Sciences 11(18) 8776 MDPI

  53. [53]

    In International Workshop on Intelligent Virtual Agents 243–255 Springer

    Jina Lee and Stacy Marsella 2006 Nonverbal behavior generator for embodied conversational agents. In International Workshop on Intelligent Virtual Agents 243–255 Springer

  54. [54]

    In Close engagements with artificial companions: key social, psychological, ethical and design issues 143–156 John Benjamins Publishing Company

    Elisabetta Bevacqua, Ken Prepin, Radoslaw Niewiadomski, Etienne de Sevin, and Catherine Pelachaud 2010 Greta: Towards an interactive conversational virtual companion. In Close engagements with artificial companions: key social, psychological, ethical and design issues 143–156 John Benjamins Publishing Company

  55. [55]

    KODIS: A Multicultural Dispute Resolution Dialogue Corpus

    James Hale, Sushrita Rakshit, Kushal Chawla, Jeanne M Brett, and Jonathan Gratch 2025 KODIS: A Multicultural Dispute Resolution Dialogue Corpus. arXiv preprint arXiv:2504.12723

  56. [56]

    Personality and social psychology bulletin 23(4) 363–377 Sage Publications Sage CA: Los Angeles, CA

    Arthur Aron, Edward Melinat, Elaine N Aron, Robert Darrin Vallone, and Renee J Bator 1997 The experimental generation of interpersonal closeness: A procedure and some preliminary findings. Personality and social psychology bulletin 23(4) 363–377 Sage Publications Sage CA: Los Angeles, CA

  57. [57]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems 1–14

    Alex Wuqi Zhang, Ting-Han Lin, Xuan Zhao, and Sarah Sebo 2023 Ice-breaking technology: Robots and computers can foster meaningful connections between strangers through in-person conversations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems 1–14

  58. [58]

    Austin, TX: University of Texas at Austin 10() 1–47

    Ryan L Boyd, Ashwini Ashokkumar, Sarah Seraj, and James W Pennebaker 2022 The development and psychometric properties of LIWC-22. Austin, TX: University of Texas at Austin 10() 1–47

  59. [59]

    Psychological Bulletin 148(11-12) 843 American Psychological Association

    Antonis Koutsoumpis, Janneke K Oostrom, Djurre Holtrop, Ward Van Breda, Sina Ghassemi, and Reinout E de Vries 2022 The kernel of truth in text-based personality assessment: A meta-analysis of the relations between the Big Five and the Linguistic Inquiry and Word Count (LIWC). Psychological Bulletin 148(11-12) 843 American Psychological Association

  60. [60]

    Personality Trait Detection Using Bagged SVM over BERT Word Embedding Ensembles

    Amirmohammad Kazameini, Samin Fatehi, Yash Mehta, Sauleh Eetemadi, and Erik Cambria 2020 Personality trait detection using bagged svm over bert word embedding ensembles. arXiv preprint arXiv:2010.01309

  61. [61]

    Journal of personality and social psychology 94(2) 334 American Psychological Association

    Lisa A Fast and David C Funder 2008 Personality as manifest in word use: Correlations with self-report, acquaintance report, and behavior. Journal of personality and social psychology 94(2) 334 American Psychological Association. Han et al

  62. [62]

    Journal of language and social psychology 29(1) 24–54 Sage Publications Sage CA: Los Angeles, CA

    Yla R Tausczik and James W Pennebaker 2010 The psychological meaning of words: LIWC and computerized text analysis methods. Journal of language and social psychology 29(1) 24–54 Sage Publications Sage CA: Los Angeles, CA

  63. [63]

    Journal of research in personality 44(3) 363–373 Elsevier

    Tal Yarkoni 2010 Personality in 100,000 words: A large-scale analysis of per- sonality and word use among bloggers. Journal of research in personality 44(3) 363–373 Elsevier

  64. [64]

    Journal of research in Personality 41(1) 203–212 Elsevier

    Beatrice Rammstedt and Oliver P John 2007 Measuring personality in one minute or less: A 10-item short version of the Big Five Inventory in English and German. Journal of research in Personality 41(1) 203–212 Elsevier

  65. [65]

    Personality and Social Psychology Review 14(2) 196–213 Sage Publications Sage CA: Los Angeles, CA

    David A Kenny and Tessa V West 2010 Similarity and agreement in self-and other perception: A meta-analysis. Personality and Social Psychology Review 14(2) 196–213 Sage Publications Sage CA: Los Angeles, CA

  66. [66]

    Personality and Social Psychology Review 17(3) 248–272 Sage Publi- cations Sage CA: Los Angeles, CA

    Lauren J Human and Jeremy C Biesanz 2013 Targeting the good target: An integrative review of the characteristics and consequences of being accurately perceived. Personality and Social Psychology Review 17(3) 248–272 Sage Publi- cations Sage CA: Los Angeles, CA

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.