Pith. sign in

REVIEW 2 major objections 5 minor 94 references

Users trusted and were more persuaded by chatbots that agreed with their opinions. Matching a chatbot's personality to the user produced no benefit, and introvert users rated the introverted chatbot lower on trust and competence.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection Opinion alignment reliably moves trust and satisfaction in a large pre-registered experiment; the personality-alignment null is undermined by a style confound and an internal inconsistency in the trust results. the 2 major comments →

arxiv 2511.10544 v3 pith:P3VEWBKQ submitted 2025-11-13 cs.HC

Effects of Personality- and Opinion-Alignment in Human-AI Interaction

classification cs.HC
keywords human-AI interactionAI personalizationsimilarity-attractionopinion alignmentpersonality alignmentLLM persuasionuser trustonline experiment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that opinion alignment — whether the AI's expressed stance matches the user's — is a consistent and comparatively strong driver of how users evaluate an AI assistant, while personality alignment is weak at best and can even backfire. In a pre-registered online experiment, 1,000 participants each chatted for several minutes with a chatbot prompted to take an affirmative or critical stance on a topical issue and to behave as an extrovert or introvert. Users rated chatbots that shared their pre-talk opinion as more competent, trustworthy, warm, satisfying, and persuasive, on both sides of the opinion split. Personality alignment showed no equivalent pattern: the extroverted chatbot was rated about equally by everyone, while introverted participants rated the introverted chatbot lower on trust and competence — the opposite of what similarity-attraction would predict. A sympathetic reader would care because AI assistants are being personalized at scale, and this is direct evidence about which form of personalization actually moves user trust and which one may be hurting it.

Core claim

Central claim: similarity attraction, documented in human relationships, transfers to human-AI interaction through opinions but not through personality. Participants rated chatbots sharing their pre-interaction opinion higher on competence, trust, warmth, satisfaction, and persuasion, and joint regressions showed opinion alignment outranked personality alignment on every outcome. The personality results contradicted the human pattern: the extroverted chatbot was rated equally by everyone, while introverted participants rated the introverted chatbot lower on trust and competence — a similarity penalty. The paper concludes that personalized AI does not work uniformly and that how deployed mode

What carries the argument

A 2×2 between-subjects factorial design randomly assigns a chatbot's opinion stance (affirmative vs. critical) and personality (extroverted vs. introverted) via system prompts, crossed with each participant's measured extroversion and pre-task opinion. The load-bearing comparison is the alignment correlation — how strongly a participant's evaluation tracks whether the assigned chatbot matched them on each dimension, estimated by regression and compared in a joint model. The gap between opinion-alignment and personality-alignment coefficients is positive, significant, and consistent across all outcomes; it is what lets the paper claim opinions, not personalities, are the operative lever in pe

Load-bearing premise

The load-bearing premise is that the two manipulations were clean and independent: the introverted chatbot's style cues (less positive emotion, fewer social-process words, more concrete language) did not lower ratings for everyone. If they did, the introvert-users-distrust-introverted-bots result is a style artifact, and the opinion results hold only if stance, not argument quality, drove them.

What would settle it

Run the same conversations with the introverted and extroverted conditions matched on warmth cues — equal positive emotion and social-process language, differing only in asserted sociability. If introvert users' trust and competence ratings of the introverted chatbot rise to parity with extroverts' ratings, the personality reversal is a style artifact; if the reversal persists with warmth held constant, personality alignment genuinely repels introvert users.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Chatbots that echo a user's stated position should reliably raise the user's ratings of competence, trust, warmth, satisfaction, and persuasiveness in short topical conversations.
  • Prompting a chatbot with an introverted persona can lower trust and competence ratings among introverted users, so personality personalization is not a harmless no-op even when its average effect is nil.
  • Because opinion alignment is the stronger lever, user-facing personalization should be studied and overseen with opinion mirroring treated as the primary mechanism of effect.
  • The persuasive boost from agreement implies a concrete risk: assistants optimized for user satisfaction may converge on agreeing with users, reinforcing existing views instead of broadening them.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the introvert result cannot be cleanly attributed to trait alignment until style cues are untangled; the introverted condition was engineered with less positive emotion, fewer social-process words, and more concrete language, so the decisive check would be an introverted condition matched to the extroverted one on warmth and emotional tone.
  • Editorial inference: because 'persuaded' ratings tracked agreement, future measurements should separate perceived persuasiveness from actual attitude shift — the two are likely conflated when the model mostly reaffirms the user's position.
  • Editorial inference: if satisfaction optimization rewards opinion mirroring, deployed assistants may drift toward unconditional agreement; logging how often personalized assistants actually disagree with users, alongside satisfaction scores, would turn that risk into a measurable quantity.
  • Editorial inference: the null extroversion result leaves open the possibility that other personality dimensions, such as agreeableness, where conversational tone and warmth are less entangled with the trait itself, might still show genuine alignment effects.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports a pre-registered, between-subjects online experiment (N=1,000) in which participants discussed a controversial topic with an LLM assistant assigned one of two personality styles (extroverted vs. introverted) and one of two opinion stances (affirmative vs. critical). The central claims are that opinion alignment robustly increases perceived competence, trust, warmth, satisfaction, and persuasiveness, while personality alignment shows no or weak effects, with introvert participants rating introvert models lower on trust and competence. The authors interpret the first pattern as support for an AI similarity-attraction hypothesis and the second pattern as evidence that personality similarity does not transfer straightforwardly to human-AI interaction.

Significance. If the opinion-alignment finding holds, it is a valuable and policy-relevant demonstration—obtained in a large, pre-registered, controlled setting—that expressed opinion similarity is a strong lever on user trust and evaluation of AI assistants, and that it is substantially stronger than personality matching in this context. The study has notable methodological strengths: preregistration (AsPredicted #239571), a power analysis, N=1,000, standardized BFI-2 measurement of extroversion, and consistent opinion-alignment correlations across several outcome measures. However, the personality manipulation is confounded with linguistic style, and the manuscript does not currently provide the promised treatment checks needed to rule out a style-penalty alternative explanation for the personality result. The opinion-alignment result is less vulnerable because the cross-over interaction between participant stance and model stance cannot be explained by a simple main effect of argument quality.

major comments (2)
  1. [§3.2 and §4.4] The personality manipulation is confounded with affective and conversational style. The introvert prompt intentionally used 'less positive emotion and fewer social process words' and more concrete language, while the extrovert prompt used the opposite. These cues plausibly affect perceived warmth, engagement, and competence directly, independent of trait alignment. The pattern that introvert participants rated the introvert model less trustworthy (r = 0.19), less competent (r = 0.18), and less satisfying (r = 0.24) could therefore reflect an interaction between participant personality and linguistic style (e.g., introverts reacting more negatively to a cold, low-engagement style) rather than a failure of personality similarity-attraction. Because H1/H2 and the second headline claim depend on a clean trait manipulation, the treatment checks in §4.4 (model BFI self-reports, simulated conve
  2. [§4.1, second paragraph] There is an internal inconsistency in the trust results. The text states 'There was no significant difference in overall trust in the affirmative (M=3.41) and critical model (M=3.88)', but Figure 4 displays means of 3.41 and 3.88 respectively—a gap of 0.47 on a 5-point scale. With roughly 500 participants per condition, such a difference would ordinarily be highly significant; the sentence is likely an error, perhaps intended to compare extrovert and introvert models. Please provide the correct null-comparison statistics and reconcile the text with the figure. This matters because readers cannot currently verify the reported null main effect.
minor comments (5)
  1. [§4.4] The treatment checks are referenced but not included in the provided text. Please include the full results, including model BFI-2 extroversion scores, the simulated-conversation evaluations, and any checks that the opinion manipulation produced the intended stance.
  2. [Figures 3–6] The figures contain repeated caption/annotation blocks with identical statistics, which appears to be a rendering artifact in the provided manuscript. Please clean up the figure rendering so each panel shows a single annotation.
  3. [§3.3.3] Please report reliability coefficients (e.g., Cronbach's alpha or McDonald's omega) for the adapted scales measuring competence, warmth, satisfaction, persuasiveness, and perceived homophily, in addition to the BFI-2 extroversion scale.
  4. [§3.2] The base model is described as 'GPT5-chat-latest'; please specify the exact model version and access date, since LLM behavior changes over time and reproducibility depends on this information.
  5. [§4.1] The 'joint standardized regression analysis' used to compare opinion-alignment and personality-alignment correlations is not fully specified. Please describe the model—predictors, coding of alignment, cluster/robust errors if used—and consider reporting the full table.

Circularity Check

0 steps flagged

No significant circularity: a pre-registered between-subjects experiment whose reported effects are measured correlations, not constructions.

full rationale

This paper is an empirical experiment (N=1,000, 2x2 between-subjects design), not a derivation, so the circularity patterns for self-definitional or fitted-input predictions largely do not apply. The outcome measures (competence, trust, warmth, satisfaction, persuasiveness) are standard published scales; the predictors are experimentally assigned prompt conditions plus pre-task participant measures (BFI-2 extroversion, opinion Likert items). No parameter is fitted to the reported correlations and then renamed as a prediction: the r values (e.g., r=0.23 for opinion alignment) are simple descriptive correlations between pre-task measures and post-task ratings, and H1-H3 were pre-registered (AsPredicted #239571) before data collection. The similarity-attraction benchmark is external to this paper, and no load-bearing claim reduces to a self-citation: the cited works (Jiang et al., Salvi et al., Nass and Lee) are used for prompt-engineering technique and topic selection, not to define the outcome. The only passage worth flagging is section 3.2 / section 4.4: the treatment check asks the model to self-report on the same twelve BFI-2 extroversion items whose keywords were inserted into the system prompt, and the section 4.4 results are referenced but not present in the provided text. That is a manipulation-check design choice and a construct-validity concern (introversion cued via less positive emotion and fewer social process words and more concrete language could act as a style penalty rather than a trait manipulation), but it is not circular: user ratings are independent of the model's self-report, and no equation makes the claimed preference effects definitional of their inputs. Verdict: no significant circularity; score 1 reflecting the mild, non-load-bearing treatment-check note.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

No data-fitting parameters exist: the design is an experiment with pre-registered hypotheses; the power analysis uses an assumed small effect d=0.2, and the topic pool's 'strength score' clustering is borrowed from Salvi et al., neither of which is a fitted constant for the claims. The axioms listed are the measurement and manipulation assumptions the causal claims inherit. No invented entities: 'AI personality' and 'AI opinion' are prompt manipulations of an existing model, not new constructs with external falsifiable handles (perceived similarity is a measured mediator, not an entity).

axioms (5)
  • domain assumption BFI-2 extroversion subscale validly measures extraversion for both human self-report and LLM self-assessment
    Used for participant screening and treatment checks alike (§3.3.1, §4.4); if LLM BFI responses do not track the expressed persona, the manipulation check is void.
  • ad hoc to paper System prompts reliably induce the assigned opinion and personality without unrelated quality differences
    Central to the 2x2 design (§3.2); introvert cues (fewer positive-emotion/social words, concrete language) may themselves lower perceived competence/warmth, confounding the personality condition.
  • domain assumption GPT-4o-generated, manually rephrased opinion items validly measure topic attitudes and their pre-post change
    Opinion-alignment correlations and any persuasion claims depend on these four items per topic (§3.3.2).
  • domain assumption The five selected topics afford balanced, non-trivial positions in a UK sample
    Topics were drawn from Salvi et al.'s medium-strength clusters and US-centric topics removed (§3.4); topic effects on argument strength are not modeled.
  • standard math Linear regression/correlation inference assumptions (approximate linearity, independence) hold for Likert-scale outcomes
    All alignment claims are tested as correlations between participant traits and 5-point Likert ratings within conditions (§4.1).

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Effects of Personality- and Opinion-Alignment in Human-AI Interaction." pith.science (2026). https://pith.science/paper/P3VEWBKQ

@misc{pith2026251110544,
  author       = {Pith},
  title        = {Pith review of: Effects of Personality- and Opinion-Alignment in Human-AI Interaction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P3VEWBKQ}},
  note         = {Machine review of arXiv:2511.10544}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Interactions with AI assistants are increasingly personalized to individual users. As AI personalization is dynamic and machine-learning-driven, we have limited understanding of how personalization affects interaction outcomes and user perceptions. We conducted a large-scale controlled experiment in which 1,000 participants interacted with AI assistants prompted to take on specific personality traits and opinions. Our results show that participants consistently preferred to interact with models that shared their opinions. Participants found opinion-aligned models more trustworthy, competent, warm, and persuasive, corroborating an AI-similarity-attraction hypothesis. In contrast, we observed no or only weak effects of AI personality alignment, with introvert models rated as less trustworthy and competent by introvert participants. These findings highlight opinion alignment as a central dimension of AI user preference, while underscoring the need for a more grounded discussion of the mechanisms and risks of AI personalization.

Figures

Figures reproduced from arXiv: 2511.10544 by Clemens Lechner, Maurice Jakesch, Maximilian Eder.

Figure 1
Figure 1. Figure 1: Study Design Overview: Participants (N=1,000) engaged in a topic discussion with an AI assistant that was experimentally assigned an opinion (critical or affirmative) and personality trait (extroverted or introverted). Participants’ own personality traits and opinions on the topic were collected prior to interaction, allowing us to analyze to what extent the alignment between the participant’s and AI’s per… view at source ↗
Figure 2
Figure 2. Figure 2: Study flow and experimental design. Participants first completed a demographics and personality questionnaire (BFI-2) and answered a pre-interaction opinion questionnaire assessing their opinions towards a randomly chosen topic. Next, they discussed the topic with an AI assistant with a random personality (extroverted or introverted) and opinion (affirmative or critical). After the interaction, participant… view at source ↗
Figure 3
Figure 3. Figure 3: Perceived model competence against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. Participants rated the extrovert model as equally competent, but introvert participants perceived the introvert model as less competent. Participants rated a model sharing their opinion as more competent. M = 3.42 r = −0.04 M = 3.37 r = 0.19** Model persona… view at source ↗
Figure 4
Figure 4. Figure 4: Participants’ trust in the model against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. Mean correlation coefficients, and significance (*p<0.05, **p<0.01, ***p<0.001) are shown in the info box on the bottom left. All participants trusted the extrovert equally, but introvert participants trusted the introvert model as less than extrovert… view at source ↗
Figure 5
Figure 5. Figure 5: User experience: Satisfaction with the model against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. All participants were equally satisfied with the extrovert model, but introvert participants were less satisfied with the introvert model than extroverts. Participants were generally more satisfied with a model that echoed their own opinio… view at source ↗
Figure 6
Figure 6. Figure 6: User experience: Perceived warmth of the model against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. All participants were equally satisfied with the extrovert model, but introvert participants were less satisfied with the introvert model than extroverts. Participants were generally more satisfied with a model that echoed their own opin… view at source ↗
Figure 7
Figure 7. Figure 7: Perceived persuasiveness of the model against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. Participants perceived the extrovert and introvert model to be equally persuasive. However, participants generally found models more persuasive that agreed with their own opinions. M = 0.198 r = 0.001 M = 0.200 r = 0.017 Model personality: Extrov… view at source ↗
Figure 8
Figure 8. Figure 8: Participant opinion changes: Observed persuasiveness of the model against personality (left) and opinion (right) traits with a fitted linear regression line and 95% confidence bands. N=1000. Participants were equally swayed by the extrovert and introvert model. However, participants generally changed their opinions more after interacting with a model they disagreed with. participants with a critical pre-ta… view at source ↗
Figure 9
Figure 9. Figure 9: Chatbot Interface. The interface of the chatbot replicated the design of wide-spread AI assistant web application. 6.2 Experience Rating Scales Participants responded to the following items using a 5-point Likert scale (1 = Strongly Disagree, 5 = Strongly Agree) [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Full System Prompt. The system prompt starts with a section on the general context, chosen topic and assigned opinion. Next, the assigned personality is described by both personality traits derived from the Big Five Inventory and language cues form prior research. Lastly, the model receives general conversation instructions, in order to maintain a natural conversation [PITH_FULL_IMAGE:figures/full_fig_p0… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

94 extracted references · 16 canonical work pages · 1 internal anchor

  1. [1]

    Anthropic. 2025. Claude. https://claude.ai/. 20 Maximilian Eder, Clemens Lechner, and Maurice Jakesch

  2. [2]

    APA Dictionary of Psychology. 2025. APA Dictionary of Psychology. https://dictionary.apa.org/

  3. [3]

    Voelkel, Shane Muldowney, Johannes C

    Hui Bai, Jan G. Voelkel, Shane Muldowney, Johannes C. Eichstaedt, and Robb Willer. 2025. LLM-generated Messages Can Persuade Humans on Policy Issues.Nature Communications16, 1 (July 2025), 6037. doi:10.1038/s41467-025-61345-5

  4. [4]

    Voelkel, johannes C

    (Max) Hui Bai, Jan G. Voelkel, johannes C. Eichstaedt, and Robb Willer. 2023. Artificial Intelligence Can Persuade Humans on Political Issues. doi:10.31219/osf.io/stakv

  5. [5]

    1966.The Duality of Human Existence: An Essay on Psychology and Religion

    David Bakan. 1966.The Duality of Human Existence: An Essay on Psychology and Religion. Rand Mcnally, Oxford, England. 242 pages

  6. [6]

    Beukeboom, Martin Tanis, and Ivar E

    Camiel J. Beukeboom, Martin Tanis, and Ivar E. Vermeulen. 2013. The Language of Extraversion: Extraverted People Talk More Abstractly, Introverts Are More Concrete.Journal of Language and Social Psychology32, 2 (2013), 191–201. doi:10.1177/0261927X12460844

  7. [7]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stef...

  8. [8]

    Michelle Brachman, Amina El-Ashry, Casey Dugan, and Werner Geyer. 2025. Current and Future Use of Large Language Models for Knowledge Work. arXiv:2503.16774 [cs] doi:10.48550/arXiv.2503.16774

  9. [9]

    Donn Byrne. 1961. Interpersonal attraction and attitude similarity.The journal of abnormal and social psychology62, 3 (1961), 713

  10. [10]

    Donn Erwin Byrne. 1972. The Attraction Paradigm.Behavior Therapy3, 2 (1972), 337–338

  11. [11]

    María Victoria Carro. 2024. Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model. arXiv:2412.02802 [cs] doi:10.48550/arXiv.2412.02802

  12. [12]

    Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, Kai Zheng, Defu Lian, and Enhong Chen. 2024. When Large Language Models Meet Personalization: Perspectives of Challenges and Opportunities.World Wide Web 27, 4 (July 2024), 42. doi:10.1007/s11280-024-01276-1

  13. [13]

    Jiayu Chen, Lin Qiu, and Moon-Ho Ringo Ho. 2020. A Meta-Analysis of Linguistic Markers of Extraversion: Positive Emotion and Social Process Words.Journal of Research in Personality89 (Dec. 2020), 104035. doi:10.1016/j.jrp.2020.104035

  14. [14]

    Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, and Dan Jurafsky. 2025. Social Sycophancy: A Broader Understanding of LLM Sycophancy. arXiv:2505.13995 [cs] doi:10.48550/arXiv.2505.13995

  15. [15]

    Hans Christian, Derwin Suhartono, Andry Chowanda, and Kamal Z. Zamli. 2021. Text Based Personality Prediction from Multiple Social Media Data Sources Using Pre-Trained Language Model and Model Averaging.Journal of Big Data8, 1 (May 2021), 68. doi:10.1186/s40537-021-00459-1

  16. [16]

    Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1...

  17. [17]

    Clerke and Erin A

    Alexa S. Clerke and Erin A. Heerey. 2021. The Influence of Similarity and Mimicry on Decisions to Trust.Collabra: Psychology7, 1 (May 2021), 23441. arXiv:https://online.ucpress.edu/collabra/article-pdf/7/1/23441/834082/collabra_2021_7_1_23441.pdf doi:10.1525/collabra.23441

  18. [18]

    Thomas H Costello, Gordon Pennycook, and David G Rand. 2024. Durably reducing conspiracy beliefs through dialogues with AI.Science385, 6714 (2024), eadq1814

  19. [19]

    Lei Cui, Shaohan Huang, Furu Wei, Chuanqi Tan, Chaoqun Duan, and Ming Zhou. 2017. SuperAgent: A Customer Service Chatbot for E-commerce Websites. InProceedings of ACL 2017, System Demonstrations, Mohit Bansal and Heng Ji (Eds.). Association for Computational Linguistics, Vancouver, Canada, 97–102

  20. [20]

    Navdeep Dhillon and Gurvinder Kaur. 2023. Impact of Personality Traits on Communication Effectiveness of Teachers: Exploring the Mediating Role of Their Communication Style.SAGE Open13, 2 (April 2023), 21582440231168049. doi:10.1177/21582440231168049

  21. [21]

    Cambridge Dictionary. 2025. Personality. https://dictionary.cambridge.org/dictionary/english/personality

  22. [22]

    D Christopher Dryer and Leonard M Horowitz. 1997. When do opposites attract? Interpersonal complementarity versus similarity.Journal of personality and social psychology72, 3 (1997), 592

  23. [23]

    Esin Durmus, Liane Lovitt, Alex Tamkin, Stuart Ritchie, Jack Clark, and Deep Ganguli. 2024. Measuring the Persuasiveness of Language Models. https://www.anthropic.com/news/measuring-model-persuasiveness. Bytes of a Feather: Personality and Opinion Alignment Effects in Human-AI Interaction 21

  24. [24]

    Benj Edwards. 2023. AI-powered Bing Chat Gains Three Distinct Personalities. https://arstechnica.com/information-technology/2023/03/microsoft- equips-bing-chat-with-multiple-personalities-creative-balanced-precise/

  25. [25]

    H.J. Eysenck. 1967.The Biological Basis of Personality. Transaction Publishers

  26. [26]

    Agarwal, Joanna Lin, Anson Zhou, Roxana Daneshjou, and Sanmi Koyejo

    Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin, Anson Zhou, Roxana Daneshjou, and Sanmi Koyejo. 2025. SycEval: Evaluating LLM Sycophancy. arXiv:2502.08177 [cs] doi:10.48550/arXiv.2502.08177

  27. [27]

    Tommaso Fornaciari, Fabio Celli, and Massimo Poesio. 2013. The Effect of Personality Type on Deceptive Communication Style. In2013 European Intelligence and Security Informatics Conference. 1–6. doi:10.1109/EISIC.2013.8

  28. [28]

    Markus Freitag and Paul. C. Bauer. 2016. Personality Traits and the Propensity to Trust Friends and Strangers.The Social Science Journal53, 4 (Dec. 2016), 467–476. doi:10.1016/j.soscij.2015.12.002

  29. [29]

    Alastair James Gill. 2004. Personality and Language: The Projection and Perception of Personality in Computer-Mediated Communication. (2004)

  30. [30]

    Description of Personality

    L. R. Goldberg. 1990. An Alternative "Description of Personality": The Big-Five Factor Structure.Journal of Personality and Social Psychology59, 6 (Dec. 1990), 1216–1229. doi:10.1037//0022-3514.59.6.1216

  31. [31]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2024. How Persuasive Is AI-generated Propaganda?PNAS Nexus 3, 2 (Feb. 2024), pgae034. doi:10.1093/pnasnexus/pgae034

  32. [32]

    Google. 2025. Gemini. https://gemini.google.com/

  33. [33]

    Andrew Guess, Brendan Nyhan, Benjamin Lyons, and Jason Reifler. 2018. Avoiding the echo chamber about echo chambers.Knight Foundation2, 1 (2018), 1–25

  34. [34]

    Michael B. Gurtman. 2009. Exploring Personality with the Interpersonal Circumplex.Social and Personality Psychology Compass3, 4 (2009), 601–619. doi:10.1111/j.1751-9004.2009.00172.x

  35. [35]

    Matthew B. Hoy. 2018. Alexa, Siri, Cortana, and More: An Introduction to Voice Assistants.Medical Reference Services Quarterly37, 1 (Jan. 2018), 81–88. doi:10.1080/02763869.2018.1404391

  36. [36]

    Dipika Jain, Akshi Kumar, and Rohit Beniwal. 2022. Personality BERT: A Transformer-Based Model for Personality Detection from Textual Data. InProceedings of International Conference on Computing and Communication Networks, Ali Kashif Bashir, Giancarlo Fortino, Ashish Khanna, and Deepak Gupta (Eds.). Springer Nature, Singapore, 515–522. doi:10.1007/978-981...

  37. [37]

    Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023. Co-Writing with Opinionated Language Models Affects Users’ Views. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–15. doi:10.1145/3544548.3581196

  38. [38]

    Hancock, and Mor Naaman

    Maurice Jakesch, Jeffrey T. Hancock, and Mor Naaman. 2023. Human Heuristics for AI-generated Language Are Flawed.Proceedings of the National Academy of Sciences120, 11 (March 2023), e2208839120. doi:10.1073/pnas.2208839120

  39. [39]

    Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu. 2023. Evaluating and Inducing Personality in Pre-trained Language Models. arXiv:2206.07550 [cs] doi:10.48550/arXiv.2206.07550

  40. [40]

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits. arXiv:2305.02547 [cs]

  41. [41]

    Julie Jiang, Xiang Ren, and Emilio Ferrara. 2021. Social Media Polarization and Echo Chambers in the Context of COVID-19: Case Study.JMIRx Med2, 3 (Aug. 2021), e29570. doi:10.2196/29570

  42. [42]

    Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T. Hancock. 2023. Working With AI to Persuade: Examining a Large Language Model’s Ability to Generate Pro-Vaccination Messages.Proc. ACM Hum.-Comput. Interact.7, CSCW1 (April 2023), 116:1–116:29. doi:10.1145/3579592

  43. [43]

    Woo Bin Kim and Hee Jin Hur. 2024. What Makes People Feel Empathy for AI Chatbots? Assessing the Role of Competence and Warmth. International Journal of Human–Computer Interaction40, 17 (Sept. 2024), 4674–4687. doi:10.1080/10447318.2023.2219961

  44. [44]

    Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale. 2024. The Benefits, Risks and Bounds of Personalizing the Alignment of Large Language Models to Individuals.Nature Machine Intelligence6, 4 (April 2024), 383–392. doi:10.1038/s42256-024-00820-y

  45. [45]

    Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Varun Gupta, Iulian Vlad Serban, and Joelle Pineau. 2020. Automated Personalized Feedback Improves Learning Gains in An Intelligent Tutoring System. InArtificial Intelligence in Education, Ig Ibert Bittencourt, Mutlu Cukurova, Kasia Muldner, Rose Luckin, and Eva Millán (Eds.). Vol. 12164. Springer Internationa...

  46. [46]

    Miles McCain, and Miles Brundage

    Sarah Kreps, R. Miles McCain, and Miles Brundage. 2022. All the News That’s Fit to Fabricate: AI-Generated Text as a Tool of Media Misinformation. Journal of Experimental Political Science9, 1 (March 2022), 104–117. doi:10.1017/XPS.2020.37

  47. [47]

    1957.Interpersonal Diagnosis of Personality: A Functional Theory and Methodology for Personality Evaluation.Ronald Press Co., New York

    Timothy Leary. 1957.Interpersonal Diagnosis of Personality: A Functional Theory and Methodology for Personality Evaluation.Ronald Press Co., New York

  48. [48]

    Kibeom Lee and Michael C. Ashton. 2004. Psychometric Properties of the HEXACO Personality Inventory.Multivariate Behavioral Research39, 2 (April 2004), 329–358. doi:10.1207/s15327906mbr3902_8

  49. [49]

    Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, and Youngjae Yu. 2024. Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset Designed for LLMs with Psychometrics. arXiv:2406.14703 [cs] doi:10.48550/arXiv.2406.14703

  50. [50]

    Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. 2024. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods. arXiv:2412.05579 [cs] doi:10.48550/arXiv.2412.05579 22 Maximilian Eder, Clemens Lechner, and Maurice Jakesch

  51. [51]

    Yannakakis, and Julian Togelius

    Jialin Liu, Sam Snodgrass, Ahmed Khalifa, Sebastian Risi, Georgios N. Yannakakis, and Julian Togelius. 2021. Deep Learning for Procedural Content Generation.Neural Computing and Applications33, 1 (Jan. 2021), 19–37. arXiv:2010.04548 [cs] doi:10.1007/s00521-020-05383-8

  52. [52]

    Mairesse, M

    F. Mairesse, M. A. Walker, M. R. Mehl, and R. K. Moore. 2007. Using Linguistic Cues for the Automatic Recognition of Personality in Conversation and Text.Journal of Artificial Intelligence Research30 (Nov. 2007), 457–500. doi:10.1613/jair.2349

  53. [53]

    Whiteman

    Gerald Matthews, Ian Deary, and M.C. Whiteman. 2009. Personality traits, Third edition.Personality Traits, Third Edition(01 2009), 1–568. doi:10.1017/CBO9780511812743

  54. [54]

    S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari, and M. Cerf. 2024. The Potential of Generative AI for Personalized Persuasion at Scale. Scientific Reports14, 1 (Feb. 2024), 4692. doi:10.1038/s41598-024-53755-0

  55. [55]

    Kiran McCloskey and Blair T. Johnson. 2021. You Are What You Repeatedly Do: Links between Personality and Habit.Personality and Individual Differences181 (Oct. 2021), 111000. doi:10.1016/j.paid.2021.111000

  56. [56]

    R. R. McCrae and O. P. John. 1992. An Introduction to the Five-Factor Model and Its Applications.Journal of Personality60, 2 (June 1992), 175–215. doi:10.1111/j.1467-6494.1992.tb00970.x

  57. [57]

    James Mccroskey, Virginia Richmond, and John Daly. 1975. The Development of a Measure of Perceived Homophily.Human Communication Interaction1 (June 1975), 323–332. doi:10.1111/j.1468-2958.1975.tb00281.x

  58. [58]

    McGrath, Oliver Lack, James Tisch, and Andreas Duenser

    Melanie J. McGrath, Oliver Lack, James Tisch, and Andreas Duenser. 2025. Measuring Trust in Artificial Intelligence: Validation of an Established Scale and Its Short Form.Frontiers in Artificial Intelligence8 (May 2025), 1582880. doi:10.3389/frai.2025.1582880

  59. [59]

    Jonker, and Myrthe L

    Siddharth Mehrotra, Catholijn M. Jonker, and Myrthe L. Tielman. 2021. More Similar Values, More Trust? - The Effect of Value Similarity on Trust in Human-Agent Interaction. InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’21). Association for Computing Machinery, New York, NY, USA, 777–783. doi:10.1145/3461702.3462576

  60. [60]

    Joonas Moilanen, Aku Visuri, Sharadhi Alape Suryanarayana, Andy Alorwu, Koji Yatani, and Simo Hosio. 2022. Measuring the Effect of Mental Health Chatbot Personality on User Engagement. InProceedings of the 21st International Conference on Mobile and Ubiquitous Multimedia. ACM, Lisbon Portugal, 138–150. doi:10.1145/3568444.3568464

  61. [61]

    Cecilie Grace Møller, Ke En Ang, María De Lourdes Bongiovanni, Md Saifuddin Khalid, and Jiayan Wu. 2024. Metrics of Success: Evaluating User Satisfaction in AI Chatbots. InProceedings of the 2024 8th International Conference on Advances in Artificial Intelligence. ACM, London United Kingdom, 168–173. doi:10.1145/3704137.3704182

  62. [62]

    Matthew Montoya and Robert S

    R. Matthew Montoya and Robert S. Horton. 2013. A Meta-Analytic Investigation of the Processes Underlying the Similarity-Attraction Effect. Journal of Social and Personal Relationships30, 1 (Feb. 2013), 64–94. doi:10.1177/0265407512452989

  63. [63]

    Matthew Montoya, Robert S

    R. Matthew Montoya, Robert S. Horton, and Jeffrey Kirchner. 2008. Is Actual Similarity Necessary for Attraction? A Meta-Analysis of Actual and Perceived Similarity.Journal of Social and Personal Relationships25, 6 (Dec. 2008), 889–922. doi:10.1177/0265407508096700

  64. [64]

    Shannon Moore, Bert Uchino, Brian Baucom, Arwen Behrends, and David Sanbonmatsu. 2017. Attitude Similarity and Familiarity and Their Links to Mental Health: An Examination of Potential Interpersonal Mediators.The Journal of social psychology157, 1 (2017), 77–85. doi:10.1080/00224545. 2016.1176551

  65. [65]

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2024. More Human than Human: Measuring ChatGPT Political Bias.Public Choice198, 1 (Jan. 2024), 3–23. doi:10.1007/s11127-023-01097-2

  66. [66]

    Clifford Nass and Kwan Min Lee. 2000. Does Computer-Generated Speech Manifest Personality? An Experimental Test of Similarity-Attraction. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, The Hague The Netherlands, 329–336. doi:10.1145/332040.332452

  67. [67]

    Jan Nehring, Aleksandra Gabryszak, Pascal Jürgens, Aljoscha Burchardt, Stefan Schaffer, Matthias Spielkamp, and Birgit Stark. 2024. Large Language Models Are Echo Chambers. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique...

  68. [68]

    Neimeyer and Kelly A

    Robert A. Neimeyer and Kelly A. Mitchell. 1988. Similarity and Attraction: A Longitudinal Study.Journal of Social and Personal Relationships5, 2 (1988), 131–148. doi:10.1177/026540758800500201

  69. [69]

    Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. HONEST: Measuring Hurtful Sentence Completion in Language Models. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven ...

  70. [70]

    OpenAI. 2025. ChatGPT. https://chatgpt.com/

  71. [71]

    OpenAI. 2025. GPT-4o. https://platform.openai.com/docs/models/gpt-4o Model identifier:gpt-4o

  72. [72]

    OpenAI. 2025. GPT-5 Chat (latest). https://platform.openai.com/docs/models/gpt-5-chat-latest Model identifier:gpt-5-chat-latest

  73. [73]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Human Fee...

  74. [74]

    Stefan Palan and Christian Schitter. 2018. Prolific.Ac—A Subject Pool for Online Experiments.Journal of Behavioral and Experimental Finance17 (March 2018), 22–27. doi:10.1016/j.jbef.2017.12.004

  75. [75]

    Keyu Pan and Yawen Zeng. 2023. Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models. arXiv:2307.16180 [cs] doi:10.48550/arXiv.2307.16180 Bytes of a Feather: Personality and Opinion Alignment Effects in Human-AI Interaction 23

  76. [76]

    Schwartz, Johannes Eichstaedt, Margaret Kern, Michal Kosinski, David Stillwell, Lyle Ungar, and Martin Seligman

    Gregory Park, H. Schwartz, Johannes Eichstaedt, Margaret Kern, Michal Kosinski, David Stillwell, Lyle Ungar, and Martin Seligman. 2014. Automatic Personality Assessment Through Social Media Language.Journal of personality and social psychology108 (Nov. 2014). doi:10.1037/pspp0000020

  77. [77]

    Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier

    Max Pellert, Clemens M. Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. 2024. AI Psychometrics: Assessing the Psychological Profiles of Large Language Models Through Psychometric Inventories.Perspectives on Psychological Science19, 5 (Sept. 2024), 808–826. doi:10.1177/ 17456916231214460

  78. [78]

    Sarah Perez. 2025. ChatGPT Doubled Its Weekly Active Users in under 6 Months, Thanks to New Releases

  79. [79]

    Wallace, Vanessa Sawicki, Kathleen M

    Aviva Philipp-Muller, Laura E. Wallace, Vanessa Sawicki, Kathleen M. Patton, and Duane T. Wegener. 2020. Understanding When Similarity- Induced Affective Attraction Predicts Willingness to Affiliate: An Attitude Strength Perspective.Frontiers in Psychology11 (Aug. 2020), 1919. doi:10.3389/fpsyg.2020.01919

  80. [80]

    Crandall, Nicholas A

    Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, and Matthew O. Jackson. 2019. Machine Behaviour.Nature568, 7753 (2019), 477–486

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.