Pith. sign in

REVIEW 7 cited by

Large Language Models Can Infer Personality from Free-Form User Interactions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13052 v1 pith:RJA7GJJM submitted 2024-05-19 cs.HC cs.AIcs.CLcs.CYcs.LG

Large Language Models Can Infer Personality from Free-Form User Interactions

classification cs.HC cs.AIcs.CLcs.CYcs.LG
keywords personalityinferencesinteractionsuseraccuracyacrosschatbotinfer
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This study investigates the capacity of Large Language Models (LLMs) to infer the Big Five personality traits from free-form user interactions. The results demonstrate that a chatbot powered by GPT-4 can infer personality with moderate accuracy, outperforming previous approaches drawing inferences from static text content. The accuracy of inferences varied across different conversational settings. Performance was highest when the chatbot was prompted to elicit personality-relevant information from users (mean r=.443, range=[.245, .640]), followed by a condition placing greater emphasis on naturalistic interaction (mean r=.218, range=[.066, .373]). Notably, the direct focus on personality assessment did not result in a less positive user experience, with participants reporting the interactions to be equally natural, pleasant, engaging, and humanlike across both conditions. A chatbot mimicking ChatGPT's default behavior of acting as a helpful assistant led to markedly inferior personality inferences and lower user experience ratings but still captured psychologically meaningful information for some of the personality traits (mean r=.117, range=[-.004, .209]). Preliminary analyses suggest that the accuracy of personality inferences varies only marginally across different socio-demographic subgroups. Our results highlight the potential of LLMs for psychological profiling based on conversational interactions. We discuss practical implications and ethical challenges associated with these findings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Are LLM Inferences Acceptable? User Reactions and Control Preferences for Inferred Personal Information

    cs.HC 2026-05 unverdicted novelty 7.0

    Users show curiosity over concern toward LLM inferences of personal information, with acceptability depending on context, alignment with expectations, and who uses the inferences rather than just the content.

  2. From Pre-trained Models to Large Language Models: A Comprehensive Survey of AI-Driven Psychological Computing

    cs.CY 2026-03 unverdicted novelty 6.0

    The paper introduces a new taxonomy that groups AI-driven psychological computing tasks by their underlying computational patterns into four categories and reviews over 300 works from the pre-trained model to LLM eras.

  3. Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities

    cs.HC 2025-07 unverdicted novelty 6.0

    A gamified system with multiple LLM agents of varied personalities gathers interaction data to produce more effective and interpretable Big Five personality assessments than single-context methods.

  4. A Survey of Large Language Models for Perception and Measurement of Human Psychology

    cs.CY 2026-05 unverdicted novelty 5.0

    A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.

  5. Cross-Lingual Attention Distillation with Personality-Informed Generative Augmentation for Multilingual Personality Recognition

    cs.CL 2026-04 unverdicted novelty 5.0

    ADAM uses personality-guided LLM augmentation and cross-lingual attention distillation to raise balanced accuracy on multilingual personality recognition to 0.6332 on Essays and 0.7448 on Kaggle, outperforming standar...

  6. Dark Personality Traits and Online Toxicity: Linking Self-Reports to Reddit Activity

    cs.CY 2025-12 conditional novelty 5.0

    Dark personality traits show weak, mostly non-significant links to toxicity in Reddit comments, while self-reported engagement with incivility strongly aligns with measurable toxic language.

  7. Dark Personality Traits and Online Toxicity: Linking Self-Reports to Reddit Activity

    cs.CY 2025-12 unverdicted novelty 5.0

    Dark personality traits predict self-reported online incivility but show no reliable link to linguistic toxicity features extracted from users' actual Reddit activity.