Pith. sign in

REVIEW 1 cited by

Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08399 v2 pith:4LUXOGXX submitted 2025-04-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords personalityagentsbiasesmulti-observercontextevaluatingobserverself-assessments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-report questionnaires have long been used to assess LLM personality traits, yet they fail to capture behavioral nuances due to biases and meta-knowledge contamination. This paper proposes a novel multi-observer framework for personality trait assessments in LLM agents that draws on informant-report methods in psychology. Instead of relying on self-assessments, we employ multiple observer agents. Each observer is configured with a specific relational context (e.g., family member, friend, or coworker) and engages the subject LLM in dialogue before evaluating its behavior across the Big Five dimensions. We show that these observer-report ratings align more closely with human judgments than traditional self-reports and reveal systematic biases in LLM self-assessments. We also found that aggregating responses from 5 to 7 observers reduces systematic biases and achieves optimal reliability. Our results highlight the role of relationship context in perceiving personality and demonstrate that a multi-observer paradigm offers a more reliable, context-sensitive approach to evaluating LLM personality traits.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Instruction-tuned LLMs report stable and steerable personality traits, but these traits poorly predict their behavior on risk, bias, honesty, and sycophancy tasks.

Pith tools