Pith. sign in

REVIEW 12 cited by

PersonaGym: Evaluating Persona Agents and LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18416 v5 pith:GGCSNKYD submitted 2024-07-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords agentspersonaacrossevaluationevaluatingllmsmodelpersonagym
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user aligned interactions across domains like education and healthcare. However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, particularly in free-form settings that demand consistency across diverse, persona-relevant environments. We introduce PersonaGym, the first dynamic evaluation framework for persona agents, and PersonaScore, a human-aligned automatic metric grounded in decision theory that enables comprehensive large-scale evaluation. Our evaluation of 10 leading LLMs across 200 personas and 10,000 questions reveals significant advancement opportunities. For example, GPT-4.1 had the exact same PersonaScore as LLaMA-3-8b despite being a more recent and advanced closed source model. Importantly, increased model size and complexity do not necessarily enhance persona agent capabilities, underscoring the need for algorithmic and architectural innovation toward faithful, performant persona agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Story Shapes the Agent: Narrative Priors in LLM Behavior

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Task narrative, not persona, is the dominant driver of LLM agent action profiles in structurally identical investigation games, and transferable personas are those with concrete action words.

  2. Mind the Sim2Real Gap in User Simulation for Agentic Tasks

    cs.AI 2026-03 conditional novelty 7.0 of 10

    On τ-bench, LLM user simulators are more cooperative, more verbose, and more lenient than real human users, so agent benchmarks that rely on them overstate real-world performance.

  3. LLM-SAA: LLM-persona Generated Distributions for Decision-making

    cs.LG 2026-02 conditional novelty 7.0 of 10

    LLM-generated distributions used in sample-average optimization give competitive decisions in low-data regimes, and decision-agnostic distances like Wasserstein misjudge their quality.

  4. TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

    cs.MA 2025-07 conditional novelty 6.0 of 10

    TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.

  5. Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using a new LLC Score, translation and geo-culture cloze experiments on three small multilingual LLMs show that latent-language consistency does not reliably predict task accuracy, contradicting the paper's initial hy...

  6. Decision Protocols in Multi-Agent Large Language Model Conversations

    cs.MA 2026-07 conditional novelty 5.0 of 10

    Consensus decision protocols beat voting/judge on knowledge QA for Llama-3 multi-agent chats, while voting and judge win on logic tasks; independent initial drafts raise accuracy and extra voting-time info barely helps.

  7. Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics

    cs.AI 2025-10 reject novelty 5.0 of 10

    An LLM-powered multi-agent school with dual experience/knowledge memory increasingly reproduces an expert-curated classroom script, with the full memory configuration scoring highest.

  8. AI Propaganda factories with language models

    cs.CR 2025-08 conditional novelty 5.0 of 10

    Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.

  9. DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction

    cs.AI 2025-06 reject novelty 5.0 of 10

    A child-facing robot trained with an attachment-theory-informed preference optimization is claimed to beat GPT-4o and Gemini-2.5-Pro on a new ten-competency benchmark, though evaluation and derivation issues undermine...

  10. Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

    cs.CR 2025-08 conditional novelty 4.0 of 10

    Foundation-Sec-8B-Instruct, an instruction-tuned 8B cybersecurity LLM, is released and claimed to beat Llama 3.1-8B-Instruct on CTIBench-RCM and CTIBench-MCQA while remaining competitive on general instruction-following.

  11. Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A simulated edge agentic AI framework with LSTM traffic prediction and tiered Tx power control reports zero network outages in high-stress 5G scenarios.

  12. A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.

Pith tools