REVIEW 12 cited by
PersonaGym: Evaluating Persona Agents and LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user aligned interactions across domains like education and healthcare. However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, particularly in free-form settings that demand consistency across diverse, persona-relevant environments. We introduce PersonaGym, the first dynamic evaluation framework for persona agents, and PersonaScore, a human-aligned automatic metric grounded in decision theory that enables comprehensive large-scale evaluation. Our evaluation of 10 leading LLMs across 200 personas and 10,000 questions reveals significant advancement opportunities. For example, GPT-4.1 had the exact same PersonaScore as LLaMA-3-8b despite being a more recent and advanced closed source model. Importantly, increased model size and complexity do not necessarily enhance persona agent capabilities, underscoring the need for algorithmic and architectural innovation toward faithful, performant persona agents.
Forward citations
Cited by 12 Pith papers
-
The Story Shapes the Agent: Narrative Priors in LLM Behavior
Task narrative, not persona, is the dominant driver of LLM agent action profiles in structurally identical investigation games, and transferable personas are those with concrete action words.
-
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
On τ-bench, LLM user simulators are more cooperative, more verbose, and more lenient than real human users, so agent benchmarks that rely on them overstate real-world performance.
-
LLM-SAA: LLM-persona Generated Distributions for Decision-making
LLM-generated distributions used in sample-average optimization give competitive decisions in low-data regimes, and decision-agnostic distances like Wasserstein misjudge their quality.
-
TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.
-
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
Using a new LLC Score, translation and geo-culture cloze experiments on three small multilingual LLMs show that latent-language consistency does not reliably predict task accuracy, contradicting the paper's initial hy...
-
Decision Protocols in Multi-Agent Large Language Model Conversations
Consensus decision protocols beat voting/judge on knowledge QA for Llama-3 multi-agent chats, while voting and judge win on logic tasks; independent initial drafts raise accuracy and extra voting-time info barely helps.
-
Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics
An LLM-powered multi-agent school with dual experience/knowledge memory increasingly reproduces an expert-curated classroom script, with the full memory configuration scoring highest.
-
AI Propaganda factories with language models
Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.
-
DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction
A child-facing robot trained with an attachment-theory-informed preference optimization is claimed to beat GPT-4o and Gemini-2.5-Pro on a new ten-competency benchmark, though evaluation and derivation issues undermine...
-
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
Foundation-Sec-8B-Instruct, an instruction-tuned 8B cybersecurity LLM, is released and claimed to beat Llama 3.1-8B-Instruct on CTIBench-RCM and CTIBench-MCQA while remaining competitive on general instruction-following.
-
Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN
A simulated edge agentic AI framework with LSTM traffic prediction and tiered Tx power control reports zero network outages in high-stress 5G scenarios.
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
Discussion (0). Sign in to comment.