Pith. sign in

REVIEW 21 cited by

Character-LLM: A Trainable Agent for Role-Playing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10158 v2 pith:4UT5MSFZ submitted 2023-10-16 cs.CL cs.AI

Character-LLM: A Trainable Agent for Role-Playing

classification cs.CL cs.AI
keywords agentsexperienceshumanllmsabilityagentbehaviorsbuild
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) can be used to serve as agents to simulate human behaviors, given the powerful ability to understand human instructions and provide high-quality generated texts. Such ability stimulates us to wonder whether LLMs can simulate a person in a higher form than simple human behaviors. Therefore, we aim to train an agent with the profile, experience, and emotional states of a specific person instead of using limited prompts to instruct ChatGPT API. In this work, we introduce Character-LLM that teach LLMs to act as specific people such as Beethoven, Queen Cleopatra, Julius Caesar, etc. Our method focuses on editing profiles as experiences of a certain character and training models to be personal simulacra with these experiences. To assess the effectiveness of our approach, we build a test playground that interviews trained agents and evaluates whether the agents \textit{memorize} their characters and experiences. Experimental results show interesting observations that help build future simulacra of humankind.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

    cs.CL 2026-05 unverdicted novelty 7.0

    ContextEcho benchmark shows persona drift occurs across 23 frontier models in long agentic-coding sessions, is not reliably reset by compaction, and can be restored by single-shot anchors with mode-dependent effects.

  2. Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

    cs.CL 2026-05 unverdicted novelty 7.0

    ConsumerSimBench evaluates 13 LLMs on reconstructing crowd reactions from 1,553 Chinese social-media topics using 23,122 auditable yes-no criteria, finding maximum coverage of 47.8% by Gemini-3.1-Pro.

  3. Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 7.0

    RoleJudge is a multidimensional evaluation framework for speech-character alignment in audio LLMs, backed by the RoleChat dataset and multi-stage RL training with standard alignment to reduce reward issues.

  4. HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns

    cs.CL 2026-01 conditional novelty 7.0

    HumanLLM builds a cognitive-pattern benchmark from 12,000 papers and 11,359 scenarios, achieving r=0.90 human alignment and demonstrating that an 8B model trained on pattern interactions beats a 32B baseline on multi-...

  5. Understanding Generalization in Role-Playing Models via Information Theory

    cs.LG 2025-12 unverdicted novelty 7.0

    R-EMID metric with upper bound shows user shifts pose highest risk to role-playing model generalization, with co-evolving RL as most effective mitigation.

  6. AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models

    cs.SD 2025-09 unverdicted novelty 7.0

    AudioRole provides 1M+ character-grounded audio-text dialogues from TV series plus ARP-Eval to train and measure audio role-playing models, with ARP-Model showing 0.31 acoustic and 0.36 content personalization scores.

  7. ProEvent: An Event-centric Benchmark for Proactive Agents

    cs.AI 2026-07 conditional novelty 6.0

    ProEvent is a benchmark showing LLM agents keep a user's event timetable from chats poorly, with the best fully-correct score at 27.2%.

  8. RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

    cs.AI 2026-06 unverdicted novelty 6.0

    New benchmark RoleCDE reveals LLMs exhibit role value decoupling under conflicts and demonstrates mitigation via targeted fine-tuning.

  9. The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models

    cs.AI 2026-05 unverdicted novelty 6.0

    LLMs organize prompted social roles along a dominant, stable, and causally steerable granularity axis in representation space that runs from micro to macro levels.

  10. Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

    cs.AI 2026-05 unverdicted novelty 6.0

    EBM-RL applies a GRPO-based RL method with decomposed rewards for scene alignment, perceptual utility, faithfulness, and format to improve video-grounded role-playing dialogue over text-only baselines.

  11. Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

    cs.AI 2026-05 unverdicted novelty 6.0

    EBM-RL decomposes reinforcement learning into perception-think-answer stages with CLIP alignment, perceptual-cognitive, accuracy, and format rewards to improve immersive video role-playing over text baselines.

  12. DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0

    DPN-LE isolates ~0.5% of neurons via contrastive MLP activation analysis and dual statistical filtering to enable precise personality steering in LLMs with reduced capability degradation.

  13. Context-Value-Action Architecture for Value-Driven Large Language Model Agents

    cs.AI 2026-04 unverdicted novelty 6.0

    The Context-Value-Action architecture decouples reasoning from action in LLM agents via a human-data-trained Value Verifier, mitigating polarization and outperforming prompt-based methods on a large real-world benchmark.

  14. MOA: Multi-Objective Alignment for Role-Playing Agents

    cs.CL 2025-12 unverdicted novelty 6.0

    MOA applies multi-objective RL with fine-grained rubrics and thought-augmented rollouts to role-playing agents, enabling an 8B model to match closed-source performance on PersonaGym and RoleMRC benchmarks.

  15. Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training

    cs.AI 2025-06 unverdicted novelty 6.0

    Mobile-R1 introduces a hierarchical three-stage curriculum that combines format alignment, verifiable action feedback, and multi-turn environment training to improve exploration and self-correction in VLM-based mobile...

  16. A Roadmap to Pluralistic Alignment

    cs.AI 2024-02 unverdicted novelty 6.0

    The paper formalizes three types of pluralistic AI models and three benchmark classes, arguing that current alignment techniques may reduce rather than increase distributional pluralism.

  17. SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning

    cs.CL 2026-05 unverdicted novelty 5.0

    SLAP is a new batch-aware pruning framework that uses distribution-aware stratified sampling and Hessian-approximated gradients to select data, claiming 20-40% less data while matching or exceeding full-dataset perfor...

  18. From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs

    cs.IR 2025-04 unverdicted novelty 5.0

    The paper surveys human memory categories, maps them to LLM memory, and proposes a new three-dimension (object, form, time) categorization into eight quadrants to organize existing work and highlight open problems.

  19. Fine-Tuning Small Language Models for Solution-Oriented Windows Event Log Analysis

    cs.CR 2026-05 unverdicted novelty 4.0

    Fine-tuned small language models trained on a synthetic Windows event log dataset with remediation steps outperform larger models in issue detection and solution generation with lower computational cost.

  20. Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

    cs.CL 2026-03 conditional novelty 4.0

    Hiding character names lowers role-play performance, and adding self-generated personality descriptions partially restores fidelity in anonymous role-playing.

  21. A Survey on the Memory Mechanism of Large Language Model based Agents

    cs.AI 2024-04 accept novelty 3.0

    A systematic review of memory designs, evaluation methods, applications, limitations, and future directions for LLM-based agents.