Pith. sign in

REVIEW 9 cited by

The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11484 v9 pith:CXMAZYST submitted 2024-07-16 cs.AI cs.CL

classification cs.AIcs.CL
keywords modelsrole-playingconsistencylanguagesurveyalignmentcharactercurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This survey explores the burgeoning field of role-playing with language models, focusing on their development from early persona-based models to advanced character-driven simulations facilitated by Large Language Models (LLMs). Initially confined to simple persona consistency due to limited model capabilities, role-playing tasks have now expanded to embrace complex character portrayals involving character consistency, behavioral alignment, and overall attractiveness. We provide a comprehensive taxonomy of the critical components in designing these systems, including data, models and alignment, agent architecture and evaluation. This survey not only outlines the current methodologies and challenges, such as managing dynamic personal profiles and achieving high-level persona consistency but also suggests avenues for future research in improving the depth and realism of role-playing applications. The goal is to guide future research by offering a structured overview of current methodologies and identifying potential areas for improvement. Related resources and papers are available at https://github.com/nuochenpku/Awesome-Role-Play-Papers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.

  2. LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    ChatAnime, a new emotionally supportive anime role-play benchmark, reports top LLMs outperforming human enthusiasts on role-playing and emotional support metrics while humans keep the diversity edge.

  3. CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A role-playing LLM that reasons about the scene and its own state before responding, trained with two semantic rewards, beats stronger baselines on role-play benchmarks.

  4. On the Adaptive Psychological Persuasion of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive preference-optimization method helps LLM persuaders choose among 11 psychological strategies, improving persuasion success on counterfactual facts while preserving general capability.

  5. Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RAR improves role-playing agents by distilling character-grounded reasoning traces and optimizing the reasoning style to fit the dialogue scene.

  6. STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    STEER-BENCH is a Reddit-derived benchmark of 5,552 multiple-choice questions on which the best of 13 large language models scores near 65 percent, versus human experts near 81 percent.

  7. Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

    cs.AI 2026-08 conditional novelty 5.0 of 10

    A multi-agent adversarial evaluation platform with six progressive attack strategies shows that role-playing LLMs degrade under sustained pressure, with automated judging correlating with human ratings.

  8. Statistical Hypothesis Testing for Auditing Robustness in Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A permutation-based hypothesis test on pairwise semantic similarities detects whether LLM outputs shift under arbitrary input or model perturbations.

  9. Personalized Image Generation from an Author Writing Style

    cs.CV 2025-07 conditional novelty 4.0 of 10

    LLM-generated text-to-image prompts derived from author style sheets produce images that ten raters judged as moderately faithful (4.08/5), but the evaluation has no control condition and the dataset link is a placeholder.

Pith tools