Pith. sign in

REVIEW 1 cited by

RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16727 v1 pith:SFSOE3CJ submitted 2024-09-25 cs.CL

classification cs.CL
keywords characterhallucinationsystemsattackframeworkgenerateshallucinationsmode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Role-playing systems powered by large language models (LLMs) have become increasingly influential in emotional communication applications. However, these systems are susceptible to character hallucinations, where the model deviates from predefined character roles and generates responses that are inconsistent with the intended persona. This paper presents the first systematic analysis of character hallucination from an attack perspective, introducing the RoleBreak framework. Our framework identifies two core mechanisms-query sparsity and role-query conflict-as key factors driving character hallucination. Leveraging these insights, we construct a novel dataset, RoleBreakEval, to evaluate existing hallucination mitigation techniques. Our experiments reveal that even enhanced models trained to minimize hallucination remain vulnerable to attacks. To address these vulnerabilities, we propose a novel defence strategy, the Narrator Mode, which generates supplemental context through narration to mitigate role-query conflicts and improve query generalization. Experimental results demonstrate that Narrator Mode significantly outperforms traditional refusal-based strategies by reducing hallucinations, enhancing fidelity to character roles and queries, and improving overall narrative coherence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Universal and Context-Independent Triggers for Precise Control of LLM Outputs

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A single trained token pair inserted around any target text forces Qwen-2 7B and Llama-3.1 8B to output that text on 54 to 75 percent of unseen prompts.

Pith tools