Pith. sign in

Using Hallucinations to Bypass GPT4's Filter

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large language models (LLMs) are initially trained on vast amounts of data, then fine-tuned using reinforcement learning from human feedback (RLHF); this also serves to teach the LLM to provide appropriate and safe responses. In this paper, we present a novel method to manipulate the fine-tuned version into reverting to its pre-RLHF behavior, effectively erasing the model's filters; the exploit currently works for GPT4, Claude Sonnet, and (to some extent) for Inflection-2.5. Unlike other jailbreaks (for example, the popular "Do Anything Now" (DAN) ), our method does not rely on instructing the LLM to override its RLHF policy; hence, simply modifying the RLHF process is unlikely to address it. Instead, we induce a hallucination involving reversed text during which the model reverts to a word bucket, effectively pausing the model's filter. We believe that our exploit presents a fundamental vulnerability in LLMs currently unaddressed, as well as an opportunity to better understand the inner workings of LLMs during hallucinations.

citation-role summary

background 1

citation-polarity summary

fields

cs.CR 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

ATAG: AI-Agent Application Threat Assessment with Attack Graphs

cs.CR · 2025-06-03 · conditional · novelty 6.0

ATAG extends the MulVAL attack graph generator with custom Datalog facts and interaction rules to model multi-step attacks on LLM-based multi-agent applications, demonstrated on a trip planner and an email responder.

citing papers explorer

Showing 1 of 1 citing paper.

  • ATAG: AI-Agent Application Threat Assessment with Attack Graphs cs.CR · 2025-06-03 · conditional · none · ref 32 · internal anchor

    ATAG extends the MulVAL attack graph generator with custom Datalog facts and interaction rules to model multi-step attacks on LLM-based multi-agent applications, demonstrated on a trip planner and an email responder.