Pith. sign in

REVIEW 8 cited by

Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06824 v5 pith:APOZRYMU submitted 2024-01-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords jailbreakingllmspatternsattacksbeenfindingslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  2. Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

    cs.CL 2024-12 conditional novelty 7.0 of 10

    Jailbreak attacks push LLM activations outside a safety boundary, mostly in low and middle layers, and a tanh-based penalty that pulls activations back inside this boundary blocks most tested attacks with under 2% uti...

  3. Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A decoding-time method called TSDI estimates and removes the context-free refusal bias caused by safety alignment, improving helpfulness while keeping safety.

  4. Fooling LLM graders into giving better grades through neural activity guided adversarial prompting

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Hidden 'high score' states in an LLM grader can be amplified by an optimized text suffix containing 'user', inflating grades and transferring across models, with a chat-template change reducing the effect.

  5. Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

    cs.CL 2025-02 conditional novelty 5.0 of 10

    TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.

  6. Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

    cs.CR 2025-01 conditional novelty 5.0 of 10

    LATPC combines variance-based selection of refusal features for adversarial training with an inference-time embedding calibrator, reducing jailbreak success while curbing over-refusal across several LLMs.

  7. Improving Multilingual Language Models by Aligning Representations through Steering

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A single learned steering vector added to one transformer layer improves multilingual task performance without fine-tuning, and transfers between related languages.

  8. Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration

    cs.CR 2025-05 conditional novelty 4.0 of 10

    Injecting a fine-tuned BERT classifier's category label into LLM prompts improves accuracy on a 150-question automotive jailbreak benchmark, but the evaluation is self-referential and lacks external validation.

Pith tools