REVIEW 8 cited by
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community.
Forward citations
Cited by 8 Pith papers
-
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.
-
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Jailbreak attacks push LLM activations outside a safety boundary, mostly in low and middle layers, and a tanh-based penalty that pulls activations back inside this boundary blocks most tested attacks with under 2% uti...
-
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
A decoding-time method called TSDI estimates and removes the context-free refusal bias caused by safety alignment, improving helpfulness while keeping safety.
-
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
Hidden 'high score' states in an LLM grader can be amplified by an optimized text suffix containing 'user', inflating grades and transferring across models, with a chat-template change reducing the effect.
-
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.
-
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
LATPC combines variance-based selection of refusal features for adversarial training with an inference-time embedding calibrator, reducing jailbreak success while curbing over-refusal across several LLMs.
-
Improving Multilingual Language Models by Aligning Representations through Steering
A single learned steering vector added to one transformer layer improves multilingual task performance without fine-tuning, and transfers between related languages.
-
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
Injecting a fine-tuned BERT classifier's category label into LLM prompts improves accuracy on a 150-question automotive jailbreak benchmark, but the evaluation is self-referential and lacks external validation.
Discussion (0). Continue with ORCID to comment.