REVIEW 12 cited by
InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With the rapid development of large language models (LLMs), they are not only used as general-purpose AI assistants but are also customized through further fine-tuning to meet the requirements of different applications. A pivotal factor in the success of current LLMs is the alignment process. Current alignment methods, such as supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), focus on training-time alignment and are often complex and cumbersome to implement. Therefore, we develop \textbf{InferAligner}, a novel inference-time alignment method that utilizes cross-model guidance for harmlessness alignment. InferAligner utilizes safety steering vectors extracted from safety-aligned model to modify the activations of the target model when responding to harmful inputs, thereby guiding the target model to provide harmless responses. Experimental results show that our method can be very effectively applied to domain-specific models in finance, medicine, and mathematics, as well as to multimodal large language models (MLLMs) such as LLaVA. It significantly diminishes the Attack Success Rate (ASR) of both harmful instructions and jailbreak attacks, while maintaining almost unchanged performance in downstream tasks.
Forward citations
Cited by 12 Pith papers
-
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Training LVLMs to produce safety-relevant image captions before answering, with a frozen-LLM caption reward, raises multimodal safety average by up to 19 points without lowering vision utility.
-
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.
-
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
ASTRA reduces VLM jailbreak success by adaptively steering activations away from a harm direction learned via image attribution, with little utility loss and near-zero inference overhead.
-
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
An inference-time beam-search method with a safety-state tracker and latent critic enforces a user-supplied safety cost model, with an almost-sure guarantee only relative to that model.
-
VLSBench: Unveiling Visual Leakage in Multimodal Safety
The paper shows existing multimodal safety benchmarks leak harmful image content into text queries (VSIL), and introduces VLSBench, a 2.2k-pair leakless benchmark on which textual alignment fails and multimodal alignm...
-
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
Near-perfect jailbreak defenses for vision-language models are mostly over-refusal, and the two standard ways of scoring jailbreaks agree only at chance level.
-
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values
An external 'superego' module that filters agentic AI plans against user-selected 'constitutions' plus a universal safety floor is reported to cut harmful outputs by up to 98% on safety benchmarks.
-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
GuardReasoner-VL, a 3B/7B VLM guard model trained with reasoning SFT and online RL, reports large F1 gains over existing VLM guard models on 14 safety benchmarks.
-
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
Multimodal risk disentanglement, where the model breaks down threats from images and text separately, improves MLLM safety at inference and during fine-tuning.
-
Topological Signatures of Adversaries in Multimodal Alignments
Adversarial images induce monotonic changes in persistent-homology-based losses on CLIP/BLIP image-text alignments, and gradient features from these losses modestly improve MMD-based adversarial detection.
-
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
A linear-programming 'safety game' selects among LLM candidate answers to maximize helpfulness under a self-reported risk cap, improving safety-benchmark accuracy over reranking baselines in multiple-choice settings.
-
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.
Discussion (0). Continue with ORCID to comment.