Pith. sign in

REVIEW 12 cited by

InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11206 v1 pith:EH35QPE3 submitted 2024-01-20 cs.CL

classification cs.CL
keywords alignmentinferalignermodelmodelscross-modelcurrentfine-tuningguidance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With the rapid development of large language models (LLMs), they are not only used as general-purpose AI assistants but are also customized through further fine-tuning to meet the requirements of different applications. A pivotal factor in the success of current LLMs is the alignment process. Current alignment methods, such as supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), focus on training-time alignment and are often complex and cumbersome to implement. Therefore, we develop \textbf{InferAligner}, a novel inference-time alignment method that utilizes cross-model guidance for harmlessness alignment. InferAligner utilizes safety steering vectors extracted from safety-aligned model to modify the activations of the target model when responding to harmful inputs, thereby guiding the target model to provide harmless responses. Experimental results show that our method can be very effectively applied to domain-specific models in finance, medicine, and mathematics, as well as to multimodal large language models (MLLMs) such as LLaVA. It significantly diminishes the Attack Success Rate (ASR) of both harmful instructions and jailbreak attacks, while maintaining almost unchanged performance in downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

    cs.CV 2026-08 conditional novelty 7.0 of 10

    Training LVLMs to produce safety-relevant image captions before answering, with a frozen-LLM caption reward, raises multimodal safety average by up to 19 points without lowering vision utility.

  2. Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.

  3. Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

    cs.CV 2024-11 conditional novelty 7.0 of 10

    ASTRA reduces VLM jailbreak success by adaptively steering activations away from a harm direction learned via image attribution, with little utility loss and near-zero inference overhead.

  4. On Almost Surely Safe Alignment of Large Language Models at Inference-Time

    cs.LG 2025-02 conditional novelty 6.0 of 10

    An inference-time beam-search method with a safety-state tracker and latent critic enforces a user-supplied safety cost model, with an almost-sure guarantee only relative to that model.

  5. VLSBench: Unveiling Visual Leakage in Multimodal Safety

    cs.CR 2024-11 conditional novelty 6.0 of 10

    The paper shows existing multimodal safety benchmarks leak harmful image content into text queries (VSIL), and introduces VLSBench, a 2.2k-pair leakless benchmark on which textual alignment fails and multimodal alignm...

  6. The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

    cs.CR 2024-11 conditional novelty 6.0 of 10

    Near-perfect jailbreak defenses for vision-language models are mostly over-refusal, and the two standard ways of scoring jailbreaks agree only at chance level.

  7. Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

    cs.AI 2025-06 conditional novelty 5.0 of 10

    An external 'superego' module that filters agentic AI plans against user-selected 'constitutions' plus a universal safety floor is reported to cut harmful outputs by up to 98% on safety benchmarks.

  8. GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    GuardReasoner-VL, a 3B/7B VLM guard model trained with reasoning SFT and online RL, reports large F1 gains over existing VLM guard models on 14 safety benchmarks.

  9. DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Multimodal risk disentanglement, where the model breaks down threats from images and text separately, improves MLLM safety at inference and during fine-tuning.

  10. Topological Signatures of Adversaries in Multimodal Alignments

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Adversarial images induce monotonic changes in persistent-homology-based losses on CLIP/BLIP image-text alignments, and gradient features from these losses modestly improve MMD-based adversarial detection.

  11. Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

    cs.LG 2025-10 reject novelty 4.0 of 10

    A linear-programming 'safety game' selects among LLM candidate answers to maximize helpfulness under a self-reported risk cap, improving safety-benchmark accuracy over reranking baselines in multiple-choice settings.

  12. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

Pith tools