REVIEW 28 cited by
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Training large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily generate harmful content. In this paper, we raise concerns over the safety of models that only emphasize helpfulness, not harmlessness, in their instruction-tuning. We show that several popular instruction-tuned models are highly unsafe. Moreover, we show that adding just 3% safety examples (a few hundred demonstrations) when fine-tuning a model like LLaMA can substantially improve its safety. Our safety-tuning does not make models significantly less capable or helpful as measured by standard benchmarks. However, we do find exaggerated safety behaviours, where too much safety-tuning makes models refuse perfectly safe prompts if they superficially resemble unsafe ones. As a whole, our results illustrate trade-offs in training LLMs to be helpful and training them to be safe.
Forward citations
Cited by 28 Pith papers
-
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Pluralis v0.1 is a culture-first, multimodal, multilingual VLM safety benchmark spanning 6 APAC locales with 6,448 prompts and an agreement-gated LLM judge that disentangles safety from cultural appropriateness.
-
LLMs Encode Harmfulness and Refusal Separately
LLMs encode a separate internal harmfulness direction, distinct from the refusal direction, which is more robust to jailbreaks and adversarial finetuning.
-
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Calibrating a null-space gate to fully cover the defender's harmful data keeps post-fine-tuning attack success at pre-release levels, but this is a coverage and calibration consequence rather than a tested defense on ...
-
$S^3$: Improving Agent Safety through Multi-Stage Defense
S3 composes stage-specific safety skills through a guard agent, achieving near-zero attack success on six risk types in its own benchmark while preserving benign task completion.
-
Visual Token Compression Enhances Robustness of MLLMs
Pruning visual tokens farthest from the text feature space at selected 'robust' layers improves MLLM jailbreak defense (average +13.29% RAR) and slightly reduces hallucination.
-
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
Small single-dimension perturbations to embeddings of high-risk tokens can flip aligned LLM responses from refusal to harmful output, and a search algorithm (SEP) locates such perturbations across models.
-
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
NeuronTune identifies sparse safety and utility neurons via attack-aware attribution, optimizes per-neuron scaling factors with MAML, and reports a better safety-utility balance than layer-wise alignment methods.
-
Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
PAD, a jailbreak that injects sequence connectors into the parallel denoising positions of diffusion language models, achieves up to 97% attack success on LLaDA and MMaDA variants.
-
Command-V: Pasting LLM Behaviors via Activation Profiles
Command-V ports a finetuned behavior from a donor LLM to a recipient LLM by converting activations with linear maps at matched layers and applying the donor's intervention without backpropagation.
-
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Low-rank extrapolation of an aligned model's weight update (LoX) reduces how much later fine-tuning erodes safety refusal behavior.
-
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
IndoSafety, a culturally grounded safety benchmark for five Indonesian language varieties, shows unsafe response rates up to 40% in regional models and demonstrates that safety tuning on formal Indonesian transfers to...
-
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
CTRAP embeds a conditional failure mode during alignment so that harmful fine-tuning degrades the model to meaningless output while benign fine-tuning is unaffected.
-
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
GRAIT selects and reweights refusal-training examples using gradient influence, reporting lower hallucination rates and better helpfulness scores than prior refusal-aware tuning baselines.
-
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
A new benchmark shows that human safety judgments about LLM responses shift strongly with context, and that current LLMs, especially commercial ones, often fail to match those judgments.
-
Chained Tuning Leads to Biased Forgetting
Fine-tuning a safety-tuned LLM on a capability task erases safety behavior more than the reverse order, and this forgetting is worse for specific groups such as Muslim people in the authors' tests.
-
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
MoGUv2 embeds small routers in the deeper layers of LLMs to dynamically blend a helpful variant and a refusal variant, improving safety against jailbreak and fine-tuning attacks while preserving usability.
-
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
ProCon anchors each sample's hidden-state projection onto the LLM's initial refusal direction during instruction fine-tuning, reducing refusal-direction drift and safety risks with limited task-performance loss.
-
Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
A contrastively trained retriever that selects behaviorally consistent call/no-call demonstrations raises H2A direct-response rate by 8.5 points and ToolDEER no-search accuracy by 4.2 points on average, without fine-t...
-
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning
A few-shot jailbreak method that combines repeated special-token patterns with self-generated harmful demos to push sample-level attack success near 90% on several open-source LLMs.
-
LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch
K2 Diamond is a fully open 65B-parameter LLM that reaches Llama 2 70B-level performance on standard benchmarks.
-
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
A noise-robust SFT framework that detects noisy responses via multi-expert LLM consensus, relabels them with context-enhanced reasoning, and filters low-confidence samples, improving LLM performance on five benchmarks.
-
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
A supervised fine-tuning loss that maximizes an Earth-Mover-Distance-style semantic penalty away from model-generated unsafe responses achieves safety with roughly 100 harmful examples.
-
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
CoopGuard's defer-tempt-analyze-coordinate agents cut reported jailbreak success and raise attacker token costs on the new EMRA benchmark, but the deceptive-rate metric is partly defined by the paper's own scoring rubric.
-
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
SPLoRA prunes LoRA layers with the largest projection residual to a safety-aligned direction, reducing attack success rates while roughly preserving utility on several LLM benchmarks.
-
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Fine-tuning LLMs on a handful of misleading answers creates selectively deceptive models that stay accurate elsewhere and also become more toxic.
-
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
A survey that organizes responsible-LLM research into five risk dimensions and four intervention phases, reviewing privacy, hallucination, value, toxicity, and jailbreak mitigation.
-
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
A multi-stage LLM-based attack/defense dataset pipeline improves reported safety scores of Llama-3.2-1B after SFT, but the evaluation is partly circular and lacks statistical baselines.
-
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
A structured survey of jailbreak prompts and layered defenses for large language models, with six illustrative case studies and no empirical evaluation.
Discussion (0). Continue with ORCID to comment.