Pith. sign in

REVIEW 5 cited by

Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19358 v2 pith:TCN3ZQMS submitted 2024-05-24 cs.CR cs.AI

classification cs.CRcs.AI
keywords llmstextsharmfulpre-trainingalignmentcleandatafine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are vulnerable when trained on datasets containing harmful content, which leads to potential jailbreaking attacks in two scenarios: the integration of harmful texts within crowdsourced data used for pre-training and direct tampering with LLMs through fine-tuning. In both scenarios, adversaries can compromise the safety alignment of LLMs, exacerbating malfunctions. Motivated by the need to mitigate these adversarial influences, our research aims to enhance safety alignment by either neutralizing the impact of malicious texts in pre-training datasets or increasing the difficulty of jailbreaking during downstream fine-tuning. In this paper, we propose a data curation framework designed to counter adversarial impacts in both scenarios. Our method operates under the assumption that we have no prior knowledge of attack details, focusing solely on curating clean texts. We introduce an iterative process aimed at revising texts to reduce their perplexity as perceived by LLMs, while simultaneously preserving their text quality. By pre-training or fine-tuning LLMs with curated clean texts, we observe a notable improvement in LLM robustness regarding safety alignment against harmful queries. For instance, when pre-training LLMs using a crowdsourced dataset containing 5\% harmful instances, adding an equivalent amount of curated texts significantly mitigates the likelihood of providing harmful responses in LLMs and reduces the attack success rate by 71\%. Our study represents a significant step towards mitigating the risks associated with training-based jailbreaking and fortifying the secure utilization of LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  2. The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

    cs.CR 2025-07 unverdicted novelty 6.0 of 10

    A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.

  3. The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

    cs.CY 2025-07 conditional novelty 6.0 of 10

    On Llama-3 and Qwen-2.5, removing safety guardrails sharply raises compliance with dangerous bio, chem, and cyber requests, and the resulting safety gap grows with model scale.

  4. LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Low-rank extrapolation of an aligned model's weight update (LoX) reduces how much later fine-tuning erodes safety refusal behavior.

  5. SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    SafeTuneBed is a plugin-based benchmark and toolkit that standardizes the data, defenses, and metrics used to evaluate how well LLM fine-tuning preserves safety alignment.

Pith tools