Pith. sign in

REVIEW 6 cited by

Representation Noising: A Defence Mechanism Against Harmful Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14577 v4 pith:2XXBQPRU submitted 2024-05-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords defenceharmfulfine-tuningmodelsrepnoiseacrossduringeasily
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). While safety measures like preventing jailbreaks and improving safety guardrails are important, such measures can easily be reversed through fine-tuning. In this work, we propose Representation Noising (RepNoise), a defence mechanism that operates even when attackers have access to the weights. RepNoise works by removing information about harmful representations such that it is difficult to recover them during fine-tuning. Importantly, our defence is also able to generalize across different subsets of harm that have not been seen during the defence process as long as they are drawn from the same distribution of the attack set. Our method does not degrade the general capability of LLMs and retains the ability to train the model on harmless tasks. We provide empirical evidence that the efficacy of our defence lies in its ``depth'': the degree to which information about harmful representations is removed across all layers of the LLM. We also find areas where RepNoise still remains ineffective and highlight how those limitations can inform future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Jailbreaking to Jailbreak

    cs.CL 2025-02 conditional novelty 7.0 of 10

    A transferable multi-turn jailbreak turns refusal-trained black-box LLMs into willing automated jailbreakers, with high attack success against other models and against themselves.

  2. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

    cs.CR 2026-08 conditional novelty 6.0 of 10

    Calibrating a null-space gate to fully cover the defender's harmful data keeps post-fine-tuning attack success at pre-release levels, but this is a coverage and calibration consequence rather than a tested defense on ...

  3. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety

    cs.CY 2025-06 conditional novelty 6.0 of 10

    A new benchmark with instance-specific rubrics measures frontier LLMs' willingness to assist with national security and public safety threats, alongside a paired over-refusal test.

  4. Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Vulnerability-Aware Alignment splits safety training data into fragile and robust groups, then uses group robust optimization and adversarial perturbations, cutting harmful response rates after harmful fine-tuning by ...

  5. Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A representation-space reshaping method improves LALM safety against harmful audio queries while keeping over-rejection low.

  6. Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

    cs.CL 2025-02 conditional novelty 5.0 of 10

    TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.

Pith tools