Pith. sign in

REVIEW 1 cited by

NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.21053 v1 pith:IPK6L6VP submitted 2025-04-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords neuronsafetyneuronsalignmentfine-tuningharmfulactivationconstraints
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Safety alignment in large language models (LLMs) is achieved through fine-tuning mechanisms that regulate neuron activations to suppress harmful content. In this work, we propose a novel approach to induce disalignment by identifying and modifying the neurons responsible for safety constraints. Our method consists of three key steps: Neuron Activation Analysis, where we examine activation patterns in response to harmful and harmless prompts to detect neurons that are critical for distinguishing between harmful and harmless inputs; Similarity-Based Neuron Identification, which systematically locates the neurons responsible for safe alignment; and Neuron Relearning for Safety Removal, where we fine-tune these selected neurons to restore the model's ability to generate previously restricted responses. Experimental results demonstrate that our method effectively removes safety constraints with minimal fine-tuning, highlighting a critical vulnerability in current alignment techniques. Our findings underscore the need for robust defenses against adversarial fine-tuning attacks on LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Mask2Shield reduces neuron-pruning attack success on ten LLMs from 80–279 to 1–44/313 by training refusal with safety neurons functionally masked while a frozen teacher preserves benign answers.

Pith tools