Pith. sign in

REVIEW 4 cited by

Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04524 v2 pith:AAY3BCOV submitted 2024-10-06 cs.CL

classification cs.CL
keywords securitystrategyllmsmethodsmodsparametersrisksfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction fine-tuning has emerged as a critical technique for customizing Large Language Models (LLMs) to specific applications. However, recent studies have highlighted significant security vulnerabilities in fine-tuned LLMs. Existing defense efforts focus more on pre-training and post-training methods, yet there remains underexplored in in-training methods. To fill this gap, we introduce a novel secure-tuning strategy called SWAT. By analyzing how module-level parameters (e.g. Q/K/V/O) affect the security feature space drift, we identify a robust subset of modules, termed Mods_Rob. Our SWAT strategy begins by warming up Mods_Rob to capture low-level features with minimal security risks, followed by training all parameters to achieve optimal task performance. Essentially, this strategy shifts the early learning burden more from global parameters to Mods_Rob, reducing update magnitudes of the non-robust subset. Across various datasets, scenarios, and LLMs, our strategy has demonstrated significant success in mitigating security risks while preserving task performance. Importantly, it can be seamlessly integrated with pre-training and post-training methods, leading to greater improvements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

    cs.CR 2026-02 conditional novelty 7.0 of 10

    TamperBench systematically benchmarks 21 open-weight LLMs against nine fine-tuning and representation-space attacks and finds all of them can be tampered into producing harmful output while retaining capability.

  2. Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A representation-space reshaping method improves LALM safety against harmful audio queries while keeping over-rejection low.

  3. CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning

    cs.CR 2025-05 conditional novelty 6.0 of 10

    CTRAP embeds a conditional failure mode during alignment so that harmful fine-tuning degrades the model to meaningless output while benign fine-tuning is unaffected.

  4. MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

    cs.CL 2025-09 conditional novelty 5.0 of 10

    MoGUv2 embeds small routers in the deeper layers of LLMs to dynamically blend a helpful variant and a refusal variant, improving safety against jailbreak and fine-tuning attacks while preserving usability.

Pith tools