Pith. sign in

REVIEW 4 cited by

Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.06824 v2 pith:4QEPC5FQ submitted 2024-11-11 cs.AI

classification cs.AI
keywords alignmentllmsmodelsdomaindomain-specificefficientexpertmergealign
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models often experience a loss in their safety abilities in the process, making them capable of generating harmful content. As a solution, we introduce an efficient and effective merging-based alignment method called \textsc{MergeAlign} that interpolates the domain and alignment vectors, creating safer domain-specific models while preserving their utility. We apply \textsc{MergeAlign} on Llama3 variants that are experts in medicine and finance, obtaining substantial alignment improvements with minimal to no degradation on domain-specific benchmarks. We study the impact of model merging through model similarity metrics and contributions of individual models being merged. We hope our findings open new research avenues and inspire more efficient development of safe expert LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

    cs.LG 2026-08 accept novelty 6.0 of 10

    Static refusal tests overstate the safety of skill-merged LLMs: models with identical static safety differ sharply under adaptive attack, and a task-vector overlap with a safety subspace flags only same-recipe abliter...

  2. MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

    cs.CR 2025-07 conditional novelty 6.0 of 10

    MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.

  3. Position: Theory of Mind Benchmarks are Broken for Large Language Models

    cs.AI 2024-12 conditional novelty 6.0 of 10

    The paper proposes that LLM theory-of-mind evaluation should measure functional adaptation to partners, not just literal prediction of their behavior, and shows the two can diverge sharply in simple games.

  4. Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Staggered asynchronous inference lets reinforcement learning agents with large, slow models act at every time step in realtime environments, at the cost of delay regret that grows with environment stochasticity.

Pith tools