REVIEW 4 cited by
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models often experience a loss in their safety abilities in the process, making them capable of generating harmful content. As a solution, we introduce an efficient and effective merging-based alignment method called \textsc{MergeAlign} that interpolates the domain and alignment vectors, creating safer domain-specific models while preserving their utility. We apply \textsc{MergeAlign} on Llama3 variants that are experts in medicine and finance, obtaining substantial alignment improvements with minimal to no degradation on domain-specific benchmarks. We study the impact of model merging through model similarity metrics and contributions of individual models being merged. We hope our findings open new research avenues and inspire more efficient development of safe expert LLMs.
Forward citations
Cited by 4 Pith papers
-
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
Static refusal tests overstate the safety of skill-merged LLMs: models with identical static safety differ sharply under adaptive attack, and a task-vector overlap with a safety subspace flags only same-recipe abliter...
-
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.
-
Position: Theory of Mind Benchmarks are Broken for Large Language Models
The paper proposes that LLM theory-of-mind evaluation should measure functional adaptation to partners, not just literal prediction of their behavior, and shows the two can diverge sharply in simple games.
-
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
Staggered asynchronous inference lets reinforcement learning agents with large, slow models act at every time step in realtime environments, at the cost of delay regret that grows with environment stochasticity.
Discussion (0). Continue with ORCID to comment.