Harmful LLM outputs rely on a compact, cross-harm set of weights distinct from benign skills; alignment compresses them, and pruning them reduces emergent misalignment.
Will releasing the weights of future large language models grant widespread access to pandemic agents? [Internet]
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
Develops a taxonomy of security interaction levels in AI/cloud infrastructure and demonstrates practical attacks exploiting isolation assumptions.
AI model evaluations for biological capabilities should prioritize high-consequence risks like pandemics, informed by life sciences dual-use experience, and occur prior to deployment to enable biosafety measures.
citing papers explorer
-
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Harmful LLM outputs rely on a compact, cross-harm set of weights distinct from benign skills; alignment compresses them, and pruning them reduces emergent misalignment.
-
Investigating The Security of Modern AI and Cloud Infrastructure
Develops a taxonomy of security interaction levels in AI/cloud infrastructure and demonstrates practical attacks exploiting isolation assumptions.
-
Prioritizing High-Consequence Biological Capabilities in Evaluations of Artificial Intelligence Models
AI model evaluations for biological capabilities should prioritize high-consequence risks like pandemics, informed by life sciences dual-use experience, and occur prior to deployment to enable biosafety measures.