REVIEW 2 cited by
Circumventing interpretability: How to defeat mind-readers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The increasing capabilities of artificial intelligence (AI) systems make it ever more important that we interpret their internals to ensure that their intentions are aligned with human values. Yet there is reason to believe that misaligned artificial intelligence will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and suggest a framework for thinking about these potential future risks.
Forward citations
Cited by 2 Pith papers
-
Obfuscated Activations Bypass LLM Latent-Space Defenses
Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.
-
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
In GPT-2 multi-document QA, the layer gap between the first correct top-1 token prediction and its stable final form is larger when relevant information is in the middle of the context.
Discussion (0). Continue with ORCID to comment.