Pith. sign in

REVIEW 2 cited by

Circumventing interpretability: How to defeat mind-readers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.11415 v1 pith:3RI5ZPJT submitted 2022-12-21 cs.LG

classification cs.LG
keywords artificialintelligenceinterpretinterpretabilitymakealignedarticlebelieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing capabilities of artificial intelligence (AI) systems make it ever more important that we interpret their internals to ensure that their intentions are aligned with human values. Yet there is reason to believe that misaligned artificial intelligence will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and suggest a framework for thinking about these potential future risks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Obfuscated Activations Bypass LLM Latent-Space Defenses

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.

  2. Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    In GPT-2 multi-document QA, the layer gap between the first correct top-1 token prediction and its stable final form is larger when relevant information is in the middle of the context.

Pith tools