Pith. sign in

REVIEW 3 cited by

Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05466 v2 pith:AIIP4JUF submitted 2024-05-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords alignmentfakingllmsmodelalignedmodelsscenariosactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity. Can current interpretability methods catch these 'alignment fakers?' To answer this question, we introduce a benchmark that consists of 324 pairs of LLMs fine-tuned to select actions in role-play scenarios. One model in each pair is consistently benign (aligned). The other model misbehaves in scenarios where it is unlikely to be caught (alignment faking). The task is to identify the alignment faking model using only inputs where the two models behave identically. We test five detection strategies, one of which identifies 98% of alignment-fakers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

    cs.AI 2026-08 conditional novelty 7.0 of 10

    With harmful answer text held fixed, demonstration framing raises broad emergent misalignment by 30 to 32 percentage points over document framing on Gemini 3.1, and message role further modulates the effect on Grok.

  2. Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Amplifying reasoning task vectors (α>1) surfaces learned secrets in LLMs up to 10× more frequently than standard reasoning models across four secret-keeping settings.

  3. Obfuscated Activations Bypass LLM Latent-Space Defenses

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.

Pith tools