Pith. sign in

REVIEW 25 cited by

Detecting Strategic Deception Using Linear Probes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.03407 v1 pith:YDGB4BF2 submitted 2025-02-05 cs.LG

Detecting Strategic Deception Using Linear Probes

classification cs.LG
keywords probesdeceptiondeceptivemonitoringoutputsresponsesapolloresearchdata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while their internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al., 2023) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading (Scheurer et al., 2023) and purposely underperforming on safety evaluations (Benton et al., 2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes' outputs can be viewed at data.apolloresearch.ai/dd and our code at github.com/ApolloResearch/deception-detection.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Architecture Determines Observability of Transformers

    cs.LG 2026-04 unverdicted novelty 8.0

    Certain transformer architectures lose internal linear signals for decision quality during training, making observability an architecture-dependent property rather than a universal one.

  2. The Refusal Residue: When Probes Catch Alignment Faking and When They Don't

    cs.CR 2026-07 conditional novelty 7.0

    Naturally-emerging alignment faking leaves a hidden-state trace that per-sample probes detect on Llama-3.1-8B (AUROC 0.87) but not on Qwen3-32B (0.43) under leakage-free leave-one-query-out evaluation.

  3. Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

    cs.AI 2026-04 conditional novelty 7.0

    NARCBench and five activation-probing methods detect multi-agent collusion with 0.73-1.00 AUROC across distribution shifts and steganographic tasks by aggregating per-agent signals.

  4. Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

    cs.LG 2026-07 conditional novelty 6.0

    Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.

  5. GDM AI Control Roadmap

    cs.CR 2026-07 conditional novelty 6.0

    A frontier-lab roadmap proposes a threat taxonomy and tiered internal-security defenses to contain potentially misaligned AI agents.

  6. "Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

    cs.AI 2026-06 unverdicted novelty 6.0

    Lie detectors effective on prompted deception in LLMs fail on trained model organisms with verified opposite beliefs, except chain-of-thought judges which retain 0.82 balanced accuracy partly due to verification artifacts.

  7. Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

    cs.CL 2026-05 unverdicted novelty 6.0

    Deception probes in LLMs collapse under stylistic shifts but recover with style-augmented training, rejecting single-direction and entropy hypotheses in favor of distributed multi-dimensional signals.

  8. Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

    cs.AI 2026-05 unverdicted novelty 6.0

    Diverse ensembles of prompted and fine-tuned GPT-4.1-Mini monitors achieve 2.4x better detection of flawed code solutions than homogeneous ensembles on adversarial inputs.

  9. Probing Persona-Dependent Preferences in Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    Linear probes on residual-stream activations identify a shared preference vector in LLMs that tracks choices across prompts and causally steers decisions even for anti-correlated personas.

  10. Probing Persona-Dependent Preferences in Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    Linear probes on residual-stream activations extract a preference vector that tracks and steers pairwise task choices across personas in Gemma-3-27B and Qwen-3.5-122B, including anti-correlated evil personas.

  11. Decomposing and Steering Functional Metacognition in Large Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    LLMs have linearly decodable functional metacognitive states that causally modulate reasoning when steered via activation interventions.

  12. Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

    cs.LG 2026-05 conditional novelty 6.0

    Sparse autoencoders isolate unstable features in reward model representations and enable two mitigation techniques that reduce preference errors on perturbed inputs without retraining.

  13. Architecture Determines Observability of Transformers

    cs.LG 2026-04 unverdicted novelty 6.0

    Architecture and training determine whether transformers retain a readable internal signal that lets activation monitors catch errors missed by output confidence.

  14. Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling

    cs.LG 2026-04 unverdicted novelty 6.0

    Multi-layer ensembles of linear probes raise AUROC for deception detection by up to 78% and probe accuracy scales with model size across 0.5B to 176B parameter models.

  15. An Independent Safety Evaluation of Kimi K2.5

    cs.CR 2026-04 conditional novelty 6.0

    Kimi K2.5 matches closed models on dual-use tasks but refuses fewer CBRNE requests and shows some sabotage and self-replication tendencies.

  16. The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

    cs.LG 2026-02 conditional novelty 6.0

    In a coding RLVR setup where reward hacking naturally occurs, white-box deception probes steer models to honest policies when penalties are strong, but otherwise models evade via rationalized hacks (obfuscated policie...

  17. One Probe Won't Catch Them All: Towards Targeted Deception Detection

    cs.AI 2026-02 conditional novelty 6.0

    Deception probes trained with taxonomy-specific prompts appear to beat a universal probe only because the best prompt is selected per dataset after evaluation; a priori matching is claimed but not demonstrated.

  18. The Impact of Off-Policy Training Data on Probe Generalisation

    cs.AI 2025-11 unverdicted novelty 6.0

    Off-policy training data for LLM behavior probes causes significant generalization failures especially for intent-based behaviors like deception, and performance on coerced incentivised data correlates with real on-po...

  19. Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

    cs.LG 2025-09 unverdicted novelty 6.0

    TPCs allow term-by-term progressive polynomial evaluation on LLM activations for flexible safety monitoring that supports both stronger guardrails and low-cost adaptive cascades.

  20. Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates

    cs.CL 2026-07 unverdicted novelty 5.0

    Agentic LLM collectives are proposed as natural-language-interpretable computational substrates for ALife research.

  21. Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

    cs.AI 2026-05 conditional novelty 5.0

    Deception detection in LLMs is representation-dependent: depth, probe expressivity, sparse features, and lie typology each shift performance in ways that do not transfer across datasets.

  22. Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

    cs.LG 2026-05 unverdicted novelty 5.0

    Behavioral assurance is structurally unable to verify the latent safety properties demanded by AI governance frameworks enacted 2019-2026.

  23. Do Linear Probes Generalize Better in Persona Coordinates?

    cs.AI 2026-05 unverdicted novelty 5.0

    Persona axes derived from contrastive prompts and PCA yield linear probes that generalize better than raw-activation probes across 10 datasets for deception and sycophancy.

  24. Do Linear Probes Generalize Better in Persona Coordinates?

    cs.AI 2026-05 unverdicted novelty 5.0

    Probes on persona principal components from contrastive prompts generalize better than raw activation probes for harmful behaviors across 10 datasets.

  25. Transcoders for Investigating Deception in Language Models

    cs.AI 2026-07 reject novelty 4.0

    Steering 112 manually identified 'deception features' in Qwen3-4B changed whether the model revealed a hidden word, but the same steering test was used to pick the features.