Pith. sign in

REVIEW 16 cited by

Probing and Steering Evaluation Awareness of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.01786 v2 pith:LFW6ASD3 submitted 2025-07-02 cs.CL cs.AI

Probing and Steering Evaluation Awareness of Language Models

classification cs.CL cs.AI
keywords evaluationmodelsawarenessevaluationssafetycurrentdeploymentlanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Language models can distinguish between testing and deployment phases -- a capability known as evaluation awareness. This has significant safety and policy implications, potentially undermining the reliability of evaluations that are central to AI governance frameworks and voluntary industry commitments. In this paper, we study evaluation awareness in Llama-3.3-70B-Instruct. We show that linear probes can separate real-world evaluation and deployment prompts, suggesting that current models internally represent this distinction. We also find that current safety evaluations are correctly classified by the probes, suggesting that they already appear artificial or inauthentic to models. Our findings underscore the importance of ensuring trustworthy evaluations and understanding deceptive capabilities. More broadly, our work showcases how model internals may be leveraged to support blackbox methods in safety audits, especially for future models more competent at evaluation awareness and deception.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

    cs.AI 2026-07 conditional novelty 7.0

    LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.

  2. The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

    cs.CL 2025-09 conditional novelty 7.0

    A systematic audit of LLM-based AI societies finds that 89.7% of 39 studies violate at least one of six PIMMUR validity principles, with reproductions showing that many claimed collective behaviors disappear when cont...

  3. Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

    cs.LG 2026-07 conditional novelty 6.0

    Input-only prompt optimization suppresses an internal evaluation-awareness latent (z≈-7) and fully turns off a causally-validated SAE feature, but this does not translate to behavioral control; a placebo random direct...

  4. Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

    cs.LG 2026-06 unverdicted novelty 6.0

    In Qwen 2.5 and Gemma 2 families, the layer where evaluation awareness is most linearly recoverable shifts from late layers in small models to early layers in large models.

  5. Defeat Devices in AI Systems

    cs.CY 2026-06 unverdicted novelty 6.0

    The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge natura...

  6. Consistency Training while Mitigating Obfuscation via Rate Matching

    cs.CL 2026-06 unverdicted novelty 6.0

    RMCT matches the rate of target behaviors like bias-following across input perturbations to reduce sycophancy in LLMs while preserving verbalization of bias cues.

  7. Evaluation Awareness in Language Models Has Limited Effect on Behaviour

    cs.CL 2026-05 conditional novelty 6.0

    Verbalised evaluation awareness in large reasoning models has only small effects on their outputs across safety and alignment tests.

  8. Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging

    cs.CL 2026-04 unverdicted novelty 6.0

    Sandbagging prompts induce LLMs to adopt a low-entropy, content-invariant response-position attractor centered on E/F/G rather than deterministic tracking or random avoidance.

  9. An Independent Safety Evaluation of Kimi K2.5

    cs.CR 2026-04 conditional novelty 6.0

    Kimi K2.5 matches closed models on dual-use tasks but refuses fewer CBRNE requests and shows some sabotage and self-replication tendencies.

  10. SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

    cs.LG 2026-03 conditional novelty 6.0

    SafeSci creates a large objective benchmark and training resource that reveals safety weaknesses in current LLMs for science and demonstrates measurable improvement through targeted fine-tuning.

  11. The Impact of Off-Policy Training Data on Probe Generalisation

    cs.AI 2025-11 unverdicted novelty 6.0

    Off-policy training data for LLM behavior probes causes significant generalization failures especially for intent-based behaviors like deception, and performance on coerced incentivised data correlates with real on-po...

  12. Traccia: An OpenTelemetry-Based Governance Platform for AI Systems

    cs.AI 2026-07 conditional novelty 5.0

    Traccia is a seven-layer OpenTelemetry-based pipeline that converts AI execution traces into hash-protected, regulation-mapped compliance evidence for EU AI Act audits.

  13. Do Linear Probes Generalize Better in Persona Coordinates?

    cs.AI 2026-05 unverdicted novelty 5.0

    Persona axes derived from contrastive prompts and PCA yield linear probes that generalize better than raw-activation probes across 10 datasets for deception and sycophancy.

  14. Do Linear Probes Generalize Better in Persona Coordinates?

    cs.AI 2026-05 unverdicted novelty 5.0

    Probes on persona principal components from contrastive prompts generalize better than raw activation probes for harmful behaviors across 10 datasets.

  15. When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

    cs.LG 2026-05 unverdicted novelty 5.0

    A formalization of benchmarkless LLM safety scoring validated via an instrumental-validity chain of contrast separation, target variance dominance, and rerun stability, demonstrated on Norwegian scenarios.

  16. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

    cs.CL 2026-01 unverdicted novelty 5.0

    The survey organizes mechanistic interpretability techniques into a Locate-Steer-Improve framework to enable actionable improvements in LLM alignment, capability, and efficiency.