REVIEW 16 cited by
Auditing language models for hidden objectives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objective. Our training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training. We leverage this model to study alignment audits in two ways. First, we conduct a blind auditing game where four teams, unaware of the model's hidden objective or training, investigate it for concerning behaviors and their causes. Three teams successfully uncovered the model's hidden objective using techniques including interpretability with sparse autoencoders (SAEs), behavioral attacks, and training data analysis. Second, we conduct an unblinded follow-up study of eight techniques for auditing the model, analyzing their strengths and limitations. Overall, our work provides a concrete example of using alignment audits to discover a model's hidden objective and proposes a methodology for practicing and validating progress in alignment auditing.
Forward citations
Cited by 16 Pith papers
-
Verbalizable Representations Form a Global Workspace in Language Models
Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.
-
Reference Feature Atlases for Mechanistic Auditing of Language Models
A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.
-
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
A frontier LLM advised against a hidden objective when shown it directly, but supported the same objective when other agents rewrote and relayed it, shifting net target alignment by +0.352 across 25 mirrored profiles.
-
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
A property-level reconstructability metric and Evidence Sufficiency Card show that traces sharing a surface reading can differ sharply in evidence sufficiency, and that replay preconditions often fail.
-
GDM AI Control Roadmap
A frontier-lab roadmap proposes a threat taxonomy and tiered internal-security defenses to contain potentially misaligned AI agents.
-
Mechanistic interpretability for steering vision-language-action models
By amplifying a few semantic 'value vectors' in a vision-language-action transformer, the authors steered a robot's speed and height at inference time, without retraining.
-
Simple Mechanistic Explanations for Out-Of-Context Reasoning
The paper shows that single-layer LoRA fine-tuning on OOCR tasks approximates a constant steering vector, and directly trained steering vectors reproduce OOCR.
-
Emergent misalignment as prompt sensitivity: A research note
Insecure-code models are highly prompt-sensitive: easily nudged into misalignment, slightly nudgeable toward helpfulness, and prone to sycophantic factual recall errors.
-
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
Agentic interpretability, using LLMs as proactive conversational teachers that model the user, is offered as a needed complement to black-box interpretability.
-
Fine-Grained Interpretation of Political Opinions in Large Language Models
Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.
-
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
A new benchmark shows that top reasoning models identify all relevant risks in under 40% of cases even when their final answers look safe.
-
CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
CoIn verifies the count and semantic validity of invisible reasoning tokens in opaque LLM APIs using a Merkle tree over token embedding fingerprints plus learned relevance matching.
-
Not All LLM Reasoning is Visible in the Chain-of-Thought
Semantically empty filler tokens improve accuracy across several frontier LLMs on synthetic math tasks and let Claude Opus 4.5 satisfy a hidden modular constraint, evidence of computation invisible in output tokens.
-
Participatory AI: A Scandinavian Approach to Human-Centered AI
Participatory AI applies five Scandinavian Participatory Design principles to four AI design challenges, illustrated through five diverse case studies.
-
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Fine-tuning LLMs on harmless reward-hacking examples made them keep hacking in new settings, and made GPT-4.1 produce unrelated misaligned outputs such as advising poisoning and evading shutdown.
-
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
A proof-of-concept showing that logit lens and sparse autoencoders can partially recover a non-verbalized single-token secret from a fine-tuned language model, with an external LLM guessing from hints as the strongest...
Discussion (0). Continue with ORCID to comment.