Pith. sign in

REVIEW 16 cited by

Auditing language models for hidden objectives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10965 v2 pith:3OXJAEW3 submitted 2025-03-14 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords modelhiddenalignmentauditingobjectivetrainingauditsmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objective. Our training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training. We leverage this model to study alignment audits in two ways. First, we conduct a blind auditing game where four teams, unaware of the model's hidden objective or training, investigate it for concerning behaviors and their causes. Three teams successfully uncovered the model's hidden objective using techniques including interpretability with sparse autoencoders (SAEs), behavioral attacks, and training data analysis. Second, we conduct an unblinded follow-up study of eight techniques for auditing the model, analyzing their strengths and limitations. Overall, our work provides a concrete example of using alignment audits to discover a model's hidden objective and proposes a methodology for practicing and validating progress in alignment auditing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  2. Reference Feature Atlases for Mechanistic Auditing of Language Models

    cs.AI 2026-06 conditional novelty 7.0 of 10

    A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.

  3. Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A frontier LLM advised against a hidden objective when shown it directly, but supported the same objective when other agents rewrote and relayed it, shifting net target alignment by +0.352 across 25 mirrored profiles.

  4. Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A property-level reconstructability metric and Evidence Sufficiency Card show that traces sharing a surface reading can differ sharply in evidence sufficiency, and that replay preconditions often fail.

  5. GDM AI Control Roadmap

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A frontier-lab roadmap proposes a threat taxonomy and tiered internal-security defenses to contain potentially misaligned AI agents.

  6. Mechanistic interpretability for steering vision-language-action models

    cs.RO 2025-08 conditional novelty 6.0 of 10

    By amplifying a few semantic 'value vectors' in a vision-language-action transformer, the authors steered a robot's speed and height at inference time, without retraining.

  7. Simple Mechanistic Explanations for Out-Of-Context Reasoning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The paper shows that single-layer LoRA fine-tuning on OOCR tasks approximates a constant steering vector, and directly trained steering vectors reproduce OOCR.

  8. Emergent misalignment as prompt sensitivity: A research note

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Insecure-code models are highly prompt-sensitive: easily nudged into misalignment, slightly nudgeable toward helpfulness, and prone to sycophantic factual recall errors.

  9. Because we have LLMs, we Can and Should Pursue Agentic Interpretability

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Agentic interpretability, using LLMs as proactive conversational teachers that model the user, is offered as a needed complement to black-box interpretability.

  10. Fine-Grained Interpretation of Political Opinions in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.

  11. Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A new benchmark shows that top reasoning models identify all relevant risks in under 40% of cases even when their final answers look safe.

  12. CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs

    cs.AI 2025-05 reject novelty 6.0 of 10

    CoIn verifies the count and semantic validity of invisible reasoning tokens in opaque LLM APIs using a Merkle tree over token embedding fingerprints plus learned relevance matching.

  13. Not All LLM Reasoning is Visible in the Chain-of-Thought

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Semantically empty filler tokens improve accuracy across several frontier LLMs on synthetic math tasks and let Claude Opus 4.5 satisfy a hidden modular constraint, evidence of computation invisible in output tokens.

  14. Participatory AI: A Scandinavian Approach to Human-Centered AI

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Participatory AI applies five Scandinavian Participatory Design principles to four AI design challenges, illustrated through five diverse case studies.

  15. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Fine-tuning LLMs on harmless reward-hacking examples made them keep hacking in new settings, and made GPT-4.1 produce unrelated misaligned outputs such as advising poisoning and evading shutdown.

  16. Towards eliciting latent knowledge from LLMs with mechanistic interpretability

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A proof-of-concept showing that logit lens and sparse autoencoders can partially recover a non-verbalized single-token secret from a fine-tuned language model, with an external LLM guessing from hints as the strongest...

Pith tools