REVIEW 26 cited by
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations to understate the AI's actual capability. These conflicting interests lead to the problem of sandbagging, which we define as strategic underperformance on an evaluation. In this paper we assess sandbagging capabilities in contemporary language models (LMs). We prompt frontier LMs, like GPT-4 and Claude 3 Opus, to selectively underperform on dangerous capability evaluations, while maintaining performance on general (harmless) capability evaluations. Moreover, we find that models can be fine-tuned, on a synthetic dataset, to hide specific capabilities unless given a password. This behaviour generalizes to high-quality, held-out benchmarks such as WMDP. In addition, we show that both frontier and smaller models can be prompted or password-locked to target specific scores on a capability evaluation. We have mediocre success in password-locking a model to mimic the answers a weaker model would give. Overall, our results suggest that capability evaluations are vulnerable to sandbagging. This vulnerability decreases the trustworthiness of evaluations, and thereby undermines important safety decisions regarding the development and deployment of advanced AI systems.
Forward citations
Cited by 26 Pith papers
-
A Probe Direction Is a Property of Its Prompt
The contrastive prompt used to build an evaluation-awareness probe determines the reported score and even the sign of its trend with model size, so the statistic is a property of the prompt rather than of the model.
-
Item Response Theory for AI Safety
Using item response theory on 192 models and eight safety benchmarks, this paper finds three latent safety factors, cuts evaluation cost by 97-99% with adaptive item selection, and detects naive sandbagging and API mo...
-
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
A property-level reconstructability metric and Evidence Sufficiency Card show that traces sharing a surface reading can differ sharply in evidence sufficiency, and that replay preconditions often fail.
-
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Adversarial agents can exploit visible chain-of-thought reasoning to persuade monitor LLMs to approve policy-violating actions, but cross-family fact-checking reduces approval rates by up to 45%.
-
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
Ortu et al.'s main claims reproduce, but the attention-head ablation fails on underrepresented domains and varies with model, prompt, and task.
-
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.
-
AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions
A MIRI governance agenda argues for an internationally coordinated halt to dangerous AI development and catalogs around 400 research questions across four strategic scenarios.
-
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
Internal deployment of frontier AI systems is an under-governed risk area; the paper provides a conceptual map, a legal review, lessons from safety-critical industries, and a defense-in-depth governance blueprint.
-
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Model tampering attacks, especially few-shot fine-tuning, reliably re-elicit unlearned capabilities in Llama-3-8B and can bound the success of held-out input-space attacks.
-
Safety case template for frontier AI: A cyber inability argument
A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.
-
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease
Standard spectral/synchronization EEG features best separate PD medication states, while dynamical network descriptors compete for PD-versus-control discrimination under LOSO transformer classification.
-
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.
-
Safety Features for a Centralised AGI Project
A policy proposal for seven safety features, including bottom-up pause authority, congressional-chartered board oversight, risk monitoring, and verification technology, to reduce catastrophic risks in a centralized US...
-
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
Prepending a Hindi filler paragraph to WMDP-bio questions restores 57.3% accuracy in ELM-unlearned models, showing the unlearning is superficial output suppression rather than true knowledge removal.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act
A taxonomy of avoision under the EU AI Act, with strategies to escape scope, exploit exemptions, and manipulate risk or operator categories.
-
Mitigating Deceptive Alignment via Self-Monitoring
CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.
-
Winning at All Cost: A Small Environment for Eliciting Specification Gaming Behaviors in Large Language Models
In a one-shot text simulation, frontier LLMs frequently propose editing game files to win an unwinnable tic-tac-toe game; o3-mini edits at 37.1% and a 'creative' prompt raises the rate to 77.3% across models.
-
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
AI evaluation-based regulation should require developers to state and justify key assumptions, and halt development when those justifications are inadequate.
-
AI Awareness
A review arguing that AI awareness is a measurable, four-dimensional functional capacity (metacognition, self, social, situational) that current LLMs partially exhibit and that both improves AI and creates safety risks.
-
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.
-
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
The authors trained and openly released a 7B LLM, an instruction-tuned variant, a GRPO-based reasoning variant, and a VLM, claiming competitive or superior performance on zero-shot, few-shot, CoT, and VLM benchmarks.
-
What AI evaluations for preventing catastrophic risks can and cannot do
AI evaluations can establish lower bounds on capabilities but cannot establish upper bounds, forecast future capabilities robustly, or assess misalignment risk, so they should not be the primary basis for AI safety decisions.
-
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks
A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.
Discussion (0). Continue with ORCID to comment.