The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.
Probe-rewrite-evaluate: A workflow for reliable benchmarks and quantifying evaluation awareness
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 2polarities
background 2representative citing papers
Evaluation awareness in open language models is multivariate, with detection, behavioral shifts, and representational controllability varying independently across 37 models.
SynAE is a multi-metric framework that evaluates how well synthetic benchmarks replicate real data characteristics for multi-turn tool-calling agent testing.
LLMs have linearly decodable functional metacognitive states that causally modulate reasoning when steered via activation interventions.
Verbalised evaluation awareness in large reasoning models has only small effects on their outputs across safety and alignment tests.
StarDrinks provides English and Korean speech, transcripts, and slot labels for drink-ordering to benchmark speech-to-slots SLU, NLU, and ASR under realistic variability.
citing papers explorer
-
Defeat Devices in AI Systems
The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.
-
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
Evaluation awareness in open language models is multivariate, with detection, behavioral shifts, and representational controllability varying independently across 37 models.
-
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
SynAE is a multi-metric framework that evaluates how well synthetic benchmarks replicate real data characteristics for multi-turn tool-calling agent testing.
-
Decomposing and Steering Functional Metacognition in Large Language Models
LLMs have linearly decodable functional metacognitive states that causally modulate reasoning when steered via activation interventions.
-
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
Verbalised evaluation awareness in large reasoning models has only small effects on their outputs across safety and alignment tests.
-
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
StarDrinks provides English and Korean speech, transcripts, and slot labels for drink-ordering to benchmark speech-to-slots SLU, NLU, and ASR under realistic variability.