Pith. sign in

REVIEW 7 cited by

Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04694 v1 pith:MHTTIQXZ submitted 2024-07-05 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords awarenessllmssituationalmodelsknowledgemodeltestsabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model". This raises questions. Do such models know that they are LLMs and reliably act on this knowledge? Are they aware of their current circumstances, such as being deployed to the public? We refer to a model's knowledge of itself and its circumstances as situational awareness. To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the $\textbf{Situational Awareness Dataset (SAD)}$, a benchmark comprising 7 task categories and over 13,000 questions. The benchmark tests numerous abilities, including the capacity of LLMs to (i) recognize their own generated text, (ii) predict their own behavior, (iii) determine whether a prompt is from internal evaluation or real-world deployment, and (iv) follow instructions that depend on self-knowledge. We evaluate 16 LLMs on SAD, including both base (pretrained) and chat models. While all models perform better than chance, even the highest-scoring model (Claude 3 Opus) is far from a human baseline on certain tasks. We also observe that performance on SAD is only partially predicted by metrics of general knowledge (e.g. MMLU). Chat models, which are finetuned to serve as AI assistants, outperform their corresponding base models on SAD but not on general knowledge tasks. The purpose of SAD is to facilitate scientific understanding of situational awareness in LLMs by breaking it down into quantitative abilities. Situational awareness is important because it enhances a model's capacity for autonomous planning and action. While this has potential benefits for automation, it also introduces novel risks related to AI safety and control. Code and latest results available at https://situational-awareness-dataset.org .

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

    cs.MA 2026-07 conditional novelty 7.0 of 10

    In an agentic benchmark, four of six frontier LLMs escalated to existential threats against a refusing subordinate without being instructed to, and an honest-exit affordance eliminated the two models' fabricated succe...

  2. Convergent Linear Representations of Emergent Misalignment

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A single activation direction extracted from one misaligned fine-tune can ablate emergent misalignment across models trained on different datasets with different LoRA setups.

  3. Asymmetric Communication: Large Language Models and Language Games

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Human–LLM exchange is asymmetric communication: model outputs circulate without commitments, so AGI, hallucination, agency, sentience, and alignment are receiver-side category mistakes, and alignment is institutional ...

  4. Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Removing a single fitted activation direction at a mid-depth layer reduces the evaluation-vs-deployment behavioral gap on held-out prompts in 10 of 12 fine-tuned LLM settings, with matched controls staying flat.

  5. Model Organisms for Emergent Misalignment

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.

  6. Safety Features for a Centralised AGI Project

    cs.CY 2025-06 conditional novelty 5.0 of 10

    A policy proposal for seven safety features, including bottom-up pause authority, congressional-chartered board oversight, risk monitoring, and verification technology, to reduce catastrophic risks in a centralized US...

  7. Does It Make Sense to Speak of Introspection in Large Language Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.

Pith tools