{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:MHTTIQXZH5WC2M7L3E3VUEZ2ZP","short_pith_number":"pith:MHTTIQXZ","schema_version":"1.0","canonical_sha256":"61e73442f93f6c2d33ebd9375a133acbf20425541c629c37f74bc74b52806b65","source":{"kind":"arxiv","id":"2407.04694","version":1},"attestation_state":"computed","paper":{"title":"Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Alexander Meinke, Bilal Chughtai, Jan Betley, Jeremy Scheurer, Kaivalya Hariharan, Marius Hobbhahn, Mikita Balesni, Owain Evans, Rudolf Laine","submitted_at":"2024-07-05T17:57:02Z","abstract_excerpt":"AI assistants such as ChatGPT are trained to respond to users by saying, \"I am a large language model\". This raises questions. Do such models know that they are LLMs and reliably act on this knowledge? Are they aware of their current circumstances, such as being deployed to the public? We refer to a model's knowledge of itself and its circumstances as situational awareness. To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the $\\textbf{Situational Awareness Dataset (SAD)}$, a benchmark co"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.04694","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-07-05T17:57:02Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"2453a0cd98ce1b95b305eb2b14f71701ab14f583a3485815a25e30e50d674277","abstract_canon_sha256":"a44fff10f8af127f567b726e788335a49a27fd709edaa94bc6eeaf7aaeffb440"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:40:37.571714Z","signature_b64":"9idkxblIHfofBsfduCMh4fwb1U8WakmDeKpyGSF2SZY7oQsW/uhILtS0/h2lLTLeS3G1eeeVJ6bukCUHSOeUDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"61e73442f93f6c2d33ebd9375a133acbf20425541c629c37f74bc74b52806b65","last_reissued_at":"2026-07-05T08:40:37.571260Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:40:37.571260Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Alexander Meinke, Bilal Chughtai, Jan Betley, Jeremy Scheurer, Kaivalya Hariharan, Marius Hobbhahn, Mikita Balesni, Owain Evans, Rudolf Laine","submitted_at":"2024-07-05T17:57:02Z","abstract_excerpt":"AI assistants such as ChatGPT are trained to respond to users by saying, \"I am a large language model\". This raises questions. Do such models know that they are LLMs and reliably act on this knowledge? Are they aware of their current circumstances, such as being deployed to the public? We refer to a model's knowledge of itself and its circumstances as situational awareness. To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the $\\textbf{Situational Awareness Dataset (SAD)}$, a benchmark co"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.04694","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.04694/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.04694","created_at":"2026-07-05T08:40:37.571322+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.04694v1","created_at":"2026-07-05T08:40:37.571322+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.04694","created_at":"2026-07-05T08:40:37.571322+00:00"},{"alias_kind":"pith_short_12","alias_value":"MHTTIQXZH5WC","created_at":"2026-07-05T08:40:37.571322+00:00"},{"alias_kind":"pith_short_16","alias_value":"MHTTIQXZH5WC2M7L","created_at":"2026-07-05T08:40:37.571322+00:00"},{"alias_kind":"pith_short_8","alias_value":"MHTTIQXZ","created_at":"2026-07-05T08:40:37.571322+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":11,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.00036","citing_title":"AI Integrity: Defending Against Backdoors and Secret Loyalties","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23583","citing_title":"Evaluation Awareness Is Not One Capability: Evidence from Open Language Models","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10740","citing_title":"When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09711","citing_title":"Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20382","citing_title":"Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29196","citing_title":"Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2410.02064","citing_title":"Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20382","citing_title":"Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2412.04984","citing_title":"Frontier Models are Capable of In-context Scheming","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13301","citing_title":"Honeypot Protocol","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06327","citing_title":"Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity","ref_index":25,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP","json":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP.json","graph_json":"https://pith.science/api/pith-number/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/graph.json","events_json":"https://pith.science/api/pith-number/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/events.json","paper":"https://pith.science/paper/MHTTIQXZ"},"agent_actions":{"view_html":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP","download_json":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP.json","view_paper":"https://pith.science/paper/MHTTIQXZ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.04694&json=true","fetch_graph":"https://pith.science/api/pith-number/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/graph.json","fetch_events":"https://pith.science/api/pith-number/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/action/storage_attestation","attest_author":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/action/author_attestation","sign_citation":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/action/citation_signature","submit_replication":"https://pith.science/pith/MHTTIQXZH5WC2M7L3E3VUEZ2ZP/action/replication_record"}},"created_at":"2026-07-05T08:40:37.571322+00:00","updated_at":"2026-07-05T08:40:37.571322+00:00"}