{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:RT5K7LJUD4E55U7EWYBVZZG7LX","short_pith_number":"pith:RT5K7LJU","schema_version":"1.0","canonical_sha256":"8cfaafad341f09ded3e4b6035ce4df5dff26cf112a98520a805b320005f67594","source":{"kind":"arxiv","id":"2507.05246","version":1},"attestation_state":"computed","paper":{"title":"When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.AI","authors_text":"David K. Elson, Erik Jenner, Heng Chen, Irhum Shafkat, Rif A. Saurous, Rohin Shah, Scott Emmons, Senthooran Rajamanoharan","submitted_at":"2025-07-07T17:54:52Z","abstract_excerpt":"While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on \"unfaithfulness\" has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rationalization in applications like auditing for bias. However, for the distinct problem of runtime monitoring to prevent severe harm, we argue the key property is not faithfulness but monitorability. To this end, we introduce a conceptual framework distinguishing CoT-as-rationalization from CoT-as-computation. We expect that certain classes of severe harm will "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.05246","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-07-07T17:54:52Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"b6806490fe0c35aa37417f80ed35b11733dd28015f9490215a3f471a7bebc160","abstract_canon_sha256":"987c5c557f13320f74a952395e93058ce770df698f9c2d4fc2cad68cab7df09d"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:33:11.160386Z","signature_b64":"cq2zP6VsPKMcBx4fuM5uQUBDn2blBQ/PQksOiPkOuZGclBROILhIhRmdzs9F4EFHazXc7rxi37C9Jr94CwuZDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"8cfaafad341f09ded3e4b6035ce4df5dff26cf112a98520a805b320005f67594","last_reissued_at":"2026-07-05T11:33:11.159942Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:33:11.159942Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.AI","authors_text":"David K. Elson, Erik Jenner, Heng Chen, Irhum Shafkat, Rif A. Saurous, Rohin Shah, Scott Emmons, Senthooran Rajamanoharan","submitted_at":"2025-07-07T17:54:52Z","abstract_excerpt":"While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on \"unfaithfulness\" has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rationalization in applications like auditing for bias. However, for the distinct problem of runtime monitoring to prevent severe harm, we argue the key property is not faithfulness but monitorability. To this end, we introduce a conceptual framework distinguishing CoT-as-rationalization from CoT-as-computation. We expect that certain classes of severe harm will "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.05246","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.05246/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.05246","created_at":"2026-07-05T11:33:11.159995+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.05246v1","created_at":"2026-07-05T11:33:11.159995+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.05246","created_at":"2026-07-05T11:33:11.159995+00:00"},{"alias_kind":"pith_short_12","alias_value":"RT5K7LJUD4E5","created_at":"2026-07-05T11:33:11.159995+00:00"},{"alias_kind":"pith_short_16","alias_value":"RT5K7LJUD4E55U7E","created_at":"2026-07-05T11:33:11.159995+00:00"},{"alias_kind":"pith_short_8","alias_value":"RT5K7LJU","created_at":"2026-07-05T11:33:11.159995+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08066","citing_title":"Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring","ref_index":85,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25013","citing_title":"Do Thinking Tokens Help with Safety?","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20560","citing_title":"How Transparent is DiffusionGemma?","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07157","citing_title":"Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25052","citing_title":"Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28742","citing_title":"CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07157","citing_title":"Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18549","citing_title":"Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12746","citing_title":"CoT-Guard: Small Models for Strong Monitoring","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03121","citing_title":"An Independent Safety Evaluation of Kimi K2.5","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11746","citing_title":"When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.06427","citing_title":"The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15726","citing_title":"LLM Reasoning Is Latent, Not the Chain of Thought","ref_index":23,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX","json":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX.json","graph_json":"https://pith.science/api/pith-number/RT5K7LJUD4E55U7EWYBVZZG7LX/graph.json","events_json":"https://pith.science/api/pith-number/RT5K7LJUD4E55U7EWYBVZZG7LX/events.json","paper":"https://pith.science/paper/RT5K7LJU"},"agent_actions":{"view_html":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX","download_json":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX.json","view_paper":"https://pith.science/paper/RT5K7LJU","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.05246&json=true","fetch_graph":"https://pith.science/api/pith-number/RT5K7LJUD4E55U7EWYBVZZG7LX/graph.json","fetch_events":"https://pith.science/api/pith-number/RT5K7LJUD4E55U7EWYBVZZG7LX/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX/action/storage_attestation","attest_author":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX/action/author_attestation","sign_citation":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX/action/citation_signature","submit_replication":"https://pith.science/pith/RT5K7LJUD4E55U7EWYBVZZG7LX/action/replication_record"}},"created_at":"2026-07-05T11:33:11.159995+00:00","updated_at":"2026-07-05T11:33:11.159995+00:00"}