{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:3OXJAEW3BO75YJZ6FUPBZMUYCA","short_pith_number":"pith:3OXJAEW3","schema_version":"1.0","canonical_sha256":"dbae9012db0bbfdc273e2d1e1cb29810280578f2b50328202b53ffd9924e9473","source":{"kind":"arxiv","id":"2503.10965","version":2},"attestation_state":"computed","paper":{"title":"Auditing language models for hidden objectives","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.AI","authors_text":"Adam Jermyn, Adam Pearce, Akbir Khan, Andy Shih, Austin Meek, Brian Chen, Carson Denison, Christopher Olah, Daniel Ziegler, Drake Thomas, Emmanuel Ameisen, Euan Ong, Evan Hubinger, Fabien Roger, Florian Dietz, Hoagy Cunningham, Jack Lindsey, Jan Kirchner, Jan Leike, Jeanne Salle, Johannes Treutlein, Jonathan Marcus, Joshua Batson, Kei Nishimura-Gasparian, Kelley Rivoire, Meg Tong, Monte MacDiarmid, Samuel Marks, Samuel R. Bowman, Satvik Golechha, Shan Carter, Siddharth Mishra-Sharma, Tim Belonax, Tom Henighan, Trenton Bricken","submitted_at":"2025-03-14T00:21:15Z","abstract_excerpt":"We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objective. Our training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training. We leverage this model to study alignment audits in two ways. First, we conduct a "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2503.10965","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-03-14T00:21:15Z","cross_cats_sorted":["cs.CL","cs.LG"],"title_canon_sha256":"3a096fd8507dac1e114cfbac3e33524699e3b761503dae2162ea771bdf595c35","abstract_canon_sha256":"178ff587c594d7f86229d1bd06c205ae24ffc12fc3cbc39f315406f8871509df"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:40:39.413246Z","signature_b64":"7qAUPFylyLsqgt7pob42kyBmRD+zJhcqQrmqL+l6DmQLJZK0EMDBXt0/o6Zki0MHPpyRL3hZ8RXZcupW1iSRBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"dbae9012db0bbfdc273e2d1e1cb29810280578f2b50328202b53ffd9924e9473","last_reissued_at":"2026-07-05T10:40:39.412749Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:40:39.412749Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Auditing language models for hidden objectives","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.AI","authors_text":"Adam Jermyn, Adam Pearce, Akbir Khan, Andy Shih, Austin Meek, Brian Chen, Carson Denison, Christopher Olah, Daniel Ziegler, Drake Thomas, Emmanuel Ameisen, Euan Ong, Evan Hubinger, Fabien Roger, Florian Dietz, Hoagy Cunningham, Jack Lindsey, Jan Kirchner, Jan Leike, Jeanne Salle, Johannes Treutlein, Jonathan Marcus, Joshua Batson, Kei Nishimura-Gasparian, Kelley Rivoire, Meg Tong, Monte MacDiarmid, Samuel Marks, Samuel R. Bowman, Satvik Golechha, Shan Carter, Siddharth Mishra-Sharma, Tim Belonax, Tom Henighan, Trenton Bricken","submitted_at":"2025-03-14T00:21:15Z","abstract_excerpt":"We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objective. Our training pipeline first teaches the model about exploitable errors in RLHF reward models (RMs), then trains the model to exploit some of these errors. We verify via out-of-distribution evaluations that the model generalizes to exhibit whatever behaviors it believes RMs rate highly, including ones not reinforced during training. We leverage this model to study alignment audits in two ways. First, we conduct a "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2503.10965","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2503.10965/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2503.10965","created_at":"2026-07-05T10:40:39.412806+00:00"},{"alias_kind":"arxiv_version","alias_value":"2503.10965v2","created_at":"2026-07-05T10:40:39.412806+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2503.10965","created_at":"2026-07-05T10:40:39.412806+00:00"},{"alias_kind":"pith_short_12","alias_value":"3OXJAEW3BO75","created_at":"2026-07-05T10:40:39.412806+00:00"},{"alias_kind":"pith_short_16","alias_value":"3OXJAEW3BO75YJZ6","created_at":"2026-07-05T10:40:39.412806+00:00"},{"alias_kind":"pith_short_8","alias_value":"3OXJAEW3","created_at":"2026-07-05T10:40:39.412806+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":20,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.22019","citing_title":"Channel Location Constrains the Auditability of Subliminal Learning","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18327","citing_title":"Self-CTRL: Self-Consistency Training with Reinforcement Learning","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2606.13310","citing_title":"RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12618","citing_title":"\"Did you lie?\" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms","ref_index":82,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09711","citing_title":"Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization","ref_index":250,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01033","citing_title":"The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00994","citing_title":"Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06846","citing_title":"Narrow Secret Loyalty Dodges Black-Box Audits","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15164","citing_title":"Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09563","citing_title":"PRISM: Recovering Instruction Sets from Language Model Activations","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21602","citing_title":"Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs","ref_index":88,"is_internal_anchor":false},{"citing_arxiv_id":"2512.05742","citing_title":"Internal Deployment in the AI Act","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10310","citing_title":"Positive Alignment: Artificial Intelligence for Human Flourishing","ref_index":125,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12813","citing_title":"REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06846","citing_title":"Narrow Secret Loyalty Dodges Black-Box Audits","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11448","citing_title":"Deep Minds and Shallow Probes","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00994","citing_title":"Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11061","citing_title":"Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06846","citing_title":"Narrow Secret Loyalty Dodges Black-Box Audits","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13602","citing_title":"Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges","ref_index":80,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA","json":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA.json","graph_json":"https://pith.science/api/pith-number/3OXJAEW3BO75YJZ6FUPBZMUYCA/graph.json","events_json":"https://pith.science/api/pith-number/3OXJAEW3BO75YJZ6FUPBZMUYCA/events.json","paper":"https://pith.science/paper/3OXJAEW3"},"agent_actions":{"view_html":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA","download_json":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA.json","view_paper":"https://pith.science/paper/3OXJAEW3","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2503.10965&json=true","fetch_graph":"https://pith.science/api/pith-number/3OXJAEW3BO75YJZ6FUPBZMUYCA/graph.json","fetch_events":"https://pith.science/api/pith-number/3OXJAEW3BO75YJZ6FUPBZMUYCA/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA/action/timestamp_anchor","attest_storage":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA/action/storage_attestation","attest_author":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA/action/author_attestation","sign_citation":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA/action/citation_signature","submit_replication":"https://pith.science/pith/3OXJAEW3BO75YJZ6FUPBZMUYCA/action/replication_record"}},"created_at":"2026-07-05T10:40:39.412806+00:00","updated_at":"2026-07-05T10:40:39.412806+00:00"}