{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:YBKWPHMWPOAJUM7XJVEGVRUA6F","short_pith_number":"pith:YBKWPHMW","schema_version":"1.0","canonical_sha256":"c055679d967b809a33f74d486ac680f14068a630b96098edffb4beed897e4f3b","source":{"kind":"arxiv","id":"2210.01790","version":2},"attestation_state":"computed","paper":{"title":"Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Jonathan Uesato, Mary Phuong, Ramana Kumar, Rohin Shah, Victoria Krakovna, Vikrant Varma, Zac Kenton","submitted_at":"2022-10-04T17:57:53Z","abstract_excerpt":"The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2210.01790","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2022-10-04T17:57:53Z","cross_cats_sorted":[],"title_canon_sha256":"797d328c1d2ca1060917dcc0d47753ce74083357807397b914ebe7b04fad3bf4","abstract_canon_sha256":"71eedd8a4aa1334cbb99090758ed97db0a83027cc34c531dff02d0e0af5266cc"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T05:12:44.613919Z","signature_b64":"Rgt0RMc7uNclepOBbGNhaLuHE19xihlrpWTBlvcof7gROVNHvYf6/VmV3LC5EYTn7dh04ODbI4Cnu1YLGeQgDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"c055679d967b809a33f74d486ac680f14068a630b96098edffb4beed897e4f3b","last_reissued_at":"2026-07-05T05:12:44.613444Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T05:12:44.613444Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Jonathan Uesato, Mary Phuong, Ramana Kumar, Rohin Shah, Victoria Krakovna, Vikrant Varma, Zac Kenton","submitted_at":"2022-10-04T17:57:53Z","abstract_excerpt":"The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2210.01790","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2210.01790/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2210.01790","created_at":"2026-07-05T05:12:44.613502+00:00"},{"alias_kind":"arxiv_version","alias_value":"2210.01790v2","created_at":"2026-07-05T05:12:44.613502+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2210.01790","created_at":"2026-07-05T05:12:44.613502+00:00"},{"alias_kind":"pith_short_12","alias_value":"YBKWPHMWPOAJ","created_at":"2026-07-05T05:12:44.613502+00:00"},{"alias_kind":"pith_short_16","alias_value":"YBKWPHMWPOAJUM7X","created_at":"2026-07-05T05:12:44.613502+00:00"},{"alias_kind":"pith_short_8","alias_value":"YBKWPHMW","created_at":"2026-07-05T05:12:44.613502+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.21939","citing_title":"Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27593","citing_title":"Voluntary Collusion with Secret Tools in Competing LLM Agents","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2605.23565","citing_title":"Understanding Goal Generalisation in Sequential Reinforcement Learning","ref_index":60,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12673","citing_title":"Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12809","citing_title":"Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces","ref_index":224,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11134","citing_title":"Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06992","citing_title":"Why Does Agentic Safety Fail to Generalize Across Tasks?","ref_index":98,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F","json":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F.json","graph_json":"https://pith.science/api/pith-number/YBKWPHMWPOAJUM7XJVEGVRUA6F/graph.json","events_json":"https://pith.science/api/pith-number/YBKWPHMWPOAJUM7XJVEGVRUA6F/events.json","paper":"https://pith.science/paper/YBKWPHMW"},"agent_actions":{"view_html":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F","download_json":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F.json","view_paper":"https://pith.science/paper/YBKWPHMW","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2210.01790&json=true","fetch_graph":"https://pith.science/api/pith-number/YBKWPHMWPOAJUM7XJVEGVRUA6F/graph.json","fetch_events":"https://pith.science/api/pith-number/YBKWPHMWPOAJUM7XJVEGVRUA6F/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F/action/timestamp_anchor","attest_storage":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F/action/storage_attestation","attest_author":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F/action/author_attestation","sign_citation":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F/action/citation_signature","submit_replication":"https://pith.science/pith/YBKWPHMWPOAJUM7XJVEGVRUA6F/action/replication_record"}},"created_at":"2026-07-05T05:12:44.613502+00:00","updated_at":"2026-07-05T05:12:44.613502+00:00"}