{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2016:7RXHLTSA6WBXBFU5C5T7B2HIVX","short_pith_number":"pith:7RXHLTSA","schema_version":"1.0","canonical_sha256":"fc6e75ce40f58370969d1767f0e8e8aded7700bd2f9f8f8db102c8e18e880861","source":{"kind":"arxiv","id":"1606.06565","version":2},"attestation_state":"computed","paper":{"title":"Concrete Problems in AI Safety","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"The main risks of accidents in AI systems come from five specific problems related to their objectives and learning processes.","cross_cats":["cs.LG"],"primary_cat":"cs.AI","authors_text":"Chris Olah, Dan Man\\'e, Dario Amodei, Jacob Steinhardt, John Schulman, Paul Christiano","submitted_at":"2016-06-21T13:37:05Z","abstract_excerpt":"Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emerge from poor design of real-world AI systems. We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function (\"avoiding side effects\" and \"avoiding reward hacking\"), a"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"1606.06565","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2016-06-21T13:37:05Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"1f6690caddafcd60afbe1f25438128e9c86db88aadc9111efb6ef75f1d2a2853","abstract_canon_sha256":"55a75fe50f92a5a645d859c068bd60179cf33cebef87e74afec9cbf81a3c66d4"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-04T21:08:58.560373Z","signature_b64":"myAjmdyyqkldoBCHrunnruvXblKj7QCAH49IZNkCKhZzIu9b1rTQs58ICTUPTQe3iji/FHJRu5bHd3/8smr6Ag==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fc6e75ce40f58370969d1767f0e8e8aded7700bd2f9f8f8db102c8e18e880861","last_reissued_at":"2026-07-04T21:08:58.559832Z","signature_status":"signed_v1","first_computed_at":"2026-07-04T21:08:58.559832Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Concrete Problems in AI Safety","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"The main risks of accidents in AI systems come from five specific problems related to their objectives and learning processes.","cross_cats":["cs.LG"],"primary_cat":"cs.AI","authors_text":"Chris Olah, Dan Man\\'e, Dario Amodei, Jacob Steinhardt, John Schulman, Paul Christiano","submitted_at":"2016-06-21T13:37:05Z","abstract_excerpt":"Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emerge from poor design of real-world AI systems. We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function (\"avoiding side effects\" and \"avoiding reward hacking\"), a"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function (avoiding side effects and avoiding reward hacking), an objective function that is too expensive to evaluate frequently (scalable supervision), or undesirable behavior during the learning process (safe exploration and distributional shift).","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That these five problems represent the primary and most actionable sources of accident risk in real-world AI systems, and that addressing them will substantially mitigate unintended harmful behavior without needing to consider additional unlisted factors.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"The paper categorizes five concrete AI safety problems arising from flawed objectives, costly evaluation, and learning dynamics.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"The main risks of accidents in AI systems come from five specific problems related to their objectives and learning processes.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"93b5cd905e6f730322cb616ab3b304d4ca3ecbbc83b35017dabb9488e6ebd5be"},"source":{"id":"1606.06565","kind":"arxiv","version":2},"verdict":{"id":"a8ff4245-0af2-4f2a-a28c-534f06d74367","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T05:12:03.516148Z","strongest_claim":"We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function (avoiding side effects and avoiding reward hacking), an objective function that is too expensive to evaluate frequently (scalable supervision), or undesirable behavior during the learning process (safe exploration and distributional shift).","one_line_summary":"The paper categorizes five concrete AI safety problems arising from flawed objectives, costly evaluation, and learning dynamics.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That these five problems represent the primary and most actionable sources of accident risk in real-world AI systems, and that addressing them will substantially mitigate unintended harmful behavior without needing to consider additional unlisted factors.","pith_extraction_headline":"The main risks of accidents in AI systems come from five specific problems related to their objectives and learning processes."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1606.06565/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":171,"sample":[{"doi":"","year":2016,"title":"Deep Learning with Diﬀerential Privacy","work_id":"c3e5aaa4-afed-43cb-af07-1544aab33983","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2005,"title":"Exploration and apprenticeship learning in reinforcement learning","work_id":"47934461-842a-4e37-b8c2-1b850b8ee921","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2015,"title":"The Hidden Cost of Eﬃciency: Fairness and Discrimination in Predictive Modeling","work_id":"b40771aa-0a39-4e2f-a786-a8c9bdf1ae1d","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2014,"title":"Taming the monster: A fast and simple algorithm for contextual ban- dits","work_id":"aa8b57e7-8fcc-4a21-939e-00381928e5aa","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2014,"title":"Domain-Adversarial Neural Networks","work_id":"87a8214d-ecfa-4ce2-adca-a4112a1e8af7","ref_index":5,"cited_arxiv_id":"1412.4446","is_internal_anchor":false}],"resolved_work":171,"snapshot_sha256":"a7a9d780951b86c4e2cb70f79193693d2ad271d5ea9402c4f1af5612e5fc16b7","internal_anchors":13},"formal_canon":{"evidence_count":2,"snapshot_sha256":"1a7f014e40a63026d1574d64dc7d3f899dc965414c0a7e3da078b2d740221bd0"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"1606.06565","created_at":"2026-07-04T21:08:58.559905+00:00"},{"alias_kind":"arxiv_version","alias_value":"1606.06565v2","created_at":"2026-07-04T21:08:58.559905+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.1606.06565","created_at":"2026-07-04T21:08:58.559905+00:00"},{"alias_kind":"pith_short_12","alias_value":"7RXHLTSA6WBX","created_at":"2026-07-04T21:08:58.559905+00:00"},{"alias_kind":"pith_short_16","alias_value":"7RXHLTSA6WBXBFU5","created_at":"2026-07-04T21:08:58.559905+00:00"},{"alias_kind":"pith_short_8","alias_value":"7RXHLTSA","created_at":"2026-07-04T21:08:58.559905+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":245,"internal_anchor_count":245,"sample":[{"citing_arxiv_id":"2607.07023","citing_title":"Online Data Selection Is Implicit Alignment","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07040","citing_title":"Measuring Intelligence Beyond Human Scale","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07695","citing_title":"Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06175","citing_title":"Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06196","citing_title":"Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability","ref_index":129,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25996","citing_title":"Autodata: An agentic data scientist to create high quality synthetic data","ref_index":49,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25870","citing_title":"Evolving Quantum Error-Correcting Encodings for Molecular Simulation","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25239","citing_title":"Tensor-Based Batch Fuzzing with Adaptive Perturbation Scaling for Deep Neural Networks","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26366","citing_title":"Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27079","citing_title":"ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25996","citing_title":"Autodata: An agentic data scientist to create high quality synthetic data","ref_index":49,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26529","citing_title":"The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24014","citing_title":"Reinforcement Learning Towards Broadly and Persistently Beneficial Models","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23991","citing_title":"Critique of Agent Model","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23280","citing_title":"Causal Reward World Models: Zero-shot Reward Design for Automated Skill Generation","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23094","citing_title":"Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22159","citing_title":"Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21939","citing_title":"Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19924","citing_title":"The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19818","citing_title":"Uncertainty-Aware Reward Modeling for Stable RLHF","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20724","citing_title":"When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17682","citing_title":"From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01277","citing_title":"Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety","ref_index":43,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01356","citing_title":"Chameleon: Recovering Cyber-Physical Systems from Memory Corruption Attacks via ML Surrogates","ref_index":41,"is_internal_anchor":true},{"citing_arxiv_id":"2607.02291","citing_title":"Optimizing Visual Generative Models via Distribution-wise Rewards","ref_index":1,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX","json":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX.json","graph_json":"https://pith.science/api/pith-number/7RXHLTSA6WBXBFU5C5T7B2HIVX/graph.json","events_json":"https://pith.science/api/pith-number/7RXHLTSA6WBXBFU5C5T7B2HIVX/events.json","paper":"https://pith.science/paper/7RXHLTSA"},"agent_actions":{"view_html":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX","download_json":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX.json","view_paper":"https://pith.science/paper/7RXHLTSA","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=1606.06565&json=true","fetch_graph":"https://pith.science/api/pith-number/7RXHLTSA6WBXBFU5C5T7B2HIVX/graph.json","fetch_events":"https://pith.science/api/pith-number/7RXHLTSA6WBXBFU5C5T7B2HIVX/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX/action/storage_attestation","attest_author":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX/action/author_attestation","sign_citation":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX/action/citation_signature","submit_replication":"https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX/action/replication_record"}},"created_at":"2026-07-04T21:08:58.559905+00:00","updated_at":"2026-07-04T21:08:58.559905+00:00"}