Pith. sign in
Pith Number

pith:7RXHLTSA

pith:2016:7RXHLTSA6WBXBFU5C5T7B2HIVX
not attested not anchored not stored refs resolved

Concrete Problems in AI Safety

Chris Olah, Dan Man\'e, Dario Amodei, Jacob Steinhardt, John Schulman, Paul Christiano

The main risks of accidents in AI systems come from five specific problems related to their objectives and learning processes.

arxiv:1606.06565 v2 · 2016-06-21 · cs.AI · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{7RXHLTSA6WBXBFU5C5T7B2HIVX}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function (avoiding side effects and avoiding reward hacking), an objective function that is too expensive to evaluate frequently (scalable supervision), or undesirable behavior during the learning process (safe exploration and distributional shift).

C2weakest assumption

That these five problems represent the primary and most actionable sources of accident risk in real-world AI systems, and that addressing them will substantially mitigate unintended harmful behavior without needing to consider additional unlisted factors.

C3one line summary

The paper categorizes five concrete AI safety problems arising from flawed objectives, costly evaluation, and learning dynamics.

References

171 extracted · 171 resolved · 13 Pith anchors

[1] Deep Learning with Differential Privacy 2016
[2] Exploration and apprenticeship learning in reinforcement learning 2005
[3] The Hidden Cost of Efficiency: Fairness and Discrimination in Predictive Modeling 2015
[4] Taming the monster: A fast and simple algorithm for contextual ban- dits 2014
[5] Domain-Adversarial Neural Networks 2014 · arXiv:1412.4446

Formal links

2 machine-checked theorem links

Cited by

245 papers in Pith

Receipt and verification
First computed 2026-07-04T21:08:58.559832Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

fc6e75ce40f58370969d1767f0e8e8aded7700bd2f9f8f8db102c8e18e880861

Aliases

arxiv: 1606.06565 · arxiv_version: 1606.06565v2 · doi: 10.48550/arxiv.1606.06565 · pith_short_12: 7RXHLTSA6WBX · pith_short_16: 7RXHLTSA6WBXBFU5 · pith_short_8: 7RXHLTSA
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/7RXHLTSA6WBXBFU5C5T7B2HIVX \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: fc6e75ce40f58370969d1767f0e8e8aded7700bd2f9f8f8db102c8e18e880861
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "55a75fe50f92a5a645d859c068bd60179cf33cebef87e74afec9cbf81a3c66d4",
    "cross_cats_sorted": [
      "cs.LG"
    ],
    "license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
    "primary_cat": "cs.AI",
    "submitted_at": "2016-06-21T13:37:05Z",
    "title_canon_sha256": "1f6690caddafcd60afbe1f25438128e9c86db88aadc9111efb6ef75f1d2a2853"
  },
  "schema_version": "1.0",
  "source": {
    "id": "1606.06565",
    "kind": "arxiv",
    "version": 2
  }
}