{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:HZYNOVGGFJRJJ3GFEQVOMTDL3H","short_pith_number":"pith:HZYNOVGG","schema_version":"1.0","canonical_sha256":"3e70d754c62a6294ecc5242ae64c6bd9dfa888ce67336725ca2041448e496a27","source":{"kind":"arxiv","id":"2310.13798","version":1},"attestation_state":"computed","paper":{"title":"Specific versus General Principles for Constitutional AI","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Amanda Askell, Andrew Callahan, Anna Chen, Anna Goldie, Avital Balwit, Azalia Mirhoseini, Brayden McLean, Cassie Evraets, Catherine Olsson, Eli Tran-Johnson, Esin Durmus, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Kaplan, Kamal Ndousse, Karina Nguyen, Nelson Elhage, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Shauna Kravec, S\\\"oren Mindermann, Thomas I. Liao, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds","submitted_at":"2023-10-20T20:12:45Z","abstract_excerpt":"Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative, replacing human feedback with feedback from AI models conditioned only on a list of written principles. We find this approach effectively prevents the expression of such behaviors. The success of simple principles motivates us to ask: can models learn general ethical behaviors from only a single written principle? To test this, we run experiments using a pr"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.13798","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2023-10-20T20:12:45Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"c2f81b46df48d135a949c730454cb725061eff16e172cc9604e91379cdc0ac23","abstract_canon_sha256":"ad77934a747c8f5dd2b21b740babd95b38cc6c1b5ee636c75ee4aa941f2a1798"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:03:34.948939Z","signature_b64":"XVmIU6xIrvWkMuc3C4QWgY1wsIiqbM713v/cChcN3SLTnpTRfqqwRADJJkZ6TEHwiRxSFU17/aJmtrhD6UwnBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3e70d754c62a6294ecc5242ae64c6bd9dfa888ce67336725ca2041448e496a27","last_reissued_at":"2026-07-05T07:03:34.948454Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:03:34.948454Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Specific versus General Principles for Constitutional AI","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Amanda Askell, Andrew Callahan, Anna Chen, Anna Goldie, Avital Balwit, Azalia Mirhoseini, Brayden McLean, Cassie Evraets, Catherine Olsson, Eli Tran-Johnson, Esin Durmus, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Kaplan, Kamal Ndousse, Karina Nguyen, Nelson Elhage, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Shauna Kravec, S\\\"oren Mindermann, Thomas I. Liao, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds","submitted_at":"2023-10-20T20:12:45Z","abstract_excerpt":"Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative, replacing human feedback with feedback from AI models conditioned only on a list of written principles. We find this approach effectively prevents the expression of such behaviors. The success of simple principles motivates us to ask: can models learn general ethical behaviors from only a single written principle? To test this, we run experiments using a pr"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.13798","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.13798/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.13798","created_at":"2026-07-05T07:03:34.948517+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.13798v1","created_at":"2026-07-05T07:03:34.948517+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.13798","created_at":"2026-07-05T07:03:34.948517+00:00"},{"alias_kind":"pith_short_12","alias_value":"HZYNOVGGFJRJ","created_at":"2026-07-05T07:03:34.948517+00:00"},{"alias_kind":"pith_short_16","alias_value":"HZYNOVGGFJRJJ3GF","created_at":"2026-07-05T07:03:34.948517+00:00"},{"alias_kind":"pith_short_8","alias_value":"HZYNOVGG","created_at":"2026-07-05T07:03:34.948517+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.24014","citing_title":"Reinforcement Learning Towards Broadly and Persistently Beneficial Models","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09475","citing_title":"Emergent alignment and the projectability of ethical personas","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28182","citing_title":"LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2401.05561","citing_title":"TrustLLM: Trustworthiness in Large Language Models","ref_index":170,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17663","citing_title":"ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data","ref_index":8,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H","json":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H.json","graph_json":"https://pith.science/api/pith-number/HZYNOVGGFJRJJ3GFEQVOMTDL3H/graph.json","events_json":"https://pith.science/api/pith-number/HZYNOVGGFJRJJ3GFEQVOMTDL3H/events.json","paper":"https://pith.science/paper/HZYNOVGG"},"agent_actions":{"view_html":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H","download_json":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H.json","view_paper":"https://pith.science/paper/HZYNOVGG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.13798&json=true","fetch_graph":"https://pith.science/api/pith-number/HZYNOVGGFJRJJ3GFEQVOMTDL3H/graph.json","fetch_events":"https://pith.science/api/pith-number/HZYNOVGGFJRJJ3GFEQVOMTDL3H/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H/action/timestamp_anchor","attest_storage":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H/action/storage_attestation","attest_author":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H/action/author_attestation","sign_citation":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H/action/citation_signature","submit_replication":"https://pith.science/pith/HZYNOVGGFJRJJ3GFEQVOMTDL3H/action/replication_record"}},"created_at":"2026-07-05T07:03:34.948517+00:00","updated_at":"2026-07-05T07:03:34.948517+00:00"}