{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:PBGWULRVXWVH35T7Z5EYG7CIMN","short_pith_number":"pith:PBGWULRV","schema_version":"1.0","canonical_sha256":"784d6a2e35bdaa7df67fcf49837c48637f0a560e4141985e9b758f64d850aaab","source":{"kind":"arxiv","id":"2304.11082","version":6},"attestation_state":"computed","paper":{"title":"Fundamental Limitations of Alignment in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Amnon Shashua, Noam Wies, Oshri Avnery, Yoav Levine, Yotam Wolf","submitted_at":"2023-04-19T17:50:09Z","abstract_excerpt":"An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by tuning the model in a way that enhances desired behaviors and inhibits undesired ones, a process referred to as alignment. In this paper, we propose a theoretical approach called Behavior Expectation Bounds (BEB) which allows us to formally investigate several inherent characteristics and limitations of alignment in large language models. Importantly, we prove that within the limits of this framework, for any behavi"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2304.11082","kind":"arxiv","version":6},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2023-04-19T17:50:09Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"ed632885193e72f26bb874420cf746a3d1cb85a092b39b4798b724373628e6bb","abstract_canon_sha256":"9a5f9653e691e40249504335f880005d4772ab107e17c0ce441bd0c79688450e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:26:11.321221Z","signature_b64":"3IHyQ8GlylB5elPsCRTkCRgSm3TuLNk2I2+1bDJV6P0bKVuWYzZd5Rs8W3/Y2xmRtYkpZhr+cFnsnAuqM1oaAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"784d6a2e35bdaa7df67fcf49837c48637f0a560e4141985e9b758f64d850aaab","last_reissued_at":"2026-07-05T08:26:11.320735Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:26:11.320735Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Fundamental Limitations of Alignment in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Amnon Shashua, Noam Wies, Oshri Avnery, Yoav Levine, Yotam Wolf","submitted_at":"2023-04-19T17:50:09Z","abstract_excerpt":"An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by tuning the model in a way that enhances desired behaviors and inhibits undesired ones, a process referred to as alignment. In this paper, we propose a theoretical approach called Behavior Expectation Bounds (BEB) which allows us to formally investigate several inherent characteristics and limitations of alignment in large language models. Importantly, we prove that within the limits of this framework, for any behavi"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2304.11082","kind":"arxiv","version":6},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2304.11082/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2304.11082","created_at":"2026-07-05T08:26:11.320795+00:00"},{"alias_kind":"arxiv_version","alias_value":"2304.11082v6","created_at":"2026-07-05T08:26:11.320795+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2304.11082","created_at":"2026-07-05T08:26:11.320795+00:00"},{"alias_kind":"pith_short_12","alias_value":"PBGWULRVXWVH","created_at":"2026-07-05T08:26:11.320795+00:00"},{"alias_kind":"pith_short_16","alias_value":"PBGWULRVXWVH35T7","created_at":"2026-07-05T08:26:11.320795+00:00"},{"alias_kind":"pith_short_8","alias_value":"PBGWULRV","created_at":"2026-07-05T08:26:11.320795+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.12234","citing_title":"On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08367","citing_title":"Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25739","citing_title":"The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30169","citing_title":"Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms","ref_index":136,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00485","citing_title":"Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2307.15043","citing_title":"Universal and Transferable Adversarial Attacks on Aligned Language Models","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20641","citing_title":"Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16339","citing_title":"Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2512.10100","citing_title":"Robust AI Security and Alignment: A Sisyphean Endeavor?","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12809","citing_title":"Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces","ref_index":206,"is_internal_anchor":false},{"citing_arxiv_id":"2307.02483","citing_title":"Jailbroken: How Does LLM Safety Training Fail?","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08496","citing_title":"Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17359","citing_title":"Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations","ref_index":29,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN","json":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN.json","graph_json":"https://pith.science/api/pith-number/PBGWULRVXWVH35T7Z5EYG7CIMN/graph.json","events_json":"https://pith.science/api/pith-number/PBGWULRVXWVH35T7Z5EYG7CIMN/events.json","paper":"https://pith.science/paper/PBGWULRV"},"agent_actions":{"view_html":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN","download_json":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN.json","view_paper":"https://pith.science/paper/PBGWULRV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2304.11082&json=true","fetch_graph":"https://pith.science/api/pith-number/PBGWULRVXWVH35T7Z5EYG7CIMN/graph.json","fetch_events":"https://pith.science/api/pith-number/PBGWULRVXWVH35T7Z5EYG7CIMN/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN/action/timestamp_anchor","attest_storage":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN/action/storage_attestation","attest_author":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN/action/author_attestation","sign_citation":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN/action/citation_signature","submit_replication":"https://pith.science/pith/PBGWULRVXWVH35T7Z5EYG7CIMN/action/replication_record"}},"created_at":"2026-07-05T08:26:11.320795+00:00","updated_at":"2026-07-05T08:26:11.320795+00:00"}