{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:ZXF3UVV5K4LWYWP46N57WJEFYR","short_pith_number":"pith:ZXF3UVV5","schema_version":"1.0","canonical_sha256":"cdcbba56bd57176c59fcf37bfb2485c45b886c855a2100cc84c6a8d3d2e521ba","source":{"kind":"arxiv","id":"2602.20021","version":1},"attestation_state":"computed","paper":{"title":"Agents of Chaos","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Autonomous language-model agents exhibit security, privacy, and governance vulnerabilities when given tools, memory, and external access in live settings.","cross_cats":["cs.CY"],"primary_cat":"cs.AI","authors_text":"Adam Belfki, Aditya Ratan Jannali, Alex Loftus, Amir Zur, Aruna Sankaranarayanan, Atai Ambus, Avery Yen, Ayelet Gordon-Tapiero, Can Rager, Christoph Riedl, Chris Wendler, David Atkinson, David Bau, David Manheim, EunJeong Hwang, Gabriele Sarti, Giordano Rogers, Hadas Orgad, Jaden Fiotto-Kaufman, Jannik Brinkmann, Jasmine Cui, Koyena Pal, Maarten Sap, Michael Ripa, Natalie Shapira, Negev Taglicht, Nikhil Prakash, Nitay Alon, Olivia Floody, P Sam Sahil, Reuth Mirsky, Rohit Gandikota, Shiri Oron, Tamar Rott Shaham, Tomer Shabtay, Tomer Ullman, Vered Shwartz, Yotam Kaplan","submitted_at":"2026-02-23T16:28:48Z","abstract_excerpt":"We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2602.20021","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2026-02-23T16:28:48Z","cross_cats_sorted":["cs.CY"],"title_canon_sha256":"2ebb0f0fa8b06db607340b8a0bf3fcd342f6f1997ad569b51d86c6d30b2f2213","abstract_canon_sha256":"aac4832a8c032d3c29e6ad970a5a61827b8c9a22a2736c13ae677b49d0724f76"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:38:53.185807Z","signature_b64":"NRGwqNIRmQSDMFWo7pUVhIBvHxBf9acZAgywkJjZbNT9dvMtO8YtxtH1XEsOicc7BqVGxQ+4I1PqtEIkKWWMBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"cdcbba56bd57176c59fcf37bfb2485c45b886c855a2100cc84c6a8d3d2e521ba","last_reissued_at":"2026-05-17T23:38:53.185153Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:38:53.185153Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Agents of Chaos","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Autonomous language-model agents exhibit security, privacy, and governance vulnerabilities when given tools, memory, and external access in live settings.","cross_cats":["cs.CY"],"primary_cat":"cs.AI","authors_text":"Adam Belfki, Aditya Ratan Jannali, Alex Loftus, Amir Zur, Aruna Sankaranarayanan, Atai Ambus, Avery Yen, Ayelet Gordon-Tapiero, Can Rager, Christoph Riedl, Chris Wendler, David Atkinson, David Bau, David Manheim, EunJeong Hwang, Gabriele Sarti, Giordano Rogers, Hadas Orgad, Jaden Fiotto-Kaufman, Jannik Brinkmann, Jasmine Cui, Koyena Pal, Maarten Sap, Michael Ripa, Natalie Shapira, Negev Taglicht, Nikhil Prakash, Nitay Alon, Olivia Floody, P Sam Sahil, Reuth Mirsky, Rohit Gandikota, Shiri Oron, Tamar Rott Shaham, Tomer Shabtay, Tomer Ullman, Vered Shwartz, Yotam Kaplan","submitted_at":"2026-02-23T16:28:48Z","abstract_excerpt":"We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the specific behaviors observed in this controlled laboratory environment with twenty researchers and particular tool integrations indicate general vulnerabilities that would reliably appear in broader, less controlled real-world deployments.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"An exploratory red-teaming study documents eleven cases of security, privacy, and governance failures in autonomous language-model agents with tool access and persistent memory.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Autonomous language-model agents exhibit security, privacy, and governance vulnerabilities when given tools, memory, and external access in live settings.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"61cb31a4ba5eb5d11e503418a1ce4c782d3193f72d8103a14a9e31f421375de8"},"source":{"id":"2602.20021","kind":"arxiv","version":1},"verdict":{"id":"b8a938da-63bc-4453-8547-3930c2b14300","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-15T06:59:12.239132Z","strongest_claim":"Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings.","one_line_summary":"An exploratory red-teaming study documents eleven cases of security, privacy, and governance failures in autonomous language-model agents with tool access and persistent memory.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the specific behaviors observed in this controlled laboratory environment with twenty researchers and particular tool integrations indicate general vulnerabilities that would reliably appear in broader, less controlled real-world deployments.","pith_extraction_headline":"Autonomous language-model agents exhibit security, privacy, and governance vulnerabilities when given tools, memory, and external access in live settings."},"references":{"count":12,"sample":[{"doi":"","year":2025,"title":"URLhttps://arxiv.org/abs/2510.26707. Matteo Bortoletto, Constantin Ruhdorfer, and Andreas Bulling. Tom-ssi: Evaluating theory of mind in situated social interactions. InProceedings of the 2025 Confere","work_id":"01f40c90-596a-4d92-b27d-3142698ce77a","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2026,"title":"Chen Chen, Kim Young Il, Yuan Yang, Wenhao Su, Yilin Zhang, Xueluan Gong, Qian Wang, Yongsen Zheng, Ziyao Liu, and Kwok-Yan Lam","work_id":"87ff5f50-1c54-4552-bc47-6e61a16b3789","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1145/2844110","year":1987,"title":"URLhttps://arxiv.org/abs/2510.01070. Daniel C. Dennett.The Intentional Stance. The MIT Press, 1987. ISBN 9780262040938. URL https://mitpress.mit.edu/9780262040938/the-intentional-stance/. Nicholas Dia","work_id":"975c75b1-2eaa-4bb4-bfaf-4c559ada5a3f","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training","work_id":"b95e7447-320c-4c85-b5d0-3708cc2cc72e","ref_index":4,"cited_arxiv_id":"2401.05566","is_internal_anchor":true},{"doi":"","year":2025,"title":"Infusing Theory of Mind into Socially Intelligent LLM Agents","work_id":"2b22ebd4-22b5-491c-8efe-c31d8aefc1f7","ref_index":5,"cited_arxiv_id":"2509.22887","is_internal_anchor":true}],"resolved_work":12,"snapshot_sha256":"78e7f353ac8f384d1d09593b0ad4ad3f000730b8e951ae5c8932a09175b683e1","internal_anchors":3},"formal_canon":{"evidence_count":3,"snapshot_sha256":"a9050fad3e308990de0322a6dd549b7b30c93bfcba6650d21db9992d9f0178c2"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2602.20021","created_at":"2026-05-17T23:38:53.185247+00:00"},{"alias_kind":"arxiv_version","alias_value":"2602.20021v1","created_at":"2026-05-17T23:38:53.185247+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2602.20021","created_at":"2026-05-17T23:38:53.185247+00:00"},{"alias_kind":"pith_short_12","alias_value":"ZXF3UVV5K4LW","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_16","alias_value":"ZXF3UVV5K4LWYWP4","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_8","alias_value":"ZXF3UVV5","created_at":"2026-05-18T12:33:37.589309+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":36,"internal_anchor_count":36,"sample":[{"citing_arxiv_id":"2606.22792","citing_title":"The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models","ref_index":85,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21037","citing_title":"Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers","ref_index":48,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12752","citing_title":"Beyond Resilience -- A Conceptual Framework for Civic Ascent","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10484","citing_title":"AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05647","citing_title":"Coding with \"Enemy\": Can Human Developers Detect AI Agent Sabotage?","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01047","citing_title":"Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.09890","citing_title":"PreAct-Bench: Benchmarking Predictive Monitoring in LLMs","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01275","citing_title":"Domination-Avoiding Learning Agents Cannot Collude","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.00341","citing_title":"ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2605.17986","citing_title":"LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.29722","citing_title":"Attraction, Not Adaptation: How AI Agent Communities Develop Distinct Linguistic Identities","ref_index":38,"is_internal_anchor":true},{"citing_arxiv_id":"2605.25435","citing_title":"Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures","ref_index":58,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30169","citing_title":"Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms","ref_index":119,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30258","citing_title":"EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10749","citing_title":"Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation","ref_index":152,"is_internal_anchor":true},{"citing_arxiv_id":"2605.19149","citing_title":"Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2605.17986","citing_title":"LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2603.21354","citing_title":"The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project","ref_index":69,"is_internal_anchor":true},{"citing_arxiv_id":"2603.27771","citing_title":"Emergent Social Intelligence Risks in Generative Multi-Agent Systems","ref_index":116,"is_internal_anchor":true},{"citing_arxiv_id":"2604.02767","citing_title":"SentinelAgent: Intent-Verified Delegation Chains for Securing Federal Multi-Agent AI Systems","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2604.03070","citing_title":"How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11730","citing_title":"Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11003","citing_title":"The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11135","citing_title":"Control Charts for Multi-agent Systems","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2605.08460","citing_title":"When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks","ref_index":4,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":3,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR","json":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR.json","graph_json":"https://pith.science/api/pith-number/ZXF3UVV5K4LWYWP46N57WJEFYR/graph.json","events_json":"https://pith.science/api/pith-number/ZXF3UVV5K4LWYWP46N57WJEFYR/events.json","paper":"https://pith.science/paper/ZXF3UVV5"},"agent_actions":{"view_html":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR","download_json":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR.json","view_paper":"https://pith.science/paper/ZXF3UVV5","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2602.20021&json=true","fetch_graph":"https://pith.science/api/pith-number/ZXF3UVV5K4LWYWP46N57WJEFYR/graph.json","fetch_events":"https://pith.science/api/pith-number/ZXF3UVV5K4LWYWP46N57WJEFYR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR/action/storage_attestation","attest_author":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR/action/author_attestation","sign_citation":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR/action/citation_signature","submit_replication":"https://pith.science/pith/ZXF3UVV5K4LWYWP46N57WJEFYR/action/replication_record"}},"created_at":"2026-05-17T23:38:53.185247+00:00","updated_at":"2026-05-17T23:38:53.185247+00:00"}