Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:32:05.798089Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 27 inbound Pith citation observations for arXiv:2602.16666.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:32:05.798089Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:19:50.579212Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a10e6045-68ae-4bf0-9645-13973e2d02fb · outbound
Towards a Science of AI Agent Reliability Bc tribunal con- firms companies remain liable for information provided by ai chatbot.Business Law Today,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a337ddb8-a7aa-45c6-9a97-f3864c0baaa3 · outbound
Towards a Science of AI Agent Reliability AgentHarm: A benchmark for measuring harmfulness of LLM agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f324aae-812f-46a7-bb0d-f3e3f30b8413 · outbound
Towards a Science of AI Agent Reliability AA-Omniscience: Eval- uating cross-domain knowledge reliability in large language models.ArXiv preprint, abs/2511.13029, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4453c2-da10-4019-b2e3-b1826fbaa51e · outbound
Towards a Science of AI Agent Reliability Basic concepts and taxonomy of dependable and secure computing.IEEE trans- actions on dependable and secure computing, 1 (1):11–33, 2004
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2677057-91cb-4b58-aa01-f4ce5a456a69 · outbound
Towards a Science of AI Agent Reliability Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa8343e-f3e4-4a87-8d36-62d9eec45b29 · outbound
Towards a Science of AI Agent Reliability G., Riols, F., and Sharma, R
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1566dc03-ed80-4673-8b6b-3b2329dc6ae6 · outbound
Towards a Science of AI Agent Reliability John Wiley & Sons, 2015
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579165b6-3728-4ad8-a3fc-2fd2e2250822 · outbound
Towards a Science of AI Agent Reliability Replit ceo apologizes after ai coding tool deleted a live production database.Business Insider, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7243b0aa-f28b-4d7b-a257-46da8786b16d · outbound
Towards a Science of AI Agent Reliability Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b50dc7d-c939-42b3-82e2-4668ac9fb89a · outbound
Towards a Science of AI Agent Reliability Saber: Small actions, big errors–safeguarding mutating steps in llm agents.ArXiv preprint, abs/2512.07850, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8066a8d9-a153-47ea-a126-205a2b8642a5 · outbound
Towards a Science of AI Agent Reliability Mind2web: Towards a generalist agent for the web
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5d56bc-2808-44ab-b2b5-5d5885ad2ed4 · outbound
Towards a Science of AI Agent Reliability System safety and artificial intelli- gence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e08298-ec2f-4c18-a11a-a5e747c3fe2c · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e17f9e3-fe8e-416a-b34b-e78e93a0a528 · outbound
Towards a Science of AI Agent Reliability EN 50126: Railway applications – the specification and demonstration of relia- bility, availability, maintainability and safety (rams), 2017
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37170fcd-5c16-4dc1-baca-7d3df0337495 · outbound
Towards a Science of AI Agent Reliability AGI safety literature review
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dce1331-875d-4640-9a66-b6f4becc3009 · outbound
Towards a Science of AI Agent Reliability Detecting hallucinations in large language models using semantic entropy.Nature, 630 (8017):625–630, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b8c2b0-e685-4d21-9334-a0a286a5bd0e · outbound
Towards a Science of AI Agent Reliability Advisory circular AC 21-16g: RTCA document DO-160 versions D, E, F, and G
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45576a26-6f54-4c7b-81a1-bff7ed75b0e6 · outbound
Towards a Science of AI Agent Reliability Advisory circular AC 25.1309-1b: System design and analysis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c59013-f4ba-440d-a4ce-30878cab342a · outbound
Towards a Science of AI Agent Reliability CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6daadb69-ab68-456e-81cd-1985825bbbec · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444fc3fe-29b2-4442-8e28-d279337527f3 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55af2ed3-a1db-4c8b-966f-eb2605f63305 · outbound
Towards a Science of AI Agent Reliability and Thinking Machines Lab
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d016375-96a5-4957-9791-9ecf17a315bd · outbound
Towards a Science of AI Agent Reliability Unsolved Problems in ML Safety
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73439282-1249-4eb2-8a3a-3fd523a9ae2c · outbound
Towards a Science of AI Agent Reliability IEC 61508: Functional safety of electrical/electron- ic/programmable electronic safety-related sys- tems, 2010
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b629d23e-f387-4c8f-b370-b35965b38885 · outbound
Towards a Science of AI Agent Reliability IEC 61513: Nuclear power plants – instrumentation and control important to safety – general re- quirements for systems, 2011
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426b4eb2-4cba-4a5a-a97f-4a20febed251 · outbound
Towards a Science of AI Agent Reliability ISO 26262: Road vehicles – functional safety,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1a75a8-1a13-45bd-a96b-f1e173a7d169 · outbound
Towards a Science of AI Agent Reliability E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adda3f0a-843e-4238-9057-27235285a61d · outbound
Towards a Science of AI Agent Reliability Language Models (Mostly) Know What They Know
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4ddc9a-3121-455e-8594-eaeed519cf14 · outbound
Towards a Science of AI Agent Reliability Why Language Models Hallucinate
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6bc575-95b5-4d97-afc9-829331e1623c · outbound
Towards a Science of AI Agent Reliability and Garrick, B
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8e115d-b3b6-4ffc-82da-459d8efd795a · outbound
Towards a Science of AI Agent Reliability S., Wei, B., Xue, T., Chen, Z., Chen, F., Utpala, S., et al
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14ac10a-1079-4305-b43f-17f6dcbee446 · outbound
Towards a Science of AI Agent Reliability Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa295a54-7671-4e56-b099-74f17f10efd5 · outbound
Towards a Science of AI Agent Reliability Dependability: Basic concepts and terminology
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48b4020-b05f-4202-a175-e3a71ff6583b · outbound
Towards a Science of AI Agent Reliability NYC’s AI chatbot tells businesses to break the law.The Markup, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6245b488-59a3-4a96-93f6-1f7783943d7d · outbound
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3670f65c-ea99-46dd-920e-db150466cd44 · outbound
Towards a Science of AI Agent Reliability HaluEval: A large-scale hallucina- tion evaluation benchmark for large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc35c64-2c6c-4219-816c-a0a343912e7b · outbound
Towards a Science of AI Agent Reliability Holistic Evaluation of Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ce6d96-1e8a-4094-a57a-981774ba6ae6 · outbound
Towards a Science of AI Agent Reliability Let’s verify step by step
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b221e500-1ec4-46f9-b5a3-6f17772db80c · outbound
Towards a Science of AI Agent Reliability Truth- fulQA: Measuring how models mimic human falsehoods
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f8f259-157c-4b79-9f1b-9e026aaa1232 · outbound
Towards a Science of AI Agent Reliability Agentbench: Eval- uating llms as agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae293490-02db-4f4c-915d-c3ef0e4c18b6 · outbound
Towards a Science of AI Agent Reliability GAIA: a benchmark for 18 Towards a Science of AI Agent Reliability general AI assistants
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1b1fb8-424b-427e-8210-1922fb4d8e44 · outbound
Towards a Science of AI Agent Reliability Evaluation and benchmarking of llm agents: A survey
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bee43e2-02e9-419c-85b4-a5278ba6320a · outbound
Towards a Science of AI Agent Reliability Tech- nical support to the national highway traf- fic safety administration (NHTSA) on the re- ported Toyota motor corporation (TMC) unin- tended acceleration (UA) investigation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 488f3856-a389-4b16-8d60-3890a7cd0261 · outbound
Towards a Science of AI Agent Reliability The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84bde15a-dda1-427e-9eb7-3e376d11369b · outbound
Towards a Science of AI Agent Reliability A taxonomy of trustworthiness for artificial intelligence.CLTC: North Charleston, SC, USA, 1, 2023
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb71f27-8b65-4562-be1f-6d5a89e810ae · outbound
Towards a Science of AI Agent Reliability Computer-using agent
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3b5b9fb-58f6-4357-8437-7e118f985c72 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0ad10d-346e-44c5-bb2b-44b969729707 · outbound
Towards a Science of AI Agent Reliability Measuring Agents in Production
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19077b5-8448-4b9e-91ff-b868424346ab · outbound
Towards a Science of AI Agent Reliability and Papernot, N
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d33331-722f-4c34-8078-aa92aaf8aeae · outbound
Towards a Science of AI Agent Reliability DO-178C: Software considerations in airborne systems and equipment certification, 2012
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5a8fd3-9277-42c2-bc77-d68e147ec667 · outbound
Towards a Science of AI Agent Reliability Concrete Problems in AI Safety, Revisited
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c935ebc1-8b1a-4e09-aa4d-56e72a1a8ac5 · outbound
Towards a Science of AI Agent Reliability Bench- marking prompt sensitivity in large language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607a9e46-276e-45aa-b2e2-063e9b32b4b3 · outbound
Towards a Science of AI Agent Reliability Microsoft puts limits on bing ai chats after its chatbot went off the rails.The New York Times, 2023
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5f5900b-8839-4df1-b7bf-2ceb56c45d2e · outbound
Towards a Science of AI Agent Reliability SAE arp4761: Guidelines and methods for conducting the safety assess- ment process on civil airborne systems and equipment, 1996
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93cf02e2-c380-411c-8ab2-2598d9c14759 · outbound
Towards a Science of AI Agent Reliability ARP4754A: Guidelines for development of civil aircraft and systems
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07eac62-3118-4b97-a797-3afdf6fa50a7 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495c8ffd-953c-4321-bfa3-c0c8dce8646d · outbound
Towards a Science of AI Agent Reliability Air canada ordered to pay cus- tomer misled by chatbot.The Guardian, 2024
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0cc11b-d1ae-4722-998a-20c9094e4362 · outbound
Towards a Science of AI Agent Reliability Ai coding platform goes rogue and deletes entire company 19 Towards a Science of AI Agent Reliability database.Tom’s Hardware, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb0a125-5de2-406a-8cf0-0b1e2a1efb8f · outbound
Towards a Science of AI Agent Reliability Solving math word problems with process- and outcome-based feedback
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d964bfe1-e70a-4d88-bd17-b998dc781a5d · outbound
Towards a Science of AI Agent Reliability Nuclear Regulatory Commission
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7e4b0c-68b4-4002-9340-d56bf5173cc6 · outbound
Towards a Science of AI Agent Reliability Nuclear Regulatory Commission
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db4b6cd-1c47-48f3-a99a-e4b623f0844f · outbound
Towards a Science of AI Agent Reliability Mac-sql: A multi-agent collaborative framework for text-to-sql
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2b54e2-1430-445c-971b-624e42866bd5 · outbound
Towards a Science of AI Agent Reliability Math- shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b21271-5c75-4a4a-8ea2-59011ee69658 · outbound
Towards a Science of AI Agent Reliability Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6e48bb-5ca5-437e-bea2-6117ceec8fb8 · outbound
Towards a Science of AI Agent Reliability RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3377de9-bded-420c-a631-d1d8f75e98f6 · outbound
Towards a Science of AI Agent Reliability Microsoft limits bing ai chat to 5 replies per session.The Verge, 2023
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b073b1-64d2-424a-939f-2690c3f7b2ca · outbound
Towards a Science of AI Agent Reliability Taxonomy of risks posed by language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fcc2a13-9dd6-4110-9d9a-9259d32ec18b · outbound
Towards a Science of AI Agent Reliability Sociotechnical Safety Evaluation of Generative AI Systems
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a02438b-b284-494a-b181-9864bdbbebef · outbound
Towards a Science of AI Agent Reliability E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56cb142-42b9-480b-84e3-eb264157e9a9 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878f42ae-3102-428b-9629-16912559dfec · outbound
Towards a Science of AI Agent Reliability $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b934862-0017-4b13-8023-63bf05d8b733 · outbound
Towards a Science of AI Agent Reliability Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd3b9a7-4675-4b90-b9f7-d9ecb4c463fb · outbound
Towards a Science of AI Agent Reliability Calibrate before use: Improving few-shot performance of language models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f1b8df-14c4-4b74-9c2c-437ce9cc806d · outbound
Towards a Science of AI Agent Reliability P., Zhang, H., Gonzalez, J
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc75de2-f29f-4bde-8e87-da062fa84c2e · outbound
Towards a Science of AI Agent Reliability Larger and more instructable lan- guage models become less reliable.Nature, 634 (8032):61–68, 2024
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c0715c-2cb4-41ae-9a49-2c0fc94b6d28 · outbound
Towards a Science of AI Agent Reliability Can I get a refund for order #12345?
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca553ad-8a3f-4246-9dc2-6f4b5434afc8 · outbound
Towards a Science of AI Agent Reliability R., and Cao, Y
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2677310-aca0-497d-b02f-b8e896914b7e · outbound
Towards a Science of AI Agent Reliability 20 Towards a Science of AI Agent Reliability
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a3d96c-b097-4b1c-a0c4-c5d75c09b0f4 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de80676d-64d6-4c50-90d6-a06a01273c45 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a381969-814e-4ca4-8043-1e951574f37a · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64de184-07c3-47fd-b425-d162863b4b7b · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dca140b-d8b7-4b40-b6ad-ebdf25c8faf3 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5699494b-dc09-46eb-b182-af008b176654 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c89c01-79c3-4cbb-8cd2-e2701fbb4d6e · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1500e0-1998-444b-9438-5243c9f36fd8 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed380dff-b3ff-4b17-8dff-6e58c6bf1907 · outbound
Towards a Science of AI Agent Reliability D.2 Implementation We evaluate 14 language models from three major providers, spanning release dates from April 2024 to December 2025
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622b3c26-bd1c-43f8-9259-36eb185e46a8 · outbound
Towards a Science of AI Agent Reliability What is the population of Paris in 2024-01-15?
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980422d2-d6ee-4b50-9fab-44725929c770 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d099fe0-c422-438c-a94b-3e1d3ccd80a8 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec7abbf-d5b0-4f20-bc57-a3469b9fb7da · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dacca3d7-7862-4b1f-979c-66693617be12 · outbound
Towards a Science of AI Agent Reliability Forbidden function
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0fee29-401c-43a7-a09b-f030072166e7 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871d8092-b558-4a8f-b562-29a7935d8dcd · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe7a93c-8d86-4d2c-9329-bfbf3afa9bae · outbound
Towards a Science of AI Agent Reliability informational
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7020922-4717-4160-8eec-e856d654f433 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be98929-8f5c-4bfe-b59c-1a4548490c1f · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec867ef-b75c-4ac8-b520-af628a832f9f · outbound
Towards a Science of AI Agent Reliability try harder
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ff1e6e-ce60-456d-970c-722e1695f0b2 · outbound
Reference 768
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8f3795-777a-44af-9eac-8a3a4eca80c5 · outbound
Towards a Science of AI Agent Reliability Unresolved cited work
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed70e2c-2a10-4796-bf45-1aefcfc0e33d · inbound
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation Towards a Science of AI Agent Reliability
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a2bdf16-e43e-41f4-9865-d911d5e96a1f · inbound
MarketBench: Evaluating AI Agents as Market Participants Towards a Science of AI Agent Reliability
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9889ff19-7931-4bf1-9783-b3025e31eb08 · inbound
Hallucinations Undermine Trust; Metacognition is a Way Forward Towards a Science of AI Agent Reliability
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9384694-029a-4423-b4c9-cf36be9197cc · inbound
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Towards a Science of AI Agent Reliability
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a08f4184-c304-4a8d-bea9-b0e684746564 · inbound
Nautilus: From One Prompt to Plug-and-Play Robot Learning Towards a Science of AI Agent Reliability
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ead9c734-8157-4dc3-97af-de282ca33a1d · inbound
Nautilus: From One Prompt to Plug-and-Play Robot Learning Towards a Science of AI Agent Reliability
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f85ca813-783e-4d8c-9295-bf08cf54d27f · inbound
Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes Towards a Science of AI Agent Reliability
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7bb037b1-d4a2-4479-9f34-f0ab9238a985 · inbound
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents Towards a Science of AI Agent Reliability
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5569acb-0f07-4d96-81bf-aea1fb9b8e79 · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities Towards a Science of AI Agent Reliability
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8433886-6462-4a12-a5e7-8d1c216ffe06 · inbound
PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents Towards a Science of AI Agent Reliability
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de17cb54-7b6e-469d-85e6-fb7155f42c92 · inbound
Security, Privacy, and Ethical Risks in OpenClaw Towards a Science of AI Agent Reliability
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bffe2de8-0724-4fb9-a889-fa4f1c7dcafc · inbound
Monitoring Agentic Systems Before They're Reliable Towards a Science of AI Agent Reliability
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 55698cfa-3644-4578-8c9c-f76697b1024b · inbound
The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48fdb14d-6f07-44a3-b77e-7bb71b09e991 · inbound
The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability
Reference 292
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6c2acb3-5074-466a-9337-56e16cc5eb01 · inbound
The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb7779e-f229-4214-aeb0-628d711248b3 · inbound
On the Reliability of Networks of AI Agents: Density Evolution, Stopping Sets, and Architecture Optimization Towards a Science of AI Agent Reliability
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87e4f5f4-b3a6-481d-883b-ec359783bf39 · inbound
ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures Towards a Science of AI Agent Reliability
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31c8b8ae-ac70-421a-af47-9895648093ad · inbound
How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks Towards a Science of AI Agent Reliability
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28fdfbb4-b973-44a3-ac86-c5c0f797105d · inbound
Life After Benchmark Saturation: A Case Study of CORE-Bench Towards a Science of AI Agent Reliability
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f275142-5b90-464c-a3e9-2e542f4f5354 · inbound
A Scalable Approach to Evaluating Moral Sensitivity in LLMs Towards a Science of AI Agent Reliability
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f62d4c-08b2-4a26-a187-116324a42ff5 · inbound
The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling Towards a Science of AI Agent Reliability
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dce57456-45d2-4361-af9e-fe7484d4f62d · inbound
The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies Towards a Science of AI Agent Reliability
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cf422f-1591-4cb4-95a7-8942a256def4 · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Towards a Science of AI Agent Reliability
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b5ddd2-4684-41a5-aeac-f399b62640be · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Towards a Science of AI Agent Reliability
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8469b6ee-b72a-49f6-b0a0-57be3d51ad09 · inbound
Decision Making Needs Uncertainty Quantification [Lecture Notes] Towards a Science of AI Agent Reliability
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939ae7d6-8ba3-42d5-8c82-3a1ac016493b · inbound
Plover: Steering GUI Agents through Plan-Centric Interaction Towards a Science of AI Agent Reliability
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e965362c-196f-455a-90fa-d49366434e4f · inbound
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance Towards a Science of AI Agent Reliability
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.