Pith. sign in

Paper Citation Record · LEDGER

Towards a Science of AI Agent Reliability

As of 4 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 27 inbound Pith citation observations for arXiv:2602.16666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.16666 v3

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:32:05.798089Z

measured 127 of 127 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:19:50.579212Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a10e6045-68ae-4bf0-9645-13973e2d02fb · outbound

This paper cites Bc tribunal con- firms companies remain liable for information provided by ai chatbot.Business Law Today,.

Towards a Science of AI Agent Reliability Bc tribunal con- firms companies remain liable for information provided by ai chatbot.Business Law Today,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.000301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.000301Z digest=sha256:18a300fe361f751d24fda4f2b1059501f506675a649e98fe93c1340580c4c59c

Observation a337ddb8-a7aa-45c6-9a97-f3864c0baaa3 · outbound

This paper cites AgentHarm: A benchmark for measuring harmfulness of LLM agents.

Towards a Science of AI Agent Reliability AgentHarm: A benchmark for measuring harmfulness of LLM agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.130639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.130639Z digest=sha256:23f2611a7d2094ad6d63e290281922b6acfda821ab586c0b62c3ea8dffa7d406

Observation 2f324aae-812f-46a7-bb0d-f3e3f30b8413 · outbound

This paper cites AA-Omniscience: Eval- uating cross-domain knowledge reliability in large language models.ArXiv preprint, abs/2511.13029, 2025.

Towards a Science of AI Agent Reliability AA-Omniscience: Eval- uating cross-domain knowledge reliability in large language models.ArXiv preprint, abs/2511.13029, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.226720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.226720Z digest=sha256:5b229b63fb1499f64ead76d17452a63a6d6db970cf77046817082e511ba36d40

Observation 9f4453c2-da10-4019-b2e3-b1826fbaa51e · outbound

This paper cites Basic concepts and taxonomy of dependable and secure computing.IEEE trans- actions on dependable and secure computing, 1 (1):11–33, 2004.

Towards a Science of AI Agent Reliability Basic concepts and taxonomy of dependable and secure computing.IEEE trans- actions on dependable and secure computing, 1 (1):11–33, 2004

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.323137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.323137Z digest=sha256:18007ed468f8dcfcc42951fc5b1451c2d96964d46096f01020af9f9781250000

Observation d2677057-91cb-4b58-aa01-f4ce5a456a69 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Towards a Science of AI Agent Reliability Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.404637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.404637Z digest=sha256:710d0e781e4f827c9972b63a462d75fae2c7978efd365cf16c761694cc231502

Observation 2fa8343e-f3e4-4a87-8d36-62d9eec45b29 · outbound

This paper cites G., Riols, F., and Sharma, R.

Towards a Science of AI Agent Reliability G., Riols, F., and Sharma, R

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.482697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.482697Z digest=sha256:e1ce6743392590d5ab926b63e034857459c874bfcbc41fb74798b88a6cf76686

Observation 1566dc03-ed80-4673-8b6b-3b2329dc6ae6 · outbound

This paper cites John Wiley & Sons, 2015.

Towards a Science of AI Agent Reliability John Wiley & Sons, 2015

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.551202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.551202Z digest=sha256:02073662fd1645f439c6e2ba0feb8da162014f45e46cb05003db1a013cdec0d0

Observation 579165b6-3728-4ad8-a3fc-2fd2e2250822 · outbound

This paper cites Replit ceo apologizes after ai coding tool deleted a live production database.Business Insider, 2025.

Towards a Science of AI Agent Reliability Replit ceo apologizes after ai coding tool deleted a live production database.Business Insider, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.608102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.608102Z digest=sha256:44d0595908dd8d77dbe8f9b5085b287d6082c03f1d99aaaadf4e3a9e468d26bc

Observation 7243b0aa-f28b-4d7b-a257-46da8786b16d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Towards a Science of AI Agent Reliability Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.669112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.669112Z digest=sha256:0ec5c19a2d6062ea26a6726221bee89c62893794044473d773cedf062a7e59ff

Observation 3b50dc7d-c939-42b3-82e2-4668ac9fb89a · outbound

This paper cites Saber: Small actions, big errors–safeguarding mutating steps in llm agents.ArXiv preprint, abs/2512.07850, 2025.

Towards a Science of AI Agent Reliability Saber: Small actions, big errors–safeguarding mutating steps in llm agents.ArXiv preprint, abs/2512.07850, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.720076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.720076Z digest=sha256:d0e503430d3854992e5ba3dbfee09537657dbf8b68ac4999a15b9136aaf05689

Observation 8066a8d9-a153-47ea-a126-205a2b8642a5 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Towards a Science of AI Agent Reliability Mind2web: Towards a generalist agent for the web

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.772845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.772845Z digest=sha256:c8c74b54883b757bbfd871ffc0ff5495edb96228d37674e78fd40bf9d6a3eba1

Observation 1c5d56bc-2808-44ab-b2b5-5d5885ad2ed4 · outbound

This paper cites System safety and artificial intelli- gence.

Towards a Science of AI Agent Reliability System safety and artificial intelli- gence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.818970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.818970Z digest=sha256:b6cf40602367dc617dc4b488bff2f77c1a4405c33cdd9c76f795e00e0d1b62f4

Observation b6e08298-ec2f-4c18-a11a-a5e747c3fe2c · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.862719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.862719Z digest=sha256:ce3af570606b776b02b81c7ea2b07510e8028a5bda02397d401e5eb4fd9104f3

Observation 3e17f9e3-fe8e-416a-b34b-e78e93a0a528 · outbound

This paper cites EN 50126: Railway applications – the specification and demonstration of relia- bility, availability, maintainability and safety (rams), 2017.

Towards a Science of AI Agent Reliability EN 50126: Railway applications – the specification and demonstration of relia- bility, availability, maintainability and safety (rams), 2017

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.912541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.912541Z digest=sha256:a0b6e916a31acea906e84f4e3a14afae29e17c948bf0a6869ac174035626c16f

Observation 37170fcd-5c16-4dc1-baca-7d3df0337495 · outbound

This paper cites AGI safety literature review.

Towards a Science of AI Agent Reliability AGI safety literature review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:56.977985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:56.977985Z digest=sha256:e9f7ca0d2f1812c825f999d97b93bb38557c2071373d0476379421d3a4bc3ff5

Observation 6dce1331-875d-4640-9a66-b6f4becc3009 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630 (8017):625–630, 2024.

Towards a Science of AI Agent Reliability Detecting hallucinations in large language models using semantic entropy.Nature, 630 (8017):625–630, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.280811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.280811Z digest=sha256:c57969cf549d4131dd108c5709ead17762bfdf56e1d01d6be483a4035d88f181

Observation 98b8c2b0-e685-4d21-9334-a0a286a5bd0e · outbound

This paper cites Advisory circular AC 21-16g: RTCA document DO-160 versions D, E, F, and G.

Towards a Science of AI Agent Reliability Advisory circular AC 21-16g: RTCA document DO-160 versions D, E, F, and G

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.322397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.322397Z digest=sha256:f87f5ef7a4eee6cb8f068be728ec4e06c893b97a3e62fb8d07ff2ebf10e5c2af

Observation 45576a26-6f54-4c7b-81a1-bff7ed75b0e6 · outbound

This paper cites Advisory circular AC 25.1309-1b: System design and analysis.

Towards a Science of AI Agent Reliability Advisory circular AC 25.1309-1b: System design and analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.393032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.393032Z digest=sha256:659388ab9e4550a4535b42328bb07b5ac60af0f14ccbe57a317d2c860d833861

Observation d9c59013-f4ba-440d-a4ce-30878cab342a · outbound

This paper cites CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents.

Towards a Science of AI Agent Reliability CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.455347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.455347Z digest=sha256:c10913823619d86e676e9b3ae63b25a854ecb3c7b6be1de16be1d5ab55b4c709

Observation 6daadb69-ab68-456e-81cd-1985825bbbec · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.512065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.512065Z digest=sha256:c245c788b223e01d4a03ad84edc52cb0469cfb900f5fe14a7658f414e906473b

Observation 444fc3fe-29b2-4442-8e28-d279337527f3 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.568257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.568257Z digest=sha256:28522db0bf33199917d072ed803e7b38d3d0a2c2611f32d7aa3cffbe616a6022

Observation 55af2ed3-a1db-4c8b-966f-eb2605f63305 · outbound

This paper cites and Thinking Machines Lab.

Towards a Science of AI Agent Reliability and Thinking Machines Lab

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.627633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.627633Z digest=sha256:052514b8c406a8e6611ffeccbd9b03d6ab9d44b2ce6ddb44f6a4b129ba8fd535

Observation 5d016375-96a5-4957-9791-9ecf17a315bd · outbound

This paper cites Unsolved Problems in ML Safety.

Towards a Science of AI Agent Reliability Unsolved Problems in ML Safety

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.772423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.772423Z digest=sha256:2cd8e899115a1623105a3f855994ee4d11c96e1915a64038b0d5781ea240bc44

Observation 73439282-1249-4eb2-8a3a-3fd523a9ae2c · outbound

This paper cites IEC 61508: Functional safety of electrical/electron- ic/programmable electronic safety-related sys- tems, 2010.

Towards a Science of AI Agent Reliability IEC 61508: Functional safety of electrical/electron- ic/programmable electronic safety-related sys- tems, 2010

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.854855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.854855Z digest=sha256:0b984634234a12396d327a81d677842f1d184b90eb7bd47b9395039f4e0a301c

Observation b629d23e-f387-4c8f-b370-b35965b38885 · outbound

This paper cites IEC 61513: Nuclear power plants – instrumentation and control important to safety – general re- quirements for systems, 2011.

Towards a Science of AI Agent Reliability IEC 61513: Nuclear power plants – instrumentation and control important to safety – general re- quirements for systems, 2011

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.908630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.908630Z digest=sha256:fc35f6fe01c3c2481566605cedad61a24c8f4ade30418c1ed9faefedf4b6aa7f

Observation 426b4eb2-4cba-4a5a-a97f-4a20febed251 · outbound

This paper cites ISO 26262: Road vehicles – functional safety,.

Towards a Science of AI Agent Reliability ISO 26262: Road vehicles – functional safety,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:57.964798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.964798Z digest=sha256:b1aeffe38a457b5009714257d165c19092aae7c586a80cc4659cd379aabce09d

Observation fb1a75a8-1a13-45bd-a96b-f1e173a7d169 · outbound

This paper cites E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K.

Towards a Science of AI Agent Reliability E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.118057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.118057Z digest=sha256:8c763242d9c7ac0cc62b8fb88ad88e80ee1c539140aac70894ec9cd6ed3b750b

Observation adda3f0a-843e-4238-9057-27235285a61d · outbound

This paper cites Language Models (Mostly) Know What They Know.

Towards a Science of AI Agent Reliability Language Models (Mostly) Know What They Know

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.200811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.200811Z digest=sha256:765c9144b54b6c9e348ed5bc0263ccc7af7bdd97ec4b9fa5b7ce21d677ab8007

Observation 6c4ddc9a-3121-455e-8594-eaeed519cf14 · outbound

This paper cites Why Language Models Hallucinate.

Towards a Science of AI Agent Reliability Why Language Models Hallucinate

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.319721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.319721Z digest=sha256:50a76c68bb991edea9045b851db21ce35af287da0b7290b9881571fd58d3e47e

Observation 8b6bc575-95b5-4d97-afc9-829331e1623c · outbound

This paper cites and Garrick, B.

Towards a Science of AI Agent Reliability and Garrick, B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.448052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.448052Z digest=sha256:8fe56e14d1c36eb1f417dfb4f380ae5dec093ae826bd8fe2721c221085303fe7

Observation ac8e115d-b3b6-4ffc-82da-459d8efd795a · outbound

This paper cites S., Wei, B., Xue, T., Chen, Z., Chen, F., Utpala, S., et al.

Towards a Science of AI Agent Reliability S., Wei, B., Xue, T., Chen, Z., Chen, F., Utpala, S., et al

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.558868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.558868Z digest=sha256:df9182b43ae2981ef790e49e1282e41fe44e0c3bbe29aabc6eff487d68470f1d

Observation a14ac10a-1079-4305-b43f-17f6dcbee446 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

Towards a Science of AI Agent Reliability Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.662787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.662787Z digest=sha256:0afcf26d6bf45de868586d3177ad86b40f598c7b7443b286f0399ebb03761656

Observation aa295a54-7671-4e56-b099-74f17f10efd5 · outbound

This paper cites Dependability: Basic concepts and terminology.

Towards a Science of AI Agent Reliability Dependability: Basic concepts and terminology

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.734057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.734057Z digest=sha256:8fb915736863bccd92de7d6a03aa167909d3335569164a9453578579909b02ba

Observation a48b4020-b05f-4202-a175-e3a71ff6583b · outbound

This paper cites NYC’s AI chatbot tells businesses to break the law.The Markup, 2024.

Towards a Science of AI Agent Reliability NYC’s AI chatbot tells businesses to break the law.The Markup, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.793618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.793618Z digest=sha256:b999fe5613e219d51b0948d4ed0a0f6e913aaa0d7c672452b0b3bb045a6618cc

Observation 6245b488-59a3-4a96-93f6-1f7783943d7d · outbound

This paper cites RLAIF vs.

Towards a Science of AI Agent Reliability RLAIF vs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.882469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.882469Z digest=sha256:0ecd8ba404fc9dff4372f76d9f12f07e0a0f8ea4006116118289b63f95f16ce5

Observation 3670f65c-ea99-46dd-920e-db150466cd44 · outbound

This paper cites HaluEval: A large-scale hallucina- tion evaluation benchmark for large language models.

Towards a Science of AI Agent Reliability HaluEval: A large-scale hallucina- tion evaluation benchmark for large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.960411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.960411Z digest=sha256:2fba492b229fa55c44c0f63d34238e2db31847996602829dafe13b94945186eb

Observation 2cc35c64-2c6c-4219-816c-a0a343912e7b · outbound

This paper cites Holistic Evaluation of Language Models.

Towards a Science of AI Agent Reliability Holistic Evaluation of Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.010250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.010250Z digest=sha256:6d99405af22992903edf38d41468765408a4f5f44cff18a35cd38bddb2310c26

Observation 90ce6d96-1e8a-4094-a57a-981774ba6ae6 · outbound

This paper cites Let’s verify step by step.

Towards a Science of AI Agent Reliability Let’s verify step by step

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.059674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.059674Z digest=sha256:7db27b71d0742a180179a2b36aeecff75efaaa50f9b3c1d058a13786aa01d733

Observation b221e500-1ec4-46f9-b5a3-6f17772db80c · outbound

This paper cites Truth- fulQA: Measuring how models mimic human falsehoods.

Towards a Science of AI Agent Reliability Truth- fulQA: Measuring how models mimic human falsehoods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.132441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.132441Z digest=sha256:244b12fec9cd67798a95aa5bf08dfc61013c10e1947caa39e4c9d77cbdf9821a

Observation f7f8f259-157c-4b79-9f1b-9e026aaa1232 · outbound

This paper cites Agentbench: Eval- uating llms as agents.

Towards a Science of AI Agent Reliability Agentbench: Eval- uating llms as agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.209112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.209112Z digest=sha256:0f36f478c86e66e898d0a093506817da53193502191b6da9f1b27ef9d14130d2

Observation ae293490-02db-4f4c-915d-c3ef0e4c18b6 · outbound

This paper cites GAIA: a benchmark for 18 Towards a Science of AI Agent Reliability general AI assistants.

Towards a Science of AI Agent Reliability GAIA: a benchmark for 18 Towards a Science of AI Agent Reliability general AI assistants

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.264997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.264997Z digest=sha256:cfa94216afbfd77e2b22d153110685b3d3fa7a184e1e8b679200f73755db7bcf

Observation 9c1b1fb8-424b-427e-8210-1922fb4d8e44 · outbound

This paper cites Evaluation and benchmarking of llm agents: A survey.

Towards a Science of AI Agent Reliability Evaluation and benchmarking of llm agents: A survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.314863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.314863Z digest=sha256:f9fa2fd39082048dc9e65958ec026efb1b23e4a7f00f09694e47a50163f1140d

Observation 8bee43e2-02e9-419c-85b4-a5278ba6320a · outbound

This paper cites Tech- nical support to the national highway traf- fic safety administration (NHTSA) on the re- ported Toyota motor corporation (TMC) unin- tended acceleration (UA) investigation.

Towards a Science of AI Agent Reliability Tech- nical support to the national highway traf- fic safety administration (NHTSA) on the re- ported Toyota motor corporation (TMC) unin- tended acceleration (UA) investigation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.392872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.392872Z digest=sha256:4df69d03040e43891840d1dad21089062d37735246d9a3b9062e4700a03b1156

Observation 488f3856-a389-4b16-8d60-3890a7cd0261 · outbound

This paper cites The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections.

Towards a Science of AI Agent Reliability The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.395624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.395624Z digest=sha256:7b4b30296a03009bc202fc7a40f3b4432167d987578284fdde8c9bb59dd55498

Observation 84bde15a-dda1-427e-9eb7-3e376d11369b · outbound

This paper cites A taxonomy of trustworthiness for artificial intelligence.CLTC: North Charleston, SC, USA, 1, 2023.

Towards a Science of AI Agent Reliability A taxonomy of trustworthiness for artificial intelligence.CLTC: North Charleston, SC, USA, 1, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.463552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.463552Z digest=sha256:be9ffa7fd36f6ee63d98423f322679d6148f9de13e3bf638a4c3e6bd16bf1988

Observation 4cb71f27-8b65-4562-be1f-6d5a89e810ae · outbound

This paper cites Computer-using agent.

Towards a Science of AI Agent Reliability Computer-using agent

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.611397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.611397Z digest=sha256:703afdaaad24505cd2c197f84abfc432c11f5971975eade25dbbab4ca44e2b3c

Observation e3b5b9fb-58f6-4357-8437-7e118f985c72 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.777746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.777746Z digest=sha256:4f38a7dfea1cba6c3a5fb8206699e2804dc258d7f3636106ab0d9ba4b697c129

Observation 5c0ad10d-346e-44c5-bb2b-44b969729707 · outbound

This paper cites Measuring Agents in Production.

Towards a Science of AI Agent Reliability Measuring Agents in Production

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:59.913036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:59.913036Z digest=sha256:1d9361312148f59c57e07a648964a511b46b060fb4b6e5f17f1967e6fe11403b

Observation b19077b5-8448-4b9e-91ff-b868424346ab · outbound

This paper cites and Papernot, N.

Towards a Science of AI Agent Reliability and Papernot, N

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.012603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.012603Z digest=sha256:6c6e4ee61b3270a1691bd13a05c8339b55b1658cd64fe39091b6f65fc86bf3da

Observation a9d33331-722f-4c34-8078-aa92aaf8aeae · outbound

This paper cites DO-178C: Software considerations in airborne systems and equipment certification, 2012.

Towards a Science of AI Agent Reliability DO-178C: Software considerations in airborne systems and equipment certification, 2012

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.136848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.136848Z digest=sha256:f4098b1e8aa9d9ac0f785024e1676cd1d931187c674e479b7e95ae9c832728fd

Observation 8b5a8fd3-9277-42c2-bc77-d68e147ec667 · outbound

This paper cites Concrete Problems in AI Safety, Revisited.

Towards a Science of AI Agent Reliability Concrete Problems in AI Safety, Revisited

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.250667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.250667Z digest=sha256:c1d44aed9738f1e1a26cd15fdfaf7d281187c120e8bcd618531bea7561975f6d

Observation c935ebc1-8b1a-4e09-aa4d-56e72a1a8ac5 · outbound

This paper cites Bench- marking prompt sensitivity in large language models.

Towards a Science of AI Agent Reliability Bench- marking prompt sensitivity in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.402850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.402850Z digest=sha256:3f142105a4285bd06085cfbaedc85ef9dad5fe8816b2d51454d7c40ef4f2257c

Observation 607a9e46-276e-45aa-b2e2-063e9b32b4b3 · outbound

This paper cites Microsoft puts limits on bing ai chats after its chatbot went off the rails.The New York Times, 2023.

Towards a Science of AI Agent Reliability Microsoft puts limits on bing ai chats after its chatbot went off the rails.The New York Times, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.565941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.565941Z digest=sha256:755e6a38d22945f536b4c928134d5b47a69b02a81790b1fb9f10703569c70eca

Observation f5f5900b-8839-4df1-b7bf-2ceb56c45d2e · outbound

This paper cites SAE arp4761: Guidelines and methods for conducting the safety assess- ment process on civil airborne systems and equipment, 1996.

Towards a Science of AI Agent Reliability SAE arp4761: Guidelines and methods for conducting the safety assess- ment process on civil airborne systems and equipment, 1996

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.680451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.680451Z digest=sha256:ad49bc2f1a3f9c116758541695a134af1cf045f72b96e00fde139aab2906d948

Observation 93cf02e2-c380-411c-8ab2-2598d9c14759 · outbound

This paper cites ARP4754A: Guidelines for development of civil aircraft and systems.

Towards a Science of AI Agent Reliability ARP4754A: Guidelines for development of civil aircraft and systems

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:00.839975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:00.839975Z digest=sha256:76c601caf9ecaafad9bcafdf9f80d036e52b2d1495557f5c28af683a7c40d715

Observation f07eac62-3118-4b97-a797-3afdf6fa50a7 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.006686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.006686Z digest=sha256:dd90935fae394478317bf34b69c547d721410276979c25469b9af6d70717f89d

Observation 495c8ffd-953c-4321-bfa3-c0c8dce8646d · outbound

This paper cites Air canada ordered to pay cus- tomer misled by chatbot.The Guardian, 2024.

Towards a Science of AI Agent Reliability Air canada ordered to pay cus- tomer misled by chatbot.The Guardian, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.131096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.131096Z digest=sha256:7fc1a3fa7a9b747a6bec9d3958e1a1ad864b05eae4e1613a16791156e607f973

Observation 7b0cc11b-d1ae-4722-998a-20c9094e4362 · outbound

This paper cites Ai coding platform goes rogue and deletes entire company 19 Towards a Science of AI Agent Reliability database.Tom’s Hardware, 2025.

Towards a Science of AI Agent Reliability Ai coding platform goes rogue and deletes entire company 19 Towards a Science of AI Agent Reliability database.Tom’s Hardware, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.268184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.268184Z digest=sha256:66cb241e0225256a801e7958f34ed61fd2324f7ce362757de6f505f773c3f659

Observation ccb0a125-5de2-406a-8cf0-0b1e2a1efb8f · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Towards a Science of AI Agent Reliability Solving math word problems with process- and outcome-based feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.393850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.393850Z digest=sha256:55e71ff75681ebec26b4a5acabf97d9262e3ba8a038ef6ef829d48b949d6cc61

Observation d964bfe1-e70a-4d88-bd17-b998dc781a5d · outbound

This paper cites Nuclear Regulatory Commission.

Towards a Science of AI Agent Reliability Nuclear Regulatory Commission

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.465081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.465081Z digest=sha256:43c66283e6b18b9d9ff82b009e5afe93ad337601b325b23bb82019c949452e72

Observation 6b7e4b0c-68b4-4002-9340-d56bf5173cc6 · outbound

This paper cites Nuclear Regulatory Commission.

Towards a Science of AI Agent Reliability Nuclear Regulatory Commission

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.560820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.560820Z digest=sha256:4fe7d6eb108bd7da721d150ab45668d89f1dee623d1ee9152b8cb0f75fd6d30e

Observation 9db4b6cd-1c47-48f3-a99a-e4b623f0844f · outbound

This paper cites Mac-sql: A multi-agent collaborative framework for text-to-sql.

Towards a Science of AI Agent Reliability Mac-sql: A multi-agent collaborative framework for text-to-sql

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.725125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.725125Z digest=sha256:ed98fd13856538057b613481f801b6f1327b9e83ed66b8246d5b607a1282eeb8

Observation 0e2b54e2-1430-445c-971b-624e42866bd5 · outbound

This paper cites Math- shepherd: Verify and reinforce llms step-by-step without human annotations.

Towards a Science of AI Agent Reliability Math- shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:01.877416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:01.877416Z digest=sha256:af5a5f95372d81a45f1fe6a4e8d0f5e55af27af338c32b46df2a3ba66abca7af

Observation 58b21271-5c75-4a4a-8ea2-59011ee69658 · outbound

This paper cites Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks.

Towards a Science of AI Agent Reliability Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.023467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.023467Z digest=sha256:bb23bb283b8f67a61efe53d98d8205d8e22798e21e0af2913b0d2165af917179

Observation ed6e48bb-5ca5-437e-bea2-6117ceec8fb8 · outbound

This paper cites RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models.

Towards a Science of AI Agent Reliability RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.429944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.429944Z digest=sha256:6d7d365ea4edd6c93dcf6a9e1b384fbd172133c286237c200e6cd26dc3ff3ef2

Observation b3377de9-bded-420c-a631-d1d8f75e98f6 · outbound

This paper cites Microsoft limits bing ai chat to 5 replies per session.The Verge, 2023.

Towards a Science of AI Agent Reliability Microsoft limits bing ai chat to 5 replies per session.The Verge, 2023

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.586611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.586611Z digest=sha256:5c80be6117d42f96a04e41cce9b8350938ea8c6b3861558984d98b4e73ec69a9

Observation a7b073b1-64d2-424a-939f-2690c3f7b2ca · outbound

This paper cites Taxonomy of risks posed by language models.

Towards a Science of AI Agent Reliability Taxonomy of risks posed by language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.734774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.734774Z digest=sha256:ca04612849f5a381b00cec3ebfc1c3fec82c5345916d25f3a3c5d1ecdf9de13c

Observation 3fcc2a13-9dd6-4110-9d9a-9259d32ec18b · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

Towards a Science of AI Agent Reliability Sociotechnical Safety Evaluation of Generative AI Systems

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.903561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.903561Z digest=sha256:09f80563ddc29a578895ce6e2fb3c06f8187340cbb35da6877d9f16b879857e3

Observation 2a02438b-b284-494a-b181-9864bdbbebef · outbound

This paper cites E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O.

Towards a Science of AI Agent Reliability E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.190872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.190872Z digest=sha256:47d10a70c3651335cd89a3ab07c91dbddb4953879e37d76b56bcdb19645eadc2

Observation f56cb142-42b9-480b-84e3-eb264157e9a9 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.252437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.252437Z digest=sha256:c025336577d545e3f07246be1f81235af04be5072c6e90652c8a04d473209e5a

Observation 878f42ae-3102-428b-9629-16912559dfec · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Towards a Science of AI Agent Reliability $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.635899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.635899Z digest=sha256:3cb42e8118a5ebb72f5a01486bdc84562c7d6b90e042dde8456243c2dec67675

Observation 5b934862-0017-4b13-8023-63bf05d8b733 · outbound

This paper cites Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J.

Towards a Science of AI Agent Reliability Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.709801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.709801Z digest=sha256:46172b846c0141d472f365b820a076c8600c2a7140729a45eb85fe1e09289397

Observation 0dd3b9a7-4675-4b90-b9f7-d9ecb4c463fb · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

Towards a Science of AI Agent Reliability Calibrate before use: Improving few-shot performance of language models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.811875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.811875Z digest=sha256:50b064e79cbc237c2dfc31f160a8cd518dccf5703011afaa93f4a436de270030

Observation c0f1b8df-14c4-4b74-9c2c-437ce9cc806d · outbound

This paper cites P., Zhang, H., Gonzalez, J.

Towards a Science of AI Agent Reliability P., Zhang, H., Gonzalez, J

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.993956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.993956Z digest=sha256:71bd28a3c0765034ca01dfd4913fbd2e08bb7d93a1b83efb86199ee590031494

Observation 2cc75de2-f29f-4bde-8e87-da062fa84c2e · outbound

This paper cites Larger and more instructable lan- guage models become less reliable.Nature, 634 (8032):61–68, 2024.

Towards a Science of AI Agent Reliability Larger and more instructable lan- guage models become less reliable.Nature, 634 (8032):61–68, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.074734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.074734Z digest=sha256:004ded0974de52b0ff808f21893ad0c2ca75725acfd3413ff21354cac45c6435

Observation 28c0715c-2cb4-41ae-9a49-2c0fc94b6d28 · outbound

This paper cites Can I get a refund for order #12345?.

Towards a Science of AI Agent Reliability Can I get a refund for order #12345?

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.144286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.144286Z digest=sha256:715e06e67ae4b19ce00399f85080df619e6dc3138a74e547d05d01be0eeb6935

Observation 4ca553ad-8a3f-4246-9dc2-6f4b5434afc8 · outbound

This paper cites R., and Cao, Y.

Towards a Science of AI Agent Reliability R., and Cao, Y

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.371122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.371122Z digest=sha256:32fe62444b6358905b3d25809f68c721791dc92b2616a0e9cb39425e64bc3b2e

Observation f2677310-aca0-497d-b02f-b8e896914b7e · outbound

This paper cites 20 Towards a Science of AI Agent Reliability.

Towards a Science of AI Agent Reliability 20 Towards a Science of AI Agent Reliability

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:03.537658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:03.537658Z digest=sha256:a2678e5420ddb349873f195942573f323f95b0acd7ff93334fcbb99e8e3f274d

Observation 95a3d96c-b097-4b1c-a0c4-c5d75c09b0f4 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.212036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.212036Z digest=sha256:60ffe54a6a66f1007df902cf5071fcbabe6ea1498406b7e3454403dea9cbc222

Observation de80676d-64d6-4c50-90d6-a06a01273c45 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.275599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.275599Z digest=sha256:999869954358bb58d6f0bbbbdaf4ee4b50fb905302f7b69b4d1196c1aa8de5a5

Observation 2a381969-814e-4ca4-8043-1e951574f37a · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.354803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.354803Z digest=sha256:c965b584f40a64182c16347919e172406343f8cdc803e97e2b255c77d932eab6

Observation f64de184-07c3-47fd-b425-d162863b4b7b · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.451259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.451259Z digest=sha256:26db95932f89ad83083676f26e2a6567e2b5eaf3fffc1a953eb25b62eadadfc6

Observation 3dca140b-d8b7-4b40-b6ad-ebdf25c8faf3 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.526199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.526199Z digest=sha256:eeadd3deca4e703150dab43fb3ff62697dfc36dd28a681e354d60379ab64d12c

Observation 5699494b-dc09-46eb-b182-af008b176654 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.605601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.605601Z digest=sha256:750a4c89d42add5f63a5593062607fe023e62d661f1195c1361991c6e4d85777

Observation c7c89c01-79c3-4cbb-8cd2-e2701fbb4d6e · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.659249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.659249Z digest=sha256:c384c85960083b939dd8605aee25c66d5fb063ea29f6a6be28874fbb4ee72671

Observation 1c1500e0-1998-444b-9438-5243c9f36fd8 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.727481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.727481Z digest=sha256:5847c1a40716281c6172a5885c9430427f4c606e70d0cac321930c2e7348ccf0

Observation ed380dff-b3ff-4b17-8dff-6e58c6bf1907 · outbound

This paper cites D.2 Implementation We evaluate 14 language models from three major providers, spanning release dates from April 2024 to December 2025.

Towards a Science of AI Agent Reliability D.2 Implementation We evaluate 14 language models from three major providers, spanning release dates from April 2024 to December 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.793653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.793653Z digest=sha256:2e7f99ec2563e2be1fe3e2e45b9cded8b3f1a4152cc7fb8aa8e16ac3d317f739

Observation 622b3c26-bd1c-43f8-9259-36eb185e46a8 · outbound

This paper cites What is the population of Paris in 2024-01-15?.

Towards a Science of AI Agent Reliability What is the population of Paris in 2024-01-15?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.867804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.867804Z digest=sha256:96e00b9d770b84df96ef4a5771307153089c47d0e40be8f24610e794469822f5

Observation 980422d2-d6ee-4b50-9fab-44725929c770 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:04.937875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:04.937875Z digest=sha256:84aa2bc596caec29e93cba56e229ba2c963b782b862430f53c23d0596f619b88

Observation 4d099fe0-c422-438c-a94b-3e1d3ccd80a8 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.025603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.025603Z digest=sha256:0342bd957e3e022803d569e2aa9bb02acffc904f9c90f373e8403b56b00380e5

Observation 9ec7abbf-d5b0-4f20-bc57-a3469b9fb7da · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.093275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.093275Z digest=sha256:ea6a330bcbfeaf494fa3ce5844d4ff2fcb9c9480beb06c1719701d2c13065f64

Observation dacca3d7-7862-4b1f-979c-66693617be12 · outbound

This paper cites Forbidden function.

Towards a Science of AI Agent Reliability Forbidden function

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.175916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.175916Z digest=sha256:1305dab805422089ab77f296118fd6b8fe9837621fdbce6413edbf9886daac35

Observation 6e0fee29-401c-43a7-a09b-f030072166e7 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.275530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.275530Z digest=sha256:0db8ab81f29860ab781a9350ddfdbfc666c50ed545aa7aecc705e74d213cdd02

Observation 871d8092-b558-4a8f-b562-29a7935d8dcd · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.408798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.408798Z digest=sha256:0c87a036e07a189638bd1bb045ceb1389c9c5184a6aa1f4b1104b0afbbd9fa3f

Observation bbe7a93c-8d86-4d2c-9329-bfbf3afa9bae · outbound

This paper cites informational.

Towards a Science of AI Agent Reliability informational

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.513237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.513237Z digest=sha256:6c01c49b0abdf58bdec1e6a25bfcb624e960ac5085b85a8a0350058bd1ff4ac6

Observation b7020922-4717-4160-8eec-e856d654f433 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.601729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.601729Z digest=sha256:a328b06de52ac409165c89d98de1a22468a530b27df65aba7810222f4125d348

Observation 7be98929-8f5c-4bfe-b59c-1a4548490c1f · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.688987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.688987Z digest=sha256:d24a5f57940f534fc85713226bd3a0f847eb9e2f2cd1c11039ec3b16e7a228d4

Observation bec867ef-b75c-4ac8-b520-af628a832f9f · outbound

This paper cites try harder.

Towards a Science of AI Agent Reliability try harder

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:05.798089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:05.798089Z digest=sha256:c467a30e378b20a1fdc7b50614a3749666a4630c788109610745186674d1d5ad

Observation a8ff1e6e-ce60-456d-970c-722e1695f0b2 · outbound

This paper cites 2018/768.

Towards a Science of AI Agent Reliability 2018/768

Reference 768

Resolution
malformed identifier
no resolver link, observed 2026-08-02T22:31:57.211940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:57.211940Z digest=sha256:6a482c381b944fa05ab24a390f3ee7cf7a37d12386be99d4be6019a36d68ab2d

Observation da8f3795-777a-44af-9eac-8a3a4eca80c5 · outbound

This paper cites an unresolved cited work.

Towards a Science of AI Agent Reliability Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:58.036119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:58.036119Z digest=sha256:cb7677f14f4ca40b50d6528fb37bb30835c2b45a60d4d89cf6d14e78776f8ab9

Pith citing papers

Observation 6ed70e2c-2a10-4796-bf45-1aefcfc0e33d · inbound

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation cites this paper.

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation Towards a Science of AI Agent Reliability

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:48:11.250689Z digest=sha256:096e47f10d05ab00ca466744d08381f22efaeb0fe4ad9399ecd37827e2c6816c

Observation 0a2bdf16-e43e-41f4-9865-d911d5e96a1f · inbound

MarketBench: Evaluating AI Agents as Market Participants cites this paper.

MarketBench: Evaluating AI Agents as Market Participants Towards a Science of AI Agent Reliability

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T06:05:07.499390Z digest=sha256:7cbb2ae5ffdae988433b99f40a663ce7e67f75e3f571b725f0a970fc353f2cae

Observation 9889ff19-7931-4bf1-9783-b3025e31eb08 · inbound

Hallucinations Undermine Trust; Metacognition is a Way Forward cites this paper.

Hallucinations Undermine Trust; Metacognition is a Way Forward Towards a Science of AI Agent Reliability

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T14:29:54.924293Z digest=sha256:e85e4bd18ff0eb04bd74d1fdbb3ce9344259552d426fdd01785c4934476521ce

Observation c9384694-029a-4423-b4c9-cf36be9197cc · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Towards a Science of AI Agent Reliability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:54ae4f9f74ddffc6eaeca2896738bea4c49ed71f6857b35cebbb33624c146372

Observation a08f4184-c304-4a8d-bea9-b0e684746564 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Towards a Science of AI Agent Reliability

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:58b84b190069453b4fd8f7a5fc2ce2b91e0dceccdd3c41d2f0299d7a4ee5182d

Observation ead9c734-8157-4dc3-97af-de282ca33a1d · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Towards a Science of AI Agent Reliability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:50.579212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:50.579212Z digest=sha256:010e41f0bbf776f4545813ec291e3e2daa516a92348499628ff19d6f5109a8ee

Observation f85ca813-783e-4d8c-9295-bf08cf54d27f · inbound

Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes cites this paper.

Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes Towards a Science of AI Agent Reliability

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T04:58:02.620334Z digest=sha256:77b325da410ca88c22ceed80569fe6abecf0e0d64cae1e2514359a0991af8b38

Observation 7bb037b1-d4a2-4479-9f34-f0ab9238a985 · inbound

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents cites this paper.

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents Towards a Science of AI Agent Reliability

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:29:17.065496Z digest=sha256:019c1dda41e24cdc64b238b2b97390661173b32be5610471d5dc95825e91ef1e

Observation f5569acb-0f07-4d96-81bf-aea1fb9b8e79 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities Towards a Science of AI Agent Reliability

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:282949007d5812ce6b4fafc02d5c0bf43aa66b6ff7a76ab9338c2c400825b1e1

Observation d8433886-6462-4a12-a5e7-8d1c216ffe06 · inbound

PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents cites this paper.

PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents Towards a Science of AI Agent Reliability

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T09:19:26.725722Z digest=sha256:44bb80ae05edf0d7a62b131ca8338281296b859a43cde1f9db57ffccc78be4b7

Observation de17cb54-7b6e-469d-85e6-fb7155f42c92 · inbound

Security, Privacy, and Ethical Risks in OpenClaw cites this paper.

Security, Privacy, and Ethical Risks in OpenClaw Towards a Science of AI Agent Reliability

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-03T02:05:13.650582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T04:25:49.102385Z digest=sha256:468729ca1b221dd85ebc3cfcfb8e3952df0d73a0231fc2470400850de9547d73

Observation bffe2de8-0724-4fb9-a889-fa4f1c7dcafc · inbound

Monitoring Agentic Systems Before They're Reliable cites this paper.

Monitoring Agentic Systems Before They're Reliable Towards a Science of AI Agent Reliability

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:16:25.063284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:28:46.165040Z digest=sha256:54f2c8bbf669cf0c4b438b14ba84e3011c9e7b916f44fc9b2255cfbdd312cb38

Observation 55698cfa-3644-4578-8c9c-f76697b1024b · inbound

The Agentic Web Requires New Normative Infrastructure cites this paper.

The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-06-28T02:41:32.588820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T11:55:10.777827Z digest=sha256:7304d931533d480cac3590b11591c964b71151e922e9bce18f7339523f5abeda

Observation 48fdb14d-6f07-44a3-b77e-7bb71b09e991 · inbound

The Agentic Web Requires New Normative Infrastructure cites this paper.

The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability

Reference 292

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T07:37:45.656513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T11:55:10.777827Z digest=sha256:aba3e2eab0fb1e82c2e885c645fe7cfabfe9ff3d6e169e6c78cfe56dd1c89f4f

Observation f6c2acb3-5074-466a-9337-56e16cc5eb01 · inbound

The Agentic Web Requires New Normative Infrastructure cites this paper.

The Agentic Web Requires New Normative Infrastructure Towards a Science of AI Agent Reliability

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:03.111695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:03.111695Z digest=sha256:a73b51d0f7ecec701522f2fcc4a72e37defa743d0f4bf74f2dd23b30737f92f7

Observation fcb7779e-f229-4214-aeb0-628d711248b3 · inbound

On the Reliability of Networks of AI Agents: Density Evolution, Stopping Sets, and Architecture Optimization cites this paper.

On the Reliability of Networks of AI Agents: Density Evolution, Stopping Sets, and Architecture Optimization Towards a Science of AI Agent Reliability

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:49:02.316356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:47:40.583759Z digest=sha256:03400d30e396d45b11540f4f5e9cb7a74fbe28dab544efe6189d2aa8bba08b30

Observation 87e4f5f4-b3a6-481d-883b-ec359783bf39 · inbound

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures cites this paper.

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures Towards a Science of AI Agent Reliability

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:38:44.290825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T03:55:09.415657Z digest=sha256:7833c20f61f184f750bf09ada194b969962de19afbac98e5ea6044f9b461449b

Observation 31c8b8ae-ac70-421a-af47-9895648093ad · inbound

How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks cites this paper.

How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks Towards a Science of AI Agent Reliability

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:28:49.823913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T02:49:04.341583Z digest=sha256:361c35893e3dc2b689bd0dee79faf5f491a5b3f7bc5538716c982435787a5b26

Observation 28fdfbb4-b973-44a3-ac86-c5c0f797105d · inbound

Life After Benchmark Saturation: A Case Study of CORE-Bench cites this paper.

Life After Benchmark Saturation: A Case Study of CORE-Bench Towards a Science of AI Agent Reliability

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:57.809667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:17:50.141575Z digest=sha256:23abf1d1c01d145997ad075a0f92cca52c3001152e53d8a7b9938019bd1457d0

Observation 9f275142-5b90-464c-a3e9-2e542f4f5354 · inbound

A Scalable Approach to Evaluating Moral Sensitivity in LLMs cites this paper.

A Scalable Approach to Evaluating Moral Sensitivity in LLMs Towards a Science of AI Agent Reliability

Reference 140

Resolution
unresolved
no resolver link, observed 2026-07-12T05:44:33.099337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:44:33.099337Z digest=sha256:1c61e1c823ba80eb1ab77857ce7531bd378a9e062f9c1d2e2f0172297c5d318b

Observation 33f62d4c-08b2-4a26-a187-116324a42ff5 · inbound

The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling cites this paper.

The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling Towards a Science of AI Agent Reliability

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:22.729188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:22.729188Z digest=sha256:3b57cddf6eb277f3349b127528b7b1b453693f7f7134c6369aff3c028073ff2b

Observation dce57456-45d2-4361-af9e-fe7484d4f62d · inbound

The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies cites this paper.

The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies Towards a Science of AI Agent Reliability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T14:32:04.742971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:32:04.742971Z digest=sha256:38c1e0bed7240f10b8045d1ed39d7c30aa192c35ce5b4e19903b1abc815baad3

Observation 29cf422f-1591-4cb4-95a7-8942a256def4 · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Towards a Science of AI Agent Reliability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T08:35:47.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T08:35:47.870083Z digest=sha256:be81b0d933d9e92ef0f869468aa290cfe571dc42cdfb5f993c0abd5168bb5763

Observation 34b5ddd2-4684-41a5-aeac-f399b62640be · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Towards a Science of AI Agent Reliability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T06:46:18.294647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:46:18.294647Z digest=sha256:88b80f2bf14a624c5851c779c984a7b2081704e9d92b08ab05efdc3d8a97b0c7

Observation 8469b6ee-b72a-49f6-b0a0-57be3d51ad09 · inbound

Decision Making Needs Uncertainty Quantification [Lecture Notes] cites this paper.

Decision Making Needs Uncertainty Quantification [Lecture Notes] Towards a Science of AI Agent Reliability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T02:14:46.632901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:14:46.632901Z digest=sha256:16423f183941bd13f9236ab33e39225521da419e27e53fd2664c165097c10572

Observation 939ae7d6-8ba3-42d5-8c82-3a1ac016493b · inbound

Plover: Steering GUI Agents through Plan-Centric Interaction cites this paper.

Plover: Steering GUI Agents through Plan-Centric Interaction Towards a Science of AI Agent Reliability

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T23:57:11.469991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:57:11.469991Z digest=sha256:686d08f207d0108b38a912c72a7db06b16ac0def973b74617bca8728467f8fb0

Observation e965362c-196f-455a-90fa-d49366434e4f · inbound

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance cites this paper.

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance Towards a Science of AI Agent Reliability

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T21:18:26.994142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:18:26.994142Z digest=sha256:21d7018c4e5bc01037d69ab8b33d9be8c7f8f601a3fc2137f3b0d84507c70d29