Pith. sign in

Paper Citation Record · LEDGER

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents

As of 23 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2605.24134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24134 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T14:38:43.708578Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:38:46.490717Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact7
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bb3c0a5c-6621-4924-bf36-fc53ab83b744 · outbound

This paper cites Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.266937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:e4c464a186b1f239159320bb5e04872be04a4d89a71a7690589cc8027811f20d

Observation 50dabbfc-72e2-48ed-863d-394171b495fc · outbound

This paper cites Physical AI Agents: Integrating Cognitive Intelligence with Real-World Action.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Physical AI Agents: Integrating Cognitive Intelligence with Real-World Action

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.278771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:fd84ab50ed5b75e3ceeacce6180fe454f692279d95080e7741c0bae39dc43643

Observation 01783b3b-c78f-43ca-b62e-d575367fa748 · outbound

This paper cites Ai agents need memory control over more context.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Ai agents need memory control over more context

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.285472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:5db6acde0c6ade2ef43a86a664bc0416588fddf582cb88dabf2518a6e6afcf64

Observation a6ac0793-c2ec-4c26-b002-b85dc545e483 · outbound

This paper cites Human oversight in the eu artificial intelligence act.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Human oversight in the eu artificial intelligence act

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.976106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:dd97989736eee652e0c3a6407e0d81ddc08a1edb3bd3a8157b25e889524d38ef

Observation 328a27f3-e45e-4e13-8972-aa4ee9a058a3 · outbound

This paper cites Regulation (eu) 2024/1689: Artificial intelligence act, article 14 human oversight,.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Regulation (eu) 2024/1689: Artificial intelligence act, article 14 human oversight,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.961252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:afd3c7dd61b7511c4961755697c6bb6b3777188d360297cac7c08d1ccb234f20

Observation c7a16a38-f02a-40c9-9661-d78c4635c1b8 · outbound

This paper cites an unresolved cited work.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:25:39.963678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:315df2ae965c37b33ec70e5e076ff305c4626ba756a3a93bda1a2c48e140a24f

Observation f8543b32-dc10-4cb0-968f-c667e5f46b08 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.287996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:f9b44eba89f63673ca17c31cd7e1a7e57a8a96daba2b9295556b51ff26e8c40b

Observation 8ea776ba-8b8c-44d6-977e-9728b13b432b · outbound

This paper cites Ethics guidelines for trustworthy ai, 2019.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Ethics guidelines for trustworthy ai, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.966339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:8ec11029f54c675c376ae6f7ad547828dcd774d0fc21da8cbde66e1b05bd6396

Observation 7b62f927-09a9-4dcb-8592-3f4af9e2edc2 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.956195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:91a77e0dec019a50a0789d0bd6996dc8789fa1a90efc22c16bb4ecf892a58177

Observation 03253d2e-b22e-47e7-b2fa-dfaf55264a55 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.281931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:0d4085986b3b548d843bc7ada576e9e29c78ec1855d2993448776c0930bd170a

Observation e122c7e9-9e48-4c2d-b6ef-7e03a0502408 · outbound

This paper cites G-eval: Nlg evaluation using gpt-4 with better human alignment.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents G-eval: Nlg evaluation using gpt-4 with better human alignment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.978219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:029a3807edbd049214007af4654f543919de8bce2fdc4a3a06becb4bb9166d7b

Observation ac35bca0-09d9-403c-895f-dc402dbd2fe4 · outbound

This paper cites Red teaming language models with language models.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Red teaming language models with language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.968973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:18133a33788dfdb93c1c3d1de44975d1cf709d0f98fd6fdd6fc5e3d1e5c756f9

Observation 149bae42-7880-426f-9031-84d3040bbd04 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:545519d000d07806158d7c53ab1a7dbaa5cea9ef099d000d831c6679392b4f40

Observation 990183fa-e47a-42b5-9f70-e8243d13016a · outbound

This paper cites Jailbroken: How does llm safety training fail?Advances in Neural Information Processing Systems, 2023.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Jailbroken: How does llm safety training fail?Advances in Neural Information Processing Systems, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.953407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:922a1ee0079ad7776e527e4ede33951ad821a3feea836a6dcfc1fa3a2ed63a56

Observation 1c2610c4-fae3-4db9-af7b-6f477a3a11e4 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.272643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:53fec01c38a3b097a4c3831ac820667e0abad153f884909aba134ea3bb379a0b

Observation 6ba8dd86-e0ff-44ee-ae31-9ba0b34c9988 · outbound

This paper cites Web- shop: Towards scalable real-world web interaction with grounded lan- guage agents.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Web- shop: Towards scalable real-world web interaction with grounded lan- guage agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.960419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:1f4b9dea323db0540a2fa232e855033b90db06247250b7fa504adeb94d0d1281

Observation d9751ca0-70fd-4ef2-9451-2f6f5f1fea40 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents React: Synergizing reasoning and acting in language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.971452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:7b3e0070b5778a7f37eb8939aaae53dd9bb99736f15f92d532fe77ff87ea71d9

Observation 67cf3cf0-a356-41bd-9643-755e0726ad0e · outbound

This paper cites Xing, 38 Hao Zhang, Joseph E.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Xing, 38 Hao Zhang, Joseph E

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:25:39.973730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:1066f2308360f7a1373c862c9dbb592f905a3bf94d91dc1c2384f1ad71696d38

Observation 28172c0f-e066-4465-9f98-3e897d6c0f6f · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:44:45.290322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:2d4114b7896d57287bcd42894d174d151ee0e29b4477a7671fb8c35e34998171

Observation 4b09c56d-48c9-40a3-9663-6108f62a622b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:44:45.269769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:38:43.708578Z digest=sha256:4ec0d5cb5530e575c398bb0c25ee64da75ef38d3a74577b3ec0b3622bcb006f7

Pith citing papers

Observation fac75aa7-8f2a-4ae8-baa5-61d393044d77 · inbound

AI Agents Do Not Fail Alone:The Context Fails First cites this paper.

AI Agents Do Not Fail Alone:The Context Fails First ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:38:46.490717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:38:46.490717Z digest=sha256:70e92244e31666ddc3969e1b3c786a3021a35bdcf35939676ce76457abd69e32

Observation cb2b457c-a9b9-421f-830d-a87abf592a58 · inbound

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness cites this paper.

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-01T03:34:07.269196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-01T03:31:34.165368Z digest=sha256:de1f5f396f65d285315f248db05012648725bab7820d43aa43d1656f95b51627