Pith. sign in

Paper Citation Record · LEDGER

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation

As of 11 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2605.10448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10448 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:55.592359Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:12:04.575194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact13
  • verified fuzzy13
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17bac3e9-2400-4289-9c67-e7de6eb105fb · outbound

This paper cites The BrowserGym Ecosystem for Web Agent Research.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation The BrowserGym Ecosystem for Web Agent Research

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.874905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:ecb8aa8b8ad3f86774c02bcbcc34e7ed85e64069ecdf26be88860d00ed4d1c87

Observation 1b876bbc-9c86-408d-8e5b-5b81ebe6e118 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:35:13.649872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:c2c4e9c5b341123e58af779a77e09b01e5b36786c457b49905945f9f67a87408

Observation 61f73cc8-9d45-4d3f-9160-045154203a55 · outbound

This paper cites Mind2Web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Mind2Web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.869481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:9d93d9a80ca121ed0e5713b4f46990af5932b773396f9780d3de507c47336618

Observation 30913e86-a02b-43d2-b37b-b9d02a88d526 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:48:05.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:9edee2966fffb02ac7c23fda19de023209f60f815c8ada7b513120981faf7ebc

Observation b336413f-c3f9-4a4d-a163-41a64cbfce97 · outbound

This paper cites URL https://cacm.acm.org/research/ datasheets-for-datasets/.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation URL https://cacm.acm.org/research/ datasheets-for-datasets/

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-12T05:06:20.544655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:22415c08bbf4f104c873b7a2cd1458389092707030df52a64e1394ab564305a4

Observation c8ea4181-b363-423a-8ace-3bf2fa8d95ae · outbound

This paper cites Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:19:37.552446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:7c836eb48eeb3c18f7062599ea89d7fd1a7d8257e1d3b8c8230a48f125dbc586

Observation 548f8a01-78e1-4594-ba10-2305f9167431 · outbound

This paper cites Webarena verified: Reliable evaluation for web agents.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Webarena verified: Reliable evaluation for web agents

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.872445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:387c2bbad827d156f07b0f82d2e3f8650a39060576ee0e05b454f9dbcedec49d

Observation 3af0832f-5189-482d-9bfd-b2cb94590fad · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Dynabench: Rethinking benchmarking in NLP

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.833771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:629bd948411b85a9fe3a4acc6375279d4236e2574266136f1b88bb2dbb7519c4

Observation 2f7b4994-2443-4319-8f51-e27681c4c782 · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.837062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:b87dbc06d995f6e7629417729bb7ee9c5028321815935495ac527cf40414ece7

Observation 2bee8ed9-77b1-4d96-b00f-6b217917f49d · outbound

This paper cites Reinforcement learning on web interfaces using workflow-guided exploration.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Reinforcement learning on web interfaces using workflow-guided exploration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.844136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:bfed8ecc4252a502a665066ed8e5d1dc7c06de71c035eb8f295067abe961d24d

Observation 9579924e-d123-450e-b424-8074c1f8cf98 · outbound

This paper cites ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.840249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:3f14a90016b1eea570197e0ab1b6f5001b0b416e634d6ef097708050f3a60ce7

Observation b67ec97d-9241-489b-94b6-cc9bb4090082 · outbound

This paper cites Agentrewardbench: Evaluating automatic evaluations of web agent trajectories.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Agentrewardbench: Evaluating automatic evaluations of web agent trajectories

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.939647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:54a07789e2eba1b168b8a8c56863cc244f27cb3e0de49de6c7fb24fe8e242dc3

Observation f94f7301-ce97-49d9-8ff7-15993292d181 · outbound

This paper cites SPHERE: An Evaluation Card for Human-AI Systems.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation SPHERE: An Evaluation Card for Human-AI Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:07.662383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:5237849357842ee4c6674e3ac483300a11ecb158ebb2e6703a01f27166c24828

Observation 86591807-bec2-497d-b971-56310094afc0 · outbound

This paper cites Manski.Partial Identification of Probability Distributions.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Manski.Partial Identification of Probability Distributions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.850059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:f03d355843fe8e806cd823734a9f7c2dd70a2750eabf8cdb2f4dd3505be40f31

Observation 1573e894-c67d-49f1-8380-eebe4f86983d · outbound

This paper cites Model Cards for Model Reporting.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Model Cards for Model Reporting

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:06:20.541842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:43920554d4cedcb0e8e8e96edc7ec99ad8c5142d2620d697a5f99ef93b35713b

Observation c4c6e752-5904-4b9a-9ba8-bc00e65e7ae6 · outbound

This paper cites Introducing SWE-bench Verified.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Introducing SWE-bench Verified

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.847178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:661b82fefc29b249b1bc269e6be5c3108fb028963729e042b92db25150cc1fb0

Observation 1943f7c6-c500-4b55-b7a1-529c5a33a49f · outbound

This paper cites Why SWE-bench verified no longer measures frontier coding capabilities.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Why SWE-bench verified no longer measures frontier coding capabilities

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.859215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:c22c834e19b465374898fb5cbb771e32f2c4398489f3c59536e8c54384eceea7

Observation 236405e5-2c75-46d2-aa4f-40eebe10e46a · outbound

This paper cites Improving reproducibility in machine learning research: A report from the NeurIPS 2019 reproducibility program.Journal of Machine Learning Research, 22(164):1–20.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Improving reproducibility in machine learning research: A report from the NeurIPS 2019 reproducibility program.Journal of Machine Learning Research, 22(164):1–20

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.866354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:48936733e1e469c20d457d4bbdcba322da6c99660860d0208d2bf1b698874288

Observation 3362711c-d4b7-46b6-95bc-891fd35ab2c2 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:24.070271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:1e03cf1c99df7f48ae297ef3852ef003e773bcb8bd808939180e82a11492b65e

Observation 5b288a5c-83b2-40ea-bfff-86c3495c4f9f · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:06:13.928391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:b46a24a32e3815535e521d35afada862eccdbd55f6e56f8b71470f1cee7dfa74

Observation 2daf2ade-6636-41b8-93db-cc41ec69d724 · outbound

This paper cites Judging the judges: A systematic study of position bias in llm-as-a-judge.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Judging the judges: A systematic study of position bias in llm-as-a-judge

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.042185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:cbe63deadf7699711fe9f6586f988452f6a1378c727a751ce104981a14828c4c

Observation 7ec99599-eb9d-4aa9-aa98-ee6eeebc3797 · outbound

This paper cites τ 3-Bench: Advancing Agent Benchmarking to Knowledge and V oice.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation τ 3-Bench: Advancing Agent Benchmarking to Knowledge and V oice

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.862489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:49aba3c0f204ece6a149a688dc84fa645b838a8de159fbbd1967ad102e1a1cb8

Observation 9879d5af-1a77-4d9c-bef4-a0ba70bc296e · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:23.972146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:7bf06a8164e330154f99640b826ac3c372529aaf26010c49099313501ca79b17

Observation 96da9926-0689-40ad-96b1-0a91b4a23dc7 · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:24.023340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:6d419289cc6f5a909840ea435c5f49a8fea857baeb6e72b2cf99a38070dc5e84

Observation ec5fb971-a6ec-4eb9-9ade-2d989e18b270 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:c6db9fedefa2169a92b52bdcd82c50af105e61a5522af0d9c2f5b0bc6739317a

Observation 65b485ef-6d41-4c0a-bb1b-7669705ee599 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:ad98b13c40e3c8763e6e31e906ea711220549779ef0a7a0d74095f7dab9a8493

Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.900869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:5a1d8890d9aa89d21850fb9d5f2fbde100ec6dec7207640ef5d33c787ea1dc14

Observation 630bdada-b3a6-444d-a1e5-5dc869b02cbe · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.853166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:f7081b71cea5c83857d025931fbb8c8c32e77da241336cf36c17a7a61d597ad1

Observation af19ef67-dc1c-4cf1-b85d-7cbd4949c9dd · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:23.924020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:0df532e23bba0644f595c6fd7b8e478f99bed97c519cb82b0a8eec9c8f3af093

Observation 2afea040-8111-4b17-aa6a-2c454a7af7ee · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:41:24.057109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:859c4e9733951a6d73bfca5f3632dc011267396fba8a0278c6c276492972efb2

Observation c81d488d-75f2-4ddc-b5bf-53517c82d157 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:41:23.882610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:67cbb4dd0a70c296d6f7ca363db21f7b64f317f7d8982bd8e6ca607160828701

Observation a2501f56-e71e-46c0-a8a3-975374a12925 · outbound

This paper cites Establishing best practices for building rigorous agentic benchmarks.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing best practices for building rigorous agentic benchmarks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:16:35.856357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:b2330b5de73ca2df6eb9a0cb84ca25cf94e511c5526a33d1ae6498dad95fd1ab

Observation 927faacd-9882-41fd-a98c-8bd197aaf590 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.002297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:889968d39a111300f8ec9f76c741a2b47d3fb9e613552360aef922aa52a6a59a

Pith citing papers

Observation bff4d13a-2789-4a36-8f58-7a24deb8574a · inbound

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation cites this paper.

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T23:12:04.575194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:12:04.575194Z digest=sha256:26515e17b3a89ae90157776cf0e9c74623de555c7014c5ca269e23204a8c6209