Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:55.592359Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2605.10448.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:05:55.592359Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T23:12:04.575194Z
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 17bac3e9-2400-4289-9c67-e7de6eb105fb · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation The BrowserGym Ecosystem for Web Agent Research
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b876bbc-9c86-408d-8e5b-5b81ebe6e118 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61f73cc8-9d45-4d3f-9160-045154203a55 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Mind2Web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30913e86-a02b-43d2-b37b-b9d02a88d526 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b336413f-c3f9-4a4d-a163-41a64cbfce97 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation URL https://cacm.acm.org/research/ datasheets-for-datasets/
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8ea4181-b363-423a-8ace-3bf2fa8d95ae · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 548f8a01-78e1-4594-ba10-2305f9167431 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Webarena verified: Reliable evaluation for web agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3af0832f-5189-482d-9bfd-b2cb94590fad · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Dynabench: Rethinking benchmarking in NLP
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f7b4994-2443-4319-8f51-e27681c4c782 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2bee8ed9-77b1-4d96-b00f-6b217917f49d · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Reinforcement learning on web interfaces using workflow-guided exploration
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9579924e-d123-450e-b424-8074c1f8cf98 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b67ec97d-9241-489b-94b6-cc9bb4090082 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Agentrewardbench: Evaluating automatic evaluations of web agent trajectories
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f94f7301-ce97-49d9-8ff7-15993292d181 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation SPHERE: An Evaluation Card for Human-AI Systems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 86591807-bec2-497d-b971-56310094afc0 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Manski.Partial Identification of Probability Distributions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1573e894-c67d-49f1-8380-eebe4f86983d · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Model Cards for Model Reporting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4c6e752-5904-4b9a-9ba8-bc00e65e7ae6 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Introducing SWE-bench Verified
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1943f7c6-c500-4b55-b7a1-529c5a33a49f · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Why SWE-bench verified no longer measures frontier coding capabilities
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 236405e5-2c75-46d2-aa4f-40eebe10e46a · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Improving reproducibility in machine learning research: A report from the NeurIPS 2019 reproducibility program.Journal of Machine Learning Research, 22(164):1–20
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3362711c-d4b7-46b6-95bc-891fd35ab2c2 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b288a5c-83b2-40ea-bfff-86c3495c4f9f · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2daf2ade-6636-41b8-93db-cc41ec69d724 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Judging the judges: A systematic study of position bias in llm-as-a-judge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ec99599-eb9d-4aa9-aa98-ee6eeebc3797 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation τ 3-Bench: Advancing Agent Benchmarking to Knowledge and V oice
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9879d5af-1a77-4d9c-bef4-a0ba70bc296e · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96da9926-0689-40ad-96b1-0a91b4a23dc7 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec5fb971-a6ec-4eb9-9ade-2d989e18b270 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 65b485ef-6d41-4c0a-bb1b-7669705ee599 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 630bdada-b3a6-444d-a1e5-5dc869b02cbe · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af19ef67-dc1c-4cf1-b85d-7cbd4949c9dd · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2afea040-8111-4b17-aa6a-2c454a7af7ee · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c81d488d-75f2-4ddc-b5bf-53517c82d157 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a2501f56-e71e-46c0-a8a3-975374a12925 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing best practices for building rigorous agentic benchmarks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 927faacd-9882-41fd-a98c-8bd197aaf590 · outbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bff4d13a-2789-4a36-8f58-7a24deb8574a · inbound
Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.