Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T05:20:01.498711Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2605.21404.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T05:20:01.498711Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:03.765535Z
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1319b642-c278-4345-a346-6dea9f29fee6 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema SWE-bench: Can Language Models Resolve Real- World GitHub Issues?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19dc41bb-0c6b-445e-a290-a7945118a295 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Introducing SWE-bench Verified
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 246da87d-6822-41d7-8959-acbeed37cf3e · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92100b1f-5529-4cc3-8027-9c8d3f0510bf · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0885c3e-0050-47f9-8b81-0ac3d0b92bd3 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Mind2Web: Towards a Generalist Agent for the Web
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4a2bd0a5-4fb1-4a8f-8816-e1a73e304e69 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema OSWorld: Benchmarking Multimodal Agents for Open- Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 988d3a46-306a-46a7-a11a-a6267d881eed · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema GAIA: A Benchmark for General AI Assistants
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b684f4ac-d9d6-4aa1-8bfc-28bf9eb51d12 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema AgentBench: Evaluating LLMs as Agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e0f895b7-277c-4f02-bce0-c5c94f159192 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema AgentBoard: An Analytical Evaluation Board of Multi-Turn LLM Agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b54df31e-e155-4fd5-aaf7-322cf87c742d · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema MLE-bench: Evaluating Machine Learning Agents on Machine Learn- ing Engineering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 42b2d626-6e83-4516-a93a-221a8542bfb4 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Evaluating Large Language Models Trained on Code
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c0f863b-150b-40ea-b177-4aef3d3619ee · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Program Synthesis with Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a0410ad2-d98a-4fa5-acbe-0db7cf0e1888 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Training Verifiers to Solve Math Word Problems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d8a25fe-c516-4d7b-a070-db0f3027c6b7 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Measuring Massive Multitask Language Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 72fcd949-b342-448f-a49e-b6c0afef3062 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Deep Reinforcement Learning that Matters
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8bdff259-05b9-4d1f-b6aa-822671318921 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Improving Reproducibility in Machine Learning Research
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0fbc7ef-815b-411b-a48e-a9f54ecb7e5c · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Model Cards for Model Reporting
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2a35313-60fa-4652-bb99-de9d653b90c0 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Datasheets for Datasets
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 85711ad8-57a5-46cd-9cbf-518d55031ab4 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7150fcf3-12ea-49ad-a8da-cc54ec50bc89 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema State of What Art? A Call for Multi-Prompt LLM Eval- uation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1fdeccb2-3166-49e3-89db-d06c18bb22ec · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3bf7f789-910d-43e1-8487-79dc26a38268 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Data Contamination: From Memorization to Exploitation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0314c8b7-13c6-4e29-b33b-d78a70855a45 · outbound
What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Proving Test Set Contamination in Black Box Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ea9a4f83-5ac8-44db-80e7-18816f51053f · inbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.