Pith. sign in

Paper Citation Record · LEDGER

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2605.21404.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21404 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T05:20:01.498711Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:03.765535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1319b642-c278-4345-a346-6dea9f29fee6 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real- World GitHub Issues?.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema SWE-bench: Can Language Models Resolve Real- World GitHub Issues?

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.780117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:e7b82764e5173099127bbb448642593ca032719b114e0104d66c380a2fef8fd0

Observation 19dc41bb-0c6b-445e-a290-a7945118a295 · outbound

This paper cites Introducing SWE-bench Verified.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Introducing SWE-bench Verified

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.777505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:424fa44dac116f3c71501892e33499b581abd175ed741ed8d85969d049004d52

Observation 246da87d-6822-41d7-8959-acbeed37cf3e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.771129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:034d45f7d61929a67c279970168526355110d14f76cb28a601daed129e472d7b

Observation 92100b1f-5529-4cc3-8027-9c8d3f0510bf · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.766818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:666edda43e182fd4444fc388f45fde92d0dd82aef6f72db68d868b34fb1f41cb

Observation b0885c3e-0050-47f9-8b81-0ac3d0b92bd3 · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Mind2Web: Towards a Generalist Agent for the Web

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.768957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:ae5912fac1277517c4576440c36c9788c9bf0df9f957e59e32895162f269277e

Observation 4a2bd0a5-4fb1-4a8f-8816-e1a73e304e69 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open- Ended Tasks in Real Computer Environments.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema OSWorld: Benchmarking Multimodal Agents for Open- Ended Tasks in Real Computer Environments

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.762101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:092d0f17b2658c6fffdb67ddfb8871aa612437971e03ad5c842648128b1b6fd7

Observation 988d3a46-306a-46a7-a11a-a6267d881eed · outbound

This paper cites GAIA: A Benchmark for General AI Assistants.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema GAIA: A Benchmark for General AI Assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.764433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:c27cdfe5aabb2c38adaf35da2e86ae17bb7295ffe5ab21ec4c73464b7d9a9eff

Observation b684f4ac-d9d6-4aa1-8bfc-28bf9eb51d12 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema AgentBench: Evaluating LLMs as Agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.773406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:f7248c45192f36965337b1302817fd0e2d36cc2beb6a82694d4fb4ad80f734b7

Observation e0f895b7-277c-4f02-bce0-c5c94f159192 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-Turn LLM Agents.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema AgentBoard: An Analytical Evaluation Board of Multi-Turn LLM Agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.775511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:16c45922dcf1e6f50117d2a475b99d8026e3685389a2fbb49c661245431c701a

Observation b54df31e-e155-4fd5-aaf7-322cf87c742d · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learn- ing Engineering.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema MLE-bench: Evaluating Machine Learning Agents on Machine Learn- ing Engineering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.782088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:9a5216c0c3c219361a0102dfa02947b3aab97674d5fec18ae5aec6fc27655946

Observation 42b2d626-6e83-4516-a93a-221a8542bfb4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:23:58.699582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:e7a2d9ab7d18cced9302ff4c540aed5941816352fcfad890736f14bfeecccdf3

Observation 5c0f863b-150b-40ea-b177-4aef3d3619ee · outbound

This paper cites Program Synthesis with Large Language Models.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Program Synthesis with Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:23:58.696044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:9750b0671d586e3dee36414f23e2663cfdebdd6ac701bb37cc3cfba507bd6097

Observation a0410ad2-d98a-4fa5-acbe-0db7cf0e1888 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:23:58.702679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:a8786f123366e09a7875ed770544a8c96c5d836e2ac87706ceaff5cea9ece454

Observation 9d8a25fe-c516-4d7b-a070-db0f3027c6b7 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Measuring Massive Multitask Language Understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.784014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:bd6857b5dce081771a854cc0588f7aa0c68e4c621cd3608738c00265b9c85b5c

Observation 72fcd949-b342-448f-a49e-b6c0afef3062 · outbound

This paper cites Deep Reinforcement Learning that Matters.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Deep Reinforcement Learning that Matters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.794891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:77bb35be836a752184697a6aa5be29cdb0e79f040aea8fecfba99973211e3321

Observation 8bdff259-05b9-4d1f-b6aa-822671318921 · outbound

This paper cites Improving Reproducibility in Machine Learning Research.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Improving Reproducibility in Machine Learning Research

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.796945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:eae04682b9913719d6831fc28c279a34af4e081c5a45be512c67d4f860cbce00

Observation b0fbc7ef-815b-411b-a48e-a9f54ecb7e5c · outbound

This paper cites Model Cards for Model Reporting.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Model Cards for Model Reporting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.790556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:65fbcbb5312f210d4182db5995fd5a788ed0d7a83821f4c6efb2774cd09c1c72

Observation d2a35313-60fa-4652-bb99-de9d653b90c0 · outbound

This paper cites Datasheets for Datasets.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Datasheets for Datasets

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.803226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:e8e502da67ba2aeee86e2441e4b291b914a1671603958a56b4a84c6e3e490319

Observation 85711ad8-57a5-46cd-9cbf-518d55031ab4 · outbound

This paper cites Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.786208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:19548d7135b0889ec70960efebffd1dcc0f860cda6d37f816367666802a5db98

Observation 7150fcf3-12ea-49ad-a8da-cc54ec50bc89 · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Eval- uation.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema State of What Art? A Call for Multi-Prompt LLM Eval- uation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.788401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:89da256be5dac0cbe95fcd3fb5920ae393c796e2f2573e5c28905ba9b2754934

Observation 1fdeccb2-3166-49e3-89db-d06c18bb22ec · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.792783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:4ca6d20bf455a529ca384d21275e2238938119347ecf11264b0875e9c5d138ca

Observation 3bf7f789-910d-43e1-8487-79dc26a38268 · outbound

This paper cites Data Contamination: From Memorization to Exploitation.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Data Contamination: From Memorization to Exploitation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.799033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:01e00ad5c0302440d90bc20dc2824f143a8748cac1285625e033afe17f16eca0

Observation 0314c8b7-13c6-4e29-b33b-d78a70855a45 · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Proving Test Set Contamination in Black Box Language Models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T05:24:39.801147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:20:01.498711Z digest=sha256:68e5c2236808626b29881f9a58e54a2899c155a25af76193cea6dea9e40060df

Pith citing papers

Observation ea9a4f83-5ac8-44db-80e7-18816f51053f · inbound

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision cites this paper.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.765535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.765535Z digest=sha256:997ea6686d3e372bb65ef86dfe02d0970d6da2ded4eeb2696477a262d7963db7