Pith. sign in

Paper Citation Record · LEDGER

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.19704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.19704 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T18:01:55.716141Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0fd833-e034-4644-8f5e-8e0f421de361 · outbound

This paper cites Assetopsbench: Benchmarking ai agents for task automation in industrial asset operations and maintenance.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Assetopsbench: Benchmarking ai agents for task automation in industrial asset operations and maintenance

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:30.300789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:712d617c74dd32e5c3057e7e89b1a3f15af4d1f0e1cd9a4356df8c4f983a2bf3

Observation 3c4eda0b-ae80-4a80-bf17-e6c75b03229d · outbound

This paper cites 2026 , url=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , url=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:4d3d64a26fe3547df24ed74d7a04935b7d116c282579f4db00973887e051a45a

Observation 22815be2-c235-482a-8200-94c1d5d31fbe · outbound

This paper cites MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:30.310286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:4edb41e02145ce30225aef860994b6d2e6ccd33afc5045284fa8586d3b3a8f95

Observation 7ba79df8-d1b7-4ca3-b2ca-fd05a87bdd1c · outbound

This paper cites 2509.17158 , archivePrefix=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2509.17158 , archivePrefix=

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.295775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:7593b42807288838704329c39b134f386b931cf0615652172b76b4297d898798

Observation 01bd111a-95b1-4b04-9342-e53a24e019e2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Advances in Neural Information Processing Systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:606f3a0e2b277f2251c75980b4c708aed4fbdf58e7f77a0f8b94c545c6a8f1d2

Observation ab7e4b51-be0e-434d-b848-d4eecba4f45a · outbound

This paper cites General Agent Evaluation.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents General Agent Evaluation

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.290814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:7878f86d4defc3575b19e951a5f508be80f3dfbc5ae4974a1a956010be7b0668

Observation 47094d9c-34d2-4008-90f0-43fa7578114a · outbound

This paper cites arXiv preprint arXiv:2603.08171 , year=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents arXiv preprint arXiv:2603.08171 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.338616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:e032693a2fb82f9d836d9f93ecae9a9378450b7a704c07f2595023ce95d63839

Observation fc703a00-ea42-42f0-bd74-c7bd8777267c · outbound

This paper cites International Conference on Learning Representations , volume=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents International Conference on Learning Representations , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:85a949731fe4625ac480b725c9bde1e7435ad5fde29edf6a7cdf79560ef29266

Observation 60a67d77-5cd2-4ca7-acfa-006c5c98a92f · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.282236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:e0177fff411e5bdc94f502fe912fcdc0e2b5c1ffce93f25855a75d717cd14a0f

Observation 440ff841-087f-4b2a-a024-ded051d40fde · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Advances in neural information processing systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:027ebd2bea1bd07027a22d2f6c6106f66f72a5d6c6d35a8fc3230902c9e43be0

Observation d7d0ca40-f91a-422a-88f6-9443fe359276 · outbound

This paper cites 2026 , note =.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:ece50bef0b5711e670f039cc2bce19e81c81e65232666644e97c7fccd83b937b

Observation 85ccfbc5-4ee4-4b33-8164-290bec07d9ee · outbound

This paper cites 2026 , note =.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:e99bd2672b47436381beae071ee75a8f63f6262a5f047e02b0aa20961b33ede8

Observation 6b2363c7-cb04-4d68-bcf8-1fbe6b4d7268 · outbound

This paper cites 2026 , note =.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:3c4497f2601482e0aed9238fb177a9b0d74b636d05248ba6455e3a4d7893a1e9

Observation a47047fe-12aa-462d-9a36-744243c73dc7 · outbound

This paper cites Skills and Knowledge Plugin.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Skills and Knowledge Plugin

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:0a2ebe56578e424220c7e6eb55c7f73f48e9b68f2bcbda1116f9b347bca8d4b1

Observation 8e1eb952-a265-4f20-ab76-309abf561c7a · outbound

This paper cites Skill-Knowledge-Augmented Agents on.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Skill-Knowledge-Augmented Agents on

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:28c6ae2394b5252068668cab42665ff965c4172a0911ae9c208e8a30a583a9cd

Observation 005900d7-7efe-43ac-bdde-b5df9c999fa5 · outbound

This paper cites Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.286771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:84ca564592fa702ca13ae84574b3ddd10de18d508658aae53b5cd827eec88ef3

Observation d423bbb8-69e1-4dfe-ae6f-39d39a3f53c1 · outbound

This paper cites 2026 , note =.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:c16de161d04b976f84fbc2f4becb202ff7e1edc7ce6993a99be0328783f1cdda

Observation e76ab5c3-f8f6-4621-82ff-5aa426845c4e · outbound

This paper cites Extending.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Extending

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:bed78255ebfec1f728393fa91b3a31ad286e28db25ac590a9830554baf5cc527

Observation 380b02eb-e37c-4ddd-a938-d11eedb96039 · outbound

This paper cites PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.332866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:3a6266975af031ab28540982b4842d24dd458884118667594ee0b54dbaf2760c

Observation 92f60392-ad13-43cf-8481-f0315401c918 · outbound

This paper cites an unresolved cited work.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:411e84ded033c0e70b0c980412ed9334decaf583b560a1f649ee49527dcd8c1a

Observation a0d7392e-9469-4ec6-a18b-4e693a86ad3c · outbound

This paper cites , institution=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents , institution=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:b646fc280d7a991aabff6f4aec5174ef53fa57e3992978c434af0a0c5aa4c059

Observation ab315bef-0c50-4bd3-b952-6c3adc9c4c75 · outbound

This paper cites Profiling and Optimizing the.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Profiling and Optimizing the

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:18b5e439257f99c3c9f7985ba7425d93f1779e58107aa8e32e8675a7625102cf

Observation 0ee71c25-740f-4531-a9f0-5ffceded1700 · outbound

This paper cites Performance Optimization of the.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Performance Optimization of the

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:acd2957f0c7c865bb85e217fd547afdb0c8c8733bf5d0ad923f23de06be5e94c

Observation ed70051b-5385-4349-8506-10e18117d46a · outbound

This paper cites Multi-Modal Agent Inference Optimization for Industrial Asset Operations: Extending.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Multi-Modal Agent Inference Optimization for Industrial Asset Operations: Extending

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:ad13ff4f5b4e732d2556df820373ee844581160fb7b1dff1c63c9a61792a00b9

Observation 9234c39c-5615-44da-bf09-5cf7d4ca80fb · outbound

This paper cites Holistic Evaluation of Language Models.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Holistic Evaluation of Language Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.320599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:d064a29dd16e94869a5acccb2912ed7ff5873ae563251c4b68f96f519111fc12

Observation 95f51d96-0c70-4ea5-b302-e9f97d5ac34d · outbound

This paper cites Proceedings of the 2021 conference of the North American chapter of the Association for Computational Linguistics: human language technologies , pages=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2021 conference of the North American chapter of the Association for Computational Linguistics: human language technologies , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:6401eae7f8efc23a16bbf1d0d3ee6a86bb12e2d3eb3a7e0ce67297be9a92e0e8

Observation ca166462-f6a9-4716-b23c-53deb31b4e73 · outbound

This paper cites Proceedings of the 58th annual meeting of the association for computational linguistics , pages=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:ec147c8007a7acd40a308f8ff9b1afaad240d4140b4ef675e61989b70bb36c8c

Observation ae3f3cad-9b7e-41ae-8ac0-260cbc075fcc · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:4103f383857fd0a42d77ce83f025b01fd2ac5438c644d31db63edd867c278d76

Observation 1aecd48d-c9ed-4246-a7b2-43ca3d38ce7a · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents AI and the Everything in the Whole Wide World Benchmark

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.305905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:a43acb139f82b0b8a1bccc212ea53f40dcb581f0a8eab85f469fc3abaaf3a526

Observation 168ddf07-d86c-4d5f-be00-b86a77fff2fa · outbound

This paper cites Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:00b6752a719e0bc4a35ed1bfe78af0f419637d097d569899de5adae8c298f10f

Observation c4d6f2a9-947a-43dc-b530-d397240f62d3 · outbound

This paper cites International conference on machine learning , pages=.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents International conference on machine learning , pages=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T18:01:55.716141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:2b2c15eb8b8f4ea4fa2c79737b88fcb5662f6292d8bb719ad1ca1601ef6f97b2

Observation 162ba649-9612-4323-a7ca-b8e49fb35ba0 · outbound

This paper cites The Benchmark Lottery.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents The Benchmark Lottery

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:30.325999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:17d04481489d3f7e36594bc7d642baad8c70f46ffbcadb2753e971ac0909dd75

Observation c67cb534-4b52-42ca-9309-07186f6b271c · outbound

This paper cites Interactive Evaluation Requires a Design Science.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Interactive Evaluation Requires a Design Science

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.272826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:c6f9831830542155225de980098220d28b9d393d6b2b091b5bef1799044d313d

Observation fdffce0d-9d50-429e-8e1c-0a6dd516e1ac · outbound

This paper cites Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.342214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:adc9bb38a66bde6717114d42d0109e1634a7df8d38a1893dae5108dd2c5c100e

Observation 749b2ad1-4c0d-4f07-b09a-7d6e23613dc9 · outbound

This paper cites The Evaluation Trap: Benchmark Design as Theoretical Commitment.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents The Evaluation Trap: Benchmark Design as Theoretical Commitment

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.314635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:67f868d61cb735786ce86fc37d149ad1f69ccaad315777bf7e167f36f6900cd6

Pith citing papers

No inbound Pith citation observations are available.