Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T18:01:55.716141Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.19704.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T18:01:55.716141Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff0fd833-e034-4644-8f5e-8e0f421de361 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Assetopsbench: Benchmarking ai agents for task automation in industrial asset operations and maintenance
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c4eda0b-ae80-4a80-bf17-e6c75b03229d · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , url=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22815be2-c235-482a-8200-94c1d5d31fbe · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ba79df8-d1b7-4ca3-b2ca-fd05a87bdd1c · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2509.17158 , archivePrefix=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01bd111a-95b1-4b04-9342-e53a24e019e2 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Advances in Neural Information Processing Systems , volume=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7e4b51-be0e-434d-b848-d4eecba4f45a · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents General Agent Evaluation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47094d9c-34d2-4008-90f0-43fa7578114a · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents arXiv preprint arXiv:2603.08171 , year=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc703a00-ea42-42f0-bd74-c7bd8777267c · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents International Conference on Learning Representations , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a67d77-5cd2-4ca7-acfa-006c5c98a92f · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 440ff841-087f-4b2a-a024-ded051d40fde · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Advances in neural information processing systems , volume=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d0ca40-f91a-422a-88f6-9443fe359276 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ccfbc5-4ee4-4b33-8164-290bec07d9ee · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2363c7-cb04-4d68-bcf8-1fbe6b4d7268 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47047fe-12aa-462d-9a36-744243c73dc7 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Skills and Knowledge Plugin
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1eb952-a265-4f20-ab76-309abf561c7a · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Skill-Knowledge-Augmented Agents on
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005900d7-7efe-43ac-bdde-b5df9c999fa5 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d423bbb8-69e1-4dfe-ae6f-39d39a3f53c1 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents 2026 , note =
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e76ab5c3-f8f6-4621-82ff-5aa426845c4e · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Extending
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380b02eb-e37c-4ddd-a938-d11eedb96039 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92f60392-ad13-43cf-8481-f0315401c918 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d7392e-9469-4ec6-a18b-4e693a86ad3c · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents , institution=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab315bef-0c50-4bd3-b952-6c3adc9c4c75 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Profiling and Optimizing the
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee71c25-740f-4531-a9f0-5ffceded1700 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Performance Optimization of the
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed70051b-5385-4349-8506-10e18117d46a · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Multi-Modal Agent Inference Optimization for Industrial Asset Operations: Extending
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9234c39c-5615-44da-bf09-5cf7d4ca80fb · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Holistic Evaluation of Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95f51d96-0c70-4ea5-b302-e9f97d5ac34d · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2021 conference of the North American chapter of the Association for Computational Linguistics: human language technologies , pages=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca166462-f6a9-4716-b23c-53deb31b4e73 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 58th annual meeting of the association for computational linguistics , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3f3cad-9b7e-41ae-8ac0-260cbc075fcc · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aecd48d-c9ed-4246-a7b2-43ca3d38ce7a · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents AI and the Everything in the Whole Wide World Benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 168ddf07-d86c-4d5f-be00-b86a77fff2fa · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d6f2a9-947a-43dc-b530-d397240f62d3 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents International conference on machine learning , pages=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162ba649-9612-4323-a7ca-b8e49fb35ba0 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents The Benchmark Lottery
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c67cb534-4b52-42ca-9309-07186f6b271c · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Interactive Evaluation Requires a Design Science
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fdffce0d-9d50-429e-8e1c-0a6dd516e1ac · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 749b2ad1-4c0d-4f07-b09a-7d6e23613dc9 · outbound
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents The Evaluation Trap: Benchmark Design as Theoretical Commitment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.