Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:01.373037Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.11323.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:01.373037Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69fc04de-924a-43d5-952a-d2f92e579858 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332b27dd-87d8-4de5-b8b6-ccdd04592b8c · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Quantifying Variance in Evaluation Benchmarks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aedec13-f0f3-4bf1-9a51-4ad62f1a451c · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b887848f-9d90-42cb-b23d-a1d1c8a8f796 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88477b10-f89f-44f7-ac71-b3c8c702842a · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db44260c-5bec-440f-9530-32e592f4064e · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 384dd9b7-c37a-4737-8b14-efc908bb4961 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f04fb3-10a1-4e55-888c-c4ce3341e668 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22278dd-6d74-44fc-994c-fc5b120df8f0 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f57a5502-39f8-48e0-a75b-a793a9ac26af · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Stop Comparing LLM Agents Without Disclosing the Harness
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da7396a-b91e-48bb-8ee5-177e866cf05f · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
Reference 1972
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18378a75-b1d5-4ccd-82b4-2da5827b080a · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31873485-17ff-467c-aa5c-148242dc43bd · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations tinyBenchmarks: evaluating LLMs with fewer examples
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b72ee5b7-3ca5-4228-bcf9-dbe4c0528c8b · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Why Do Multi-Agent LLM Systems Fail?
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af3943e-88d0-445b-85a2-27bff4a3f0e3 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Holistic Evaluation of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8b6ed6-75a9-46f3-94ad-8de40ec916a2 · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 647d245e-3e52-47dc-af7b-04d21a150fec · outbound
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.