Pith. sign in

Paper Citation Record · LEDGER

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2608.02444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02444 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:04.462120Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6bc585d6-d3f4-4f24-a23c-d177e6a5f6aa · outbound

This paper cites Efficient benchmarking of AI agents,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Efficient benchmarking of AI agents,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.474902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.474902Z digest=sha256:d73eabf1c0190df51a3c6e2b39881c49f1033c7999fcfecf3e74888ede99266f

Observation 92eeb214-c9da-4e4e-98e0-76bf282c353a · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues?.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision SWE-bench: Can language models resolve real-world GitHub issues?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.581554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.581554Z digest=sha256:bb2d2c1d1a64087982ec10fdcc246a3036c26383c3953b5a917d466491e34a76

Observation 79952cdf-d754-4245-a16b-8bf184e87cbf · outbound

This paper cites SWE-bench leaderboards,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision SWE-bench leaderboards,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.644754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.644754Z digest=sha256:5f63aedd5ee2d14c5c6d7677b6ffcd238c6e8bd1957fda7715f544699fbd56ce

Observation 5ccd03e1-e641-4d5b-ae85-d74efbe8308e · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.792257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.792257Z digest=sha256:d79618e3fd3459ce8f9db8054593b30555da3a9ae641b52a9a98730b5209e219

Observation 1b8568ab-4219-4297-b317-7b68302d4ed2 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.939210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.939210Z digest=sha256:a4db257047a5891c045de64306e3e5ff198926a4b96e4869915691f4961850e9

Observation 74d1fca5-746a-4604-8d0c-8c63bee87aec · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.008966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.008966Z digest=sha256:410f39e3f9d8761639c5dd588bc9a4f884fc72f3753a8bed56b71a0ab8ecbe05

Observation 3f7c8a70-366e-4211-9552-44826efe4c5d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.134876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.134876Z digest=sha256:7d12ae4b0468d143eb4b33fcfa287e9f02dbb2a6624ac56085eff03381f79526

Observation 3cc61c9a-68e1-4bcb-9572-d3e069f682ba · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.235500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.235500Z digest=sha256:79f318f84390576a3bcddce62627316aa36cc8790295d40c4f1e07229e483fe3

Observation 8b46a2fb-923c-4a78-afb6-ffb4da4e48dc · outbound

This paper cites Holistic Agent Leaderboard: The missing infrastructure for AI agent evaluation,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Holistic Agent Leaderboard: The missing infrastructure for AI agent evaluation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.346033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.346033Z digest=sha256:63e62509c191c18f8f744e435abe3f5ab5ef4411281844d3ad8e5c05690b2848

Observation 734ac525-f565-4975-ae5d-2ef05c0342ce · outbound

This paper cites Sequential tests of statistical hypotheses,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Sequential tests of statistical hypotheses,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.458781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.458781Z digest=sha256:f10421a2bb94b0f9b6b0949e3adc84f3cf6793d62dc3a873b16b7da7400902da

Observation 6431829b-6e8f-4f81-b69c-a24188277490 · outbound

This paper cites Time-uniform, nonparametric, nonasymptotic confidence sequences,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Time-uniform, nonparametric, nonasymptotic confidence sequences,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.564760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.564760Z digest=sha256:a476e00011c9dbdda71dc3d813e5fb04cc4d58cb69d9a5532c4c9ec7a8a67124

Observation a94a15f9-e1c9-4cf2-bcac-fa24abe10bc4 · outbound

This paper cites Selectivenet: A deep neural network with an integrated reject option,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Selectivenet: A deep neural network with an integrated reject option,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.684572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.684572Z digest=sha256:289b9611608f2402c05f5af3066d7793aa8ed787fe4025743a121b042e6bde53

Observation 39e41e85-b040-4894-9449-735bd43a3d09 · outbound

This paper cites Consistent estimators for learning to defer to an expert,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Consistent estimators for learning to defer to an expert,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.781995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.781995Z digest=sha256:7923e69317e6c840ea9e6b953dde7898f9e736463c2ece6b56899dfab7f4499c

Observation 3ffb8890-b1a4-49d8-bc03-2ecdbb92b577 · outbound

This paper cites Conformal Risk Control.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Conformal Risk Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.947149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.947149Z digest=sha256:a85a39f8a4e461a03bdd205053fa676c695cd335e908f1f5cc53a18c03d80f82

Observation 7792a26e-f3aa-4191-abc4-f47f294f57c8 · outbound

This paper cites Holistic evaluation of language models,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Holistic evaluation of language models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.030711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.030711Z digest=sha256:a68137eed7811eca598cd48191d7b43e1f4fa284ede9a897e8eab55a7734321e

Observation 8d58f2f3-9dca-4162-ac89-0903f8683b9c · outbound

This paper cites Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.152718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.152718Z digest=sha256:8c16d42403a6f0c9bf6e77e4cbfbb6cd8173f648aa339c72027272ad9485dcaf

Observation 97598b0f-a068-43aa-84ad-40580d288ceb · outbound

This paper cites Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.239193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.239193Z digest=sha256:f6e7a335432f3efc8bec6e38adba975c0c07184ad1d78bdb031ed6f2bd9ff06f

Observation 7499cc94-7d84-4097-bd2b-d41870d8d26f · outbound

This paper cites Efficient Evaluation of LLM Performance with Statistical Guarantees.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Efficient Evaluation of LLM Performance with Statistical Guarantees

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.321698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.321698Z digest=sha256:4340dc18ccd99a2f47818def3b063261a940e26ef323e7697193e171dd31571b

Observation b31fe661-ddad-45a7-86c4-106aa6de11f4 · outbound

This paper cites Probability inequalities for the sum in sampling without replacement,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Probability inequalities for the sum in sampling without replacement,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.447337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.447337Z digest=sha256:17cebe29e5db6dfda10fec61c9b731c6404ec0201efe631efb420617ebe0bbaf

Observation 3f59e237-2679-4419-913c-2c0e66cc0c2f · outbound

This paper cites On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.554167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.554167Z digest=sha256:57135f7065251526cc458721fb151860384e8e78db7c59cc4b95f1929cb38602

Observation 761f7a42-9a06-4c22-945a-1902bccad3ac · outbound

This paper cites AI Agents That Matter.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AI Agents That Matter

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.668280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.668280Z digest=sha256:06f950b627419243c888636a42c8be59d26c9246d5a7588add1b5603c757dec2

Observation ea9a4f83-5ac8-44db-80e7-18816f51053f · outbound

This paper cites What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.765535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.765535Z digest=sha256:997ea6686d3e372bb65ef86dfe02d0970d6da2ded4eeb2696477a262d7963db7

Observation 511290d2-3acb-4861-a129-130a80233df0 · outbound

This paper cites General Agent Evaluation.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision General Agent Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.880384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.880384Z digest=sha256:c619b102c0cde46c008823018e80651a4bf6936f679486c045b75e042fdc2f2f

Observation 72ebbd39-fd75-4c41-9c4c-640a81ef4233 · outbound

This paper cites A2Perf: Real-World Autonomous Agents Benchmark.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision A2Perf: Real-World Autonomous Agents Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.013592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.013592Z digest=sha256:464d4bf0dcb6106ea6573b525c76fb41f8eebcfd9f658c576f5b74792d2a664e

Observation 05d92a8d-e366-44b7-87f8-f158dcc5a18a · outbound

This paper cites ProSoftArena: Benchmarking hierarchical capabilities of multimodal agents in professional software environments,.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision ProSoftArena: Benchmarking hierarchical capabilities of multimodal agents in professional software environments,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.129035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.129035Z digest=sha256:fb02efa010729f57841bca850b9f7e9b1fc72f2eb715af46df834903c6e0edc0

Observation c77656e4-d561-4551-af1c-1a4ce5632115 · outbound

This paper cites Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.202588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.202588Z digest=sha256:97b6481f9efc7ecf43254c0de6241f5caf1ba0d7c27c042b09da392665d7429f

Observation 1cda3a64-7c95-415c-9335-232e8daec7c4 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.302624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.302624Z digest=sha256:54ffdadb843e7eb4cf43b41bc86fbc0aec2230ec869f1860b2691e5686937c9f

Observation 831ec9eb-dd1d-4816-ab7f-e3b29418f223 · outbound

This paper cites ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.462120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.462120Z digest=sha256:d8f5ee641c3a1dfde0208b7f727f2467bbceb283d72e4f2f2203148ea817b029

Observation b2ea0bdf-45b3-4e53-9e15-d98befaa2c74 · outbound

This paper cites 7076–7087.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision 7076–7087

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:02.865413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:02.865413Z digest=sha256:755e3a2d72bc4e58e1d490272ee56f0f65001cac3898c1975c692f9f0bad63a6

Observation ce3acd5e-8496-42bb-9064-eeb2690e0555 · outbound

This paper cites Available: https://www.swebench.com/.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Available: https://www.swebench.com/

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:01.738273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:01.738273Z digest=sha256:3a7673d6c85be16997c88674955aa1e6d3c6403c49e3c3c7cb18f388aebf2d69

Pith citing papers

No inbound Pith citation observations are available.