Pith. sign in

Paper Citation Record · LEDGER

An End-to-End Agent Auditing Engine

As of 11 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07346 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:52:48.176696Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5625cc56-5b20-443e-bde5-a3cc049f14d4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

An End-to-End Agent Auditing Engine Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.085855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.085855Z digest=sha256:222f0ea470421ff3db3dd157378e606bab1acbbd9c1e18c654b9f8ab701e62d3

Observation 1a54921d-cc84-44e7-8ee1-3c99ee4bfb3d · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

An End-to-End Agent Auditing Engine SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.098487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.098487Z digest=sha256:37bddcd400aa3b39e64e44bce6be633704d06af6796d2f1ea35c76423fef674b

Observation db198a5b-35bc-4d38-a853-3cf7de2b12b1 · outbound

This paper cites Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI.

An End-to-End Agent Auditing Engine Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.102698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.102698Z digest=sha256:e92e4687e226829420f1b3aa453451f639720328f93462123bedfefbc0a415f1

Observation 3eb138b2-b36f-4c69-8065-0659313dbba7 · outbound

This paper cites AgentQuest: A modular benchmark framework to measure progress and improve LLM agents.

An End-to-End Agent Auditing Engine AgentQuest: A modular benchmark framework to measure progress and improve LLM agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.539262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.105912Z digest=sha256:96d3bd0ebaeeae6392f5804ce8d6a933a43e70fa167d8f05ee2dc3accbafcf3d

Observation 65bdc42a-287d-4d87-bf08-93c141cd6b4b · outbound

This paper cites doi: 10.18653/v1/ 2024.naacl-demo.19.

An End-to-End Agent Auditing Engine doi: 10.18653/v1/ 2024.naacl-demo.19

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-10T05:52:48.108994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.108994Z digest=sha256:e1a89159fa956576968fb41d9f563eab57447e7eb8413c580702b5ad6be486b6

Observation 04d89948-296e-4634-a31b-577117c37761 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

An End-to-End Agent Auditing Engine Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.112150Z digest=sha256:c271203bcdb36f594c0dd964ca649ac5e1fb7cfa86692c2256f84193e5a208b3

Observation 408085aa-2f01-4960-bb53-84efac2b1c56 · outbound

This paper cites Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.

An End-to-End Agent Auditing Engine Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.122129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.122129Z digest=sha256:e0184c2bbd8fa2b3ba3ca1097896f4cb3b67dafe1365df9390042270c58db69f

Observation 3817d661-2696-4aea-93b4-b902cef122b0 · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html.

An End-to-End Agent Auditing Engine URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.128965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.128965Z digest=sha256:924cdefb7d3dd410c13c806f895d50a5d52195560f8119cce195af120b29543b

Observation d5c090e1-e6ea-44b7-b0d9-017520822c2a · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

An End-to-End Agent Auditing Engine Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.136857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.136857Z digest=sha256:59ffe94c4cbd722b34a97dc1cdef638870cdc75f4aed3e2483060fdc4387c669

Observation f3c75470-fc93-44e4-9e36-be13b64de80c · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

An End-to-End Agent Auditing Engine GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.144277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.144277Z digest=sha256:303c93dc83ef79c9cc4da247735e163d2a21418611b6781a8617e510c88f830a

Observation 139e6e5a-fff7-4f4d-87e6-4455a1230951 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

An End-to-End Agent Auditing Engine GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.147679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.147679Z digest=sha256:ac3a5b4e37b9710b6953f96f569a24d6785de61343983e69da137435dda73d2f

Observation 1b5f96d7-880a-4d91-b219-8ca16ce2a11f · outbound

This paper cites Challenging big-bench tasks and whether chain-of- thought can solve them.

An End-to-End Agent Auditing Engine Challenging big-bench tasks and whether chain-of- thought can solve them

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.528742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.154853Z digest=sha256:9cecf3a228bca5704da29824478b72306db66f02a1d794c749e63b86f3dddbba

Observation 11cf1d65-33b9-46f8-8219-5cb159b44821 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

An End-to-End Agent Auditing Engine Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.518772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.158585Z digest=sha256:98d50f380488c2c6f9e5d9c58e0c692711416115590396cd30a9e85e9970f6a0

Observation dc862431-79e6-4274-bd6b-2e62e90a1f64 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

An End-to-End Agent Auditing Engine AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.161998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.161998Z digest=sha256:bc40651e6f91926596959ce392f270541751e42e17044dd4950665d4a0717327

Observation 29cd5a68-1edd-45f6-b352-5db9ba220843 · outbound

This paper cites URL https: //aclanthology.org/2025.acl-long.1355/.

An End-to-End Agent Auditing Engine URL https: //aclanthology.org/2025.acl-long.1355/

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.165568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.165568Z digest=sha256:b6244a0d3dbdbe7922f409d151c299148ba16cdc621792a14f7a47e0c6eea481

Observation 4cd9036f-7a42-413c-be52-d7ab72ab06a0 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

An End-to-End Agent Auditing Engine $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.169156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.169156Z digest=sha256:09cb3d7e72ca7343785538c6a655baca787dfa4c5f1394e823268c1c8bcadc54

Observation 967a22cf-b289-4992-9ea8-4cdfa8037e74 · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

An End-to-End Agent Auditing Engine Agieval: A human-centric benchmark for evaluating foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.507977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.173232Z digest=sha256:249dc2c6f463ffdf57f3fbc628f2169a699eef9839727e7331db909e77d49117

Observation 1193309a-467b-4d65-9667-62c9f12175ac · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

An End-to-End Agent Auditing Engine Webarena: A realistic web environment for building autonomous agents

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.497017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.176696Z digest=sha256:ad6c1a3ede6e53deb85764c53480416e17c8578e95375f4dff50c6ff7a5acf20

Observation 539b6acb-6091-478a-b9be-ec9b3644bcb4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

An End-to-End Agent Auditing Engine Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.094149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.094149Z digest=sha256:f97b88c208e16fcd37f8bf74ff463aad08514bd0338109160b3b2b5dfbea4205

Observation 0348cd40-320d-423a-bdfc-bc61eea4c093 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

An End-to-End Agent Auditing Engine Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.115676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.115676Z digest=sha256:64ba67420137295f4381cb7f44b161c41fefb86ed1b72e6621d495339f05f65c

Observation fc4b48c1-a2ab-40ad-b5d2-933213c61d98 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

An End-to-End Agent Auditing Engine Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.089974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.089974Z digest=sha256:8c6552047c8c966f59b7ee735809f39a24aec61d77f7f0c8c0f1ff8102307d07

Observation ed8b3f05-daae-4ebf-92d7-bf465ad0e3ba · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

An End-to-End Agent Auditing Engine TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.125316Z digest=sha256:6cd059c4a1e10834311d6388c0fca204ee84b3590f795f14832d4bae8ad8b642

Observation aaeede7f-642d-4d82-8cfc-e039574f507f · outbound

This paper cites Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents.

An End-to-End Agent Auditing Engine Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-10T05:52:48.244282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.151220Z digest=sha256:11ce0e0591abf1d198c07b7dc1e0a507ae7ebfc4d19708eb6fd5514fbab54383

Observation a8fbcf01-2c69-4ddb-8a7a-dcadbdc527ce · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

An End-to-End Agent Auditing Engine $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.081468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.081468Z digest=sha256:7dbce83fb75f2a6b7f2a8215bdb9435fb83a7ec333e22f1d2596cc46289a3773

Observation f020150e-188f-4d8b-bf9e-c106e2db9075 · outbound

This paper cites General Agent Evaluation.

An End-to-End Agent Auditing Engine General Agent Evaluation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.077579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.077579Z digest=sha256:2e80033ecccbb33be28f64b1e23dd8410305f45b029e5a673a500f919022466a

Pith citing papers

No inbound Pith citation observations are available.