Pith. sign in

Paper Citation Record · LEDGER

An End-to-End Agent Auditing Engine

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07346 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:52:48.176696Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5625cc56-5b20-443e-bde5-a3cc049f14d4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

An End-to-End Agent Auditing Engine Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.085855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.085855Z digest=sha256:29ba48e92cecb76a83a2bb975abcd744bb58408cde26a64878458f6f29e33e28

Observation 1a54921d-cc84-44e7-8ee1-3c99ee4bfb3d · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

An End-to-End Agent Auditing Engine SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.098487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.098487Z digest=sha256:49bbdee4ea953424f0966c14f2a1e04406eec0ae6499100bf9b2df11cb39f799

Observation db198a5b-35bc-4d38-a853-3cf7de2b12b1 · outbound

This paper cites Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI.

An End-to-End Agent Auditing Engine Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.102698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.102698Z digest=sha256:42ee6c64fa6e88bc864809e2d2a0adbf357213be58fea09073e34d6e302ae0b4

Observation 3eb138b2-b36f-4c69-8065-0659313dbba7 · outbound

This paper cites AgentQuest: A modular benchmark framework to measure progress and improve LLM agents.

An End-to-End Agent Auditing Engine AgentQuest: A modular benchmark framework to measure progress and improve LLM agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.539262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.105912Z digest=sha256:1b8ca0363a16f38bf27268455b6cb2540da78d2041268d048731bc4b722076b9

Observation 65bdc42a-287d-4d87-bf08-93c141cd6b4b · outbound

This paper cites doi: 10.18653/v1/ 2024.naacl-demo.19.

An End-to-End Agent Auditing Engine doi: 10.18653/v1/ 2024.naacl-demo.19

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-10T05:52:48.108994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.108994Z digest=sha256:8816e995880e0b6de9152215c674312786ee8350fbb8808b9cde19aa254d34cf

Observation 04d89948-296e-4634-a31b-577117c37761 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

An End-to-End Agent Auditing Engine Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.112150Z digest=sha256:3f1b071ff868c56ad40868c0504220cba96af80890348dd19f3b365f681585c2

Observation 408085aa-2f01-4960-bb53-84efac2b1c56 · outbound

This paper cites Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.

An End-to-End Agent Auditing Engine Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.122129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.122129Z digest=sha256:39c21608d14658e0d1c2274a29d15bd9a5f831637da67ae4de093c3b30f9854e

Observation 3817d661-2696-4aea-93b4-b902cef122b0 · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html.

An End-to-End Agent Auditing Engine URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.128965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.128965Z digest=sha256:7329bcdfcb90b9c3ea63a92fa7ab0115a558f53cbbb80fb15b49047ec42b527a

Observation d5c090e1-e6ea-44b7-b0d9-017520822c2a · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

An End-to-End Agent Auditing Engine Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.136857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.136857Z digest=sha256:6f05e3b3fb3ff6ddabaa18a1e6087fdc3c0e924a7a93cb49e9260551dc3181ec

Observation f3c75470-fc93-44e4-9e36-be13b64de80c · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

An End-to-End Agent Auditing Engine GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.144277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.144277Z digest=sha256:1d79c1941e5cfea58ff47a2ffb95a4972f8d3c43933871047f856caee63d15b7

Observation 139e6e5a-fff7-4f4d-87e6-4455a1230951 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

An End-to-End Agent Auditing Engine GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.147679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.147679Z digest=sha256:f1a6e647f45811a35015b58cdb739c23b998f104c5cfa63614eb6885538966fe

Observation 1b5f96d7-880a-4d91-b219-8ca16ce2a11f · outbound

This paper cites Challenging big-bench tasks and whether chain-of- thought can solve them.

An End-to-End Agent Auditing Engine Challenging big-bench tasks and whether chain-of- thought can solve them

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.528742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.154853Z digest=sha256:5e83060c19bc958538b614542150fd5e3ce2a032422cdea3b8f0f50f76918686

Observation 11cf1d65-33b9-46f8-8219-5cb159b44821 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

An End-to-End Agent Auditing Engine Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.518772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.158585Z digest=sha256:3fb712e0d4a0c973badd0703d8f2d1a934ed51f9474401000a4ce549b99054ea

Observation dc862431-79e6-4274-bd6b-2e62e90a1f64 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

An End-to-End Agent Auditing Engine AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.161998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.161998Z digest=sha256:f21416b28025219d674662a31fb289ce21266d4fd3625a8ae70121732fa8d7a5

Observation 29cd5a68-1edd-45f6-b352-5db9ba220843 · outbound

This paper cites URL https: //aclanthology.org/2025.acl-long.1355/.

An End-to-End Agent Auditing Engine URL https: //aclanthology.org/2025.acl-long.1355/

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.165568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.165568Z digest=sha256:4b72c21a532c2c7458d6d76dd6f9e9cb34b8aa7554cf3a69c8e668fa256ce621

Observation 4cd9036f-7a42-413c-be52-d7ab72ab06a0 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

An End-to-End Agent Auditing Engine $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.169156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.169156Z digest=sha256:b4b12c0096de0ad54033214a03f6937a249716913eaa96f89d961830f89df655

Observation 967a22cf-b289-4992-9ea8-4cdfa8037e74 · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

An End-to-End Agent Auditing Engine Agieval: A human-centric benchmark for evaluating foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.507977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.173232Z digest=sha256:8404d4bf9a7b168d4d7eb80dbd3eceef14941ab88dd133aee6fe0f12ed147cc9

Observation 1193309a-467b-4d65-9667-62c9f12175ac · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

An End-to-End Agent Auditing Engine Webarena: A realistic web environment for building autonomous agents

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.497017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.176696Z digest=sha256:7b0bfdbaaaffa6ab164ab93b991612b6177ea27e9bad173be3c4b7fc0feb9ec9

Observation 539b6acb-6091-478a-b9be-ec9b3644bcb4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

An End-to-End Agent Auditing Engine Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.094149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.094149Z digest=sha256:e5b7ad966bdbaebe1f18fd94d7f8374a8794b1b2f8737090854870675b2a450d

Observation 0348cd40-320d-423a-bdfc-bc61eea4c093 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

An End-to-End Agent Auditing Engine Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.115676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.115676Z digest=sha256:5c55b7854a8e65e8b098678beb2e5d9185180178f63b3938d3b0ddd9f5f241ff

Observation fc4b48c1-a2ab-40ad-b5d2-933213c61d98 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

An End-to-End Agent Auditing Engine Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.089974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.089974Z digest=sha256:9800833aa2ca571435c516c659d0a3d025004f65fe729e2944912136fc1e6380

Observation ed8b3f05-daae-4ebf-92d7-bf465ad0e3ba · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

An End-to-End Agent Auditing Engine TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.125316Z digest=sha256:98e6494561dbbed9d47889abe2821d39c787de0dc3b03e0ee0d57e6317b114ae

Observation aaeede7f-642d-4d82-8cfc-e039574f507f · outbound

This paper cites Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents.

An End-to-End Agent Auditing Engine Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-10T05:52:48.244282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.151220Z digest=sha256:0a75987d47f96d237a21fe8e0f12daa1d13cd52ba4dbd37b1bec2ca615d35dae

Observation a8fbcf01-2c69-4ddb-8a7a-dcadbdc527ce · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

An End-to-End Agent Auditing Engine $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.081468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.081468Z digest=sha256:4a45f793c7dfe6b26d751c68a5b62db3d2057e8107a1735d3afbadb6f421c353

Observation f020150e-188f-4d8b-bf9e-c106e2db9075 · outbound

This paper cites General Agent Evaluation.

An End-to-End Agent Auditing Engine General Agent Evaluation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.077579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.077579Z digest=sha256:1e5628527bcb241ba37723486c09c431914e11deb4bfcc3c976cbbbef9c4cb8b

Pith citing papers

No inbound Pith citation observations are available.