Pith. sign in

Paper Citation Record · LEDGER

$A^2E$ : An End-to-End Agent Auditing Engine

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07346 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:33.473138Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6fe09d6a-105d-4c2e-bfa8-75abf599e15d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

$A^2E$ : An End-to-End Agent Auditing Engine Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.364048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.364048Z digest=sha256:2b77f0c649336c355db29983229b34f6b50ec010df7c6fea1cfb83efb438cf55

Observation 85a655a0-27ff-43a3-b41b-e5aa50d2e13d · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

$A^2E$ : An End-to-End Agent Auditing Engine SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.378120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.378120Z digest=sha256:db6ab5a77bc2bd2ed35db852b91f174d369f6a272bb781a0f1e6718a7c22376a

Observation e6711cd4-f1d9-46dc-a6ed-309f45919e40 · outbound

This paper cites Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI.

$A^2E$ : An End-to-End Agent Auditing Engine Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.382742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.382742Z digest=sha256:d988c6bf02fa602fab4fc1fc981ade232254eec660dc6a6ca5c5b467742f98f1

Observation f996032c-bcb6-4b08-a30a-74b0902148f9 · outbound

This paper cites AgentQuest: A modular benchmark framework to measure progress and improve LLM agents.

$A^2E$ : An End-to-End Agent Auditing Engine AgentQuest: A modular benchmark framework to measure progress and improve LLM agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.387042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.387042Z digest=sha256:985c04a41f590bf8f20ad663a39cda35aed4bbf1989b1aee13e166fa175403bd

Observation a4cbc644-2304-4364-adaa-14a7c54358d8 · outbound

This paper cites doi: 10.18653/v1/ 2024.naacl-demo.19.

$A^2E$ : An End-to-End Agent Auditing Engine doi: 10.18653/v1/ 2024.naacl-demo.19

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:17:33.391178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.391178Z digest=sha256:c57c29eb9a98745ae1a4e7228ef7e677fe5763432a0d95cdccb4ab189235baa6

Observation c4a28fc7-d2f8-46bc-ac35-0cef38a36aaa · outbound

This paper cites Measuring Massive Multitask Language Understanding.

$A^2E$ : An End-to-End Agent Auditing Engine Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.396604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.396604Z digest=sha256:8b2520408577d16dc2db5054b6b28ae78ac25b72c9e376d94a4ff09439fd6a16

Observation a2c4bcd3-83f1-4f2e-a3e5-40bdbe1aa901 · outbound

This paper cites Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.

$A^2E$ : An End-to-End Agent Auditing Engine Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.410111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.410111Z digest=sha256:d3b6357c3051995a80be6938aa2875e6128f8d3587880dc0352018c7b4213b84

Observation 512a202a-feb8-4092-ba15-74460ccc9ec9 · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html.

$A^2E$ : An End-to-End Agent Auditing Engine URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.419003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.419003Z digest=sha256:b019e1349ce189c738ad56c26da1ba2850bc0b978c3aae26de2825ea58b80a05

Observation 98d1442c-0757-4d19-a99e-2f968ece2607 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

$A^2E$ : An End-to-End Agent Auditing Engine Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.427518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.427518Z digest=sha256:a874b6e64225f0779691cfdb325d611a917aa04c0e29c67e3f30eae633a6717c

Observation 8b0c0ebd-ed70-43cb-9e45-bb396e22b387 · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

$A^2E$ : An End-to-End Agent Auditing Engine GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.436059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.436059Z digest=sha256:63072c03e40c0719a6630f1115cf72565817924a91cb9fb32f58b9a3bb295168

Observation ac6c6bc6-ca76-4e15-bb68-576ff13e1a30 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

$A^2E$ : An End-to-End Agent Auditing Engine GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.439391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.439391Z digest=sha256:f97bdb9f2500a56a842d1b7da3531a217f3eb2c909dafaee6f0e659ae21d692b

Observation 37be3db4-e960-4275-9864-699086af7118 · outbound

This paper cites Challenging big-bench tasks and whether chain-of- thought can solve them.

$A^2E$ : An End-to-End Agent Auditing Engine Challenging big-bench tasks and whether chain-of- thought can solve them

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.447362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.447362Z digest=sha256:a6acda4c01ebd93f4f9f01b84d1390b96bf6c9b64f11c2c2901af868d1e0d866

Observation f0c1eef5-2f85-49d6-8dc4-feea77916cff · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

$A^2E$ : An End-to-End Agent Auditing Engine Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:33.855686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:33.451697Z digest=sha256:0498b57da81777c3f866c9e76399bbd6c96c8f93d0d6bb064067b8c514bbda79

Observation 4ffeb436-0111-4cb0-82b1-bac09091c6ad · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

$A^2E$ : An End-to-End Agent Auditing Engine AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.455842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.455842Z digest=sha256:c5c7d00321c9151f76f87cdac24d17f82be1e97728bde98e6ed65179df097f4f

Observation bdb53c5e-d863-49c1-af9d-57a763577a09 · outbound

This paper cites URL https: //aclanthology.org/2025.acl-long.1355/.

$A^2E$ : An End-to-End Agent Auditing Engine URL https: //aclanthology.org/2025.acl-long.1355/

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.460021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.460021Z digest=sha256:315e2cea49c7eda123b13986b501a29d25a69502af41c0bb7b5db82027ae466f

Observation 3e200386-5886-46de-b720-1f9424d4d30a · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

$A^2E$ : An End-to-End Agent Auditing Engine $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.464279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.464279Z digest=sha256:4d8df62fbcf42640d26724d3e840daa61d0495c8b82951b00f3290c0c9d0f0c6

Observation 2506623f-d569-485a-be0b-65c7e4978297 · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

$A^2E$ : An End-to-End Agent Auditing Engine Agieval: A human-centric benchmark for evaluating foundation models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.468800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.468800Z digest=sha256:0d52fae7b3cc7832f883a3547abd09428af3439df344f7842224cac86f3f08ce

Observation ea7a3de9-cef9-4144-a134-ac56f1cae46a · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

$A^2E$ : An End-to-End Agent Auditing Engine Webarena: A realistic web environment for building autonomous agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.473138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.473138Z digest=sha256:3ea7222ee6e8a8632e1a704f87dc0a795a5fc52747ca9382c2b150dfa7a4db9e

Observation 64f5c445-0214-420d-8cf9-1135ee61d448 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

$A^2E$ : An End-to-End Agent Auditing Engine Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.373391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.373391Z digest=sha256:a86579c39ebf55ce01b1340344f383592201231911187ad4a2e06e2ffb2a18ac

Observation 1bd01000-952b-4576-b0f2-9d6ec3671495 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

$A^2E$ : An End-to-End Agent Auditing Engine Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.401885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.401885Z digest=sha256:957076991c90f7a772a6505488850ee435ddc7f7ab20ee67069a50496b19a204

Observation 55335bf2-2044-4324-8be3-91ddff90797e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

$A^2E$ : An End-to-End Agent Auditing Engine Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.368857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.368857Z digest=sha256:234786c0bcb08dd4ae400bb2eda4c42954779cd7ddf125c8c8f49ea9c70eb70a

Observation 3e8bc028-a3c5-4c36-851a-ed1ca8373b0c · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

$A^2E$ : An End-to-End Agent Auditing Engine TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.414314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.414314Z digest=sha256:f00d1a50e5e789bde2886a112641be41b6a8db97b8a44104ccfbaebccb7438f7

Observation 1c1740ba-fc59-48e7-98d1-19196a995173 · outbound

This paper cites Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents.

$A^2E$ : An End-to-End Agent Auditing Engine Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.443440Z digest=sha256:01b4873259206c1553304b3a42ad344d4789e072b8c30ca5ff5d56d58de68b1e

Observation 54d93f2c-240b-424d-85ab-f2ca83b10900 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

$A^2E$ : An End-to-End Agent Auditing Engine $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.358874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.358874Z digest=sha256:eaa94db952b4ffe598596da028bc5a9ffdbeffe350a5a46d01b921ecc372e9e0

Observation b3c35b48-fa35-4cad-8a83-57672f87bca2 · outbound

This paper cites General Agent Evaluation.

$A^2E$ : An End-to-End Agent Auditing Engine General Agent Evaluation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:33.355201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:33.355201Z digest=sha256:42b701df4e6647e3bc94ce6df2694315ff4b76c044ba52617036dd9ccefc65dc

Pith citing papers

No inbound Pith citation observations are available.