Pith. sign in

Paper Citation Record · LEDGER

AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2401.13178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.13178 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:31.241598Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f76ff87a-3d72-4686-8967-5ffdc37f8e4c · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.581967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:b81dd7e4715546459f68809c7839a3a71cde666195e5a3886ecdfb27d0f074fe

Observation 8fb94b4b-3c84-4b6a-85c1-06bcb76127a7 · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:29:27.515329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:e51ef3e1fb31fba76a1613bf10f423cc9d1eb1e1bcf9d309fd0b02c3643e1e10

Observation f6202b31-7864-4934-af61-b7d2544c0570 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.416304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:0be7bb6706c99c4f8244613e62ca4c256c82c20348b48f8b1d802bc9d5d56c42

Observation 78642a46-cfb1-4bda-b33f-2afb0c9e06cf · inbound

Make Planning Research Rigorous Again! cites this paper.

Make Planning Research Rigorous Again! AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.241598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:32:31.241598Z digest=sha256:753d6fc1a25f544d28e43afb3d7f26e127d6714e8f84d435e5d393ecf1024b56

Observation 07df1bbd-0184-437d-ac61-5330f2269e25 · inbound

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback cites this paper.

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:55.654047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:55.654047Z digest=sha256:95cad88b1252490a6ef4d195c16b437da48c38014faa48db141e035f21bddcb2

Observation bc6f3ffc-5f94-4b77-aaed-b64b99c6b270 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.228552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:9be5b9cba9941e6c5cbbd606e6bccdd7c2e28d5ddf6d7cb2d6283cd7e5b8a4c6

Observation 581b98ec-4edd-4610-96ee-2b871f2fa734 · inbound

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems cites this paper.

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:59.063879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:59.063879Z digest=sha256:e181b23ac4cc43821ade7f49edfadbef8d353d2487449d7ae2ce5dcb96c74f03

Observation 3a1d57a7-a4b6-40c4-a890-fc0c4be9b11a · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.104695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.104695Z digest=sha256:7e448a155b67590f2e26c47b3958beb7c53398e631c2a2812f12ea0cb0b32e1b

Observation d775f98c-79ce-4fda-ab4b-d2fc7398bc4f · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.905361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.905361Z digest=sha256:9bab835836a34c6329826945c15bd74fad3faf942957400a1a9e6c61dd62a480

Observation 59d6be4b-64f8-49cf-977e-956a9ec79f4d · inbound

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models cites this paper.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.781635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.781635Z digest=sha256:05e7fb434b390741d15c71d26f3eec95ccb0c13bc891881eecfe8c05c43a6942

Observation ff4c10d5-7239-4858-9037-39477c88bd91 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.699335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.699335Z digest=sha256:9d50a652633fe83e6099c5de1629dfa9c39d88645790cccd3275a0bbdd89d943

Observation f1ba59cb-94fc-4e95-8896-a3cff2265162 · inbound

Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents cites this paper.

Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T01:01:38.735705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:01:38.735705Z digest=sha256:7eed115f9e1696f8a98e57d8500ca1455c23787ee84d05b52fe9ee52b80ecab4

Observation 14eba690-db7b-4d04-92b7-fd4fb6d5f73e · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.243421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:b4556dd0ac6488bb56749416820b64956079be14e0bbd16c11cade1cd567b79f

Observation e9607fd6-e224-478a-ba99-489caa8123e5 · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:18:20.359761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:cb5fb7cd2ea3002f6067fb8ffb1f0e18e030d9ae5239ee4302de93416f846e05

Observation 9a1aa186-e393-40e7-ac4f-7548b4f78508 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:57:24.553918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:dbffd5fa7f40b2b69737bcc1c7cd7f2721ff4adfa6be414ee11a05eec9a2877c

Observation aa8b3b2f-1d2b-428a-8fa9-a65bb2bb6c90 · inbound

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill cites this paper.

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:02.516415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:36:10.725278Z digest=sha256:10d0389153d39b4a500147036d9d623875870202119a274b6637816d486654e2

Observation 6045f6b9-eb30-42c8-a7bb-07f156e290bd · inbound

Holistic Evaluation and Failure Diagnosis of AI Agents cites this paper.

Holistic Evaluation and Failure Diagnosis of AI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.221202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:17:28.794622Z digest=sha256:83f8faf6ca3c3fa931e33c71a27360e6fc5e209d507fcc248178a8243c77213b

Observation 2aa9f43d-5d04-4293-b1fc-66dad36290ca · inbound

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design cites this paper.

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:23:03.208096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:22:40.544305Z digest=sha256:81211b8e645725cd334e7256b18969c8dc025e45e96066ac26f51d5776cb2fc5

Observation ddced076-1a45-4ff7-b2d4-29f0fe116742 · inbound

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design cites this paper.

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:34:59.702391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:24:57.376385Z digest=sha256:a1873fce21a76040eef26d139480bf34e4cbd64a3f8811e3e7510862898be675

Observation ba622d8f-9203-4871-876f-a11f9308063c · inbound

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents cites this paper.

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:44.106800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:34:43.686189Z digest=sha256:f95973e31fd3e3319a5f988287fc4b606bf83e7c2ed37f1cc9d4e58aacac388a

Observation 1508867e-de20-4be9-b4e3-fbfadd577f49 · inbound

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents cites this paper.

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.825258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T17:52:14.944875Z digest=sha256:7ad684c7a78b380ddef8827a12d8b61d5faf1d704832235e1c6e415651e55a28

Observation 6de4b676-0d25-434b-804d-9b7c2be9a0eb · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.094083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b40c23a09d55d420fcfd8dcd4fddb3babd0ce4aff1a7ee15e52191096eba308b

Observation 702924c7-5ae9-4c68-8ece-9d8b6d499abe · inbound

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows cites this paper.

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.284768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:42:20.527063Z digest=sha256:d513d3ac97b213c0df56fc49a162631e6a079bacf736b76860b73a42cd5b4432

Observation aefd2fa4-d369-4759-95c5-8aa39d835a58 · inbound

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents cites this paper.

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.634908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:28:08.382201Z digest=sha256:10cefc8c2f110d5fd15bc89e81f78d44b4212613151d460d598ede0b716e6e9b

Observation 929c8d43-2632-4205-9d91-610cbc9c1482 · inbound

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness cites this paper.

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.015798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:04:29.791558Z digest=sha256:b0eb367bede867d7d7296227b8084f624bf67eafcfdd23e701750995b1ac1774

Observation a38e9815-8bdd-43b8-be5b-62b0e1aa79ac · inbound

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development cites this paper.

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:41:32.038494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:41:32.038494Z digest=sha256:418ae6d08710f7b3c18d2962bb669e96eb7f5a3e94f52f12071dd78d7083d78a

Observation 285ed43e-09ab-44c5-bc7e-271c70d3b5ce · inbound

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows cites this paper.

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:35:43.053786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:35:43.053786Z digest=sha256:ee3768b18482cdc05ae79dd506aafbeeae7a0bb40801ef6188b1d3e533a2e8dc

Observation 078adbd5-018a-4b12-9862-500131aaa417 · inbound

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems cites this paper.

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T16:10:44.381736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T16:10:44.381736Z digest=sha256:d2f8e5b239a4a5c0ce3cf59d2dcbe3a4183fdaf9d4f653c0112286196860ddf8

Observation b53841d8-7de1-4d38-8c00-2d2f42968e7d · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.674463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.674463Z digest=sha256:20244a7ea53aba8f83882c691eee5aae8d3f42339c1311bc00eae70c7cb0ba2d

Observation ecde07a1-275c-4b25-ba4d-c7856806209c · inbound

Beyond Component Testing: Validating Agentic AI Systems cites this paper.

Beyond Component Testing: Validating Agentic AI Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T07:41:24.314016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:41:24.314016Z digest=sha256:9158ae0c5845642aab47606b449b6c77702d01bf883df0bb34341f7696c786fe

Observation b9eb8009-1da9-4cfd-b2d0-f72eff0b0898 · inbound

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning cites this paper.

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:44.867963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:16:44.867963Z digest=sha256:a556aed07f7d8f8af90b63dae286ed2d013c03b1a2e8f3148007b0647ba7e724