Pith. sign in

Paper Citation Record · LEDGER

Evaluation and Benchmarking of LLM Agents: A Survey

As of 7 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 6 inbound Pith citation observations for arXiv:2507.21504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21504 v1

Coverage vector

measured 100 of 143 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.813422Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:59:52.218232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:27:56.022361Z

Reference resolution

100 of 143 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01cdcaa9-3927-49bf-beb2-97395950eaba · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.482897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.482897Z digest=sha256:03ed53eb7c3fc0f045aff175edbe91a7ec129461ae5d14bc89ef99e0294cea02

Observation 99349306-a237-44b4-b97b-f48a3ee2428a · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.486623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.486623Z digest=sha256:93799b6508954e8eb5ef3a4d52a7727dc75171d8dc6a1ce3096fe3cfc18ea561

Observation d6b75862-6f1c-4a82-b403-cf6c41ea3fb6 · outbound

This paper cites 2024.Inspect AI: Framework for Large Language Model Evaluations.

Evaluation and Benchmarking of LLM Agents: A Survey 2024.Inspect AI: Framework for Large Language Model Evaluations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.489909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.489909Z digest=sha256:e71ef68f16876445fec39291b517056d0976a3dd47ae074cf1d491d26b2a76b2

Observation 3acd0a5f-50e6-42ab-b9c6-5501b546d41d · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.493073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.493073Z digest=sha256:907ff777efc63edbf0e23d47d32785f167f473bae9a32e6efc40845ac581189e

Observation a0183b9e-8c90-4431-93d2-6f8bc626cc76 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.495973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.495973Z digest=sha256:1b2489f88beab63f23b67bb562bf66e2b17855375e98e9abfe1332e50677f418

Observation 3e084299-b3f2-4353-8702-1ac4c1ff8a31 · outbound

This paper cites Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs.

Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.502718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.502718Z digest=sha256:b0571ee46f1c10204c7c3ce7a36c63cbbbcedabdca25b02c23f8cc33683c4f73

Observation f046b40a-e071-4cab-b409-b5d75cd66385 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Evaluation and Benchmarking of LLM Agents: A Survey ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.506012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.506012Z digest=sha256:9704c9484a097449f9260e79834b906558d15253b6107bfdd5a88b26fdc254a8

Observation 6b32d450-e913-47cb-8d89-384fcf71762e · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.509324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.509324Z digest=sha256:03d707387b47cb9f17c787c5beff09501340efc1e6b04566b0ebb53fd7827db3

Observation e0ce98c8-6a51-450f-a48f-dfd9539403e7 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.512128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.512128Z digest=sha256:d8aaf6ee78fec11fcd914f0acf96b8222c2ead7be40a3b037363b2b5f8c38e18

Observation a5fc4b80-d188-49ad-920f-98e5b6681c0a · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.517960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.517960Z digest=sha256:edd7d379572069b07da57b2006dede6e3a85a0bc386aafed3de2d8a58ea06938

Observation 2094d729-a247-4398-969c-1bfd1da0a94a · outbound

This paper cites ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery.

Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.520560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.520560Z digest=sha256:a5a1366e08be6e705130c91e64b8fc86fe10f9e80448aca9692dab906a12c9b8

Observation 724229a0-0d2e-477b-bc0a-f9fcc9135313 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.523530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.523530Z digest=sha256:e1a1f5d5581e806431eead9471f94028e3aebfcf0e6e68b74fb09c3a0bbe4563

Observation 33451500-2922-4100-8d9f-25563c283b73 · outbound

This paper cites AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.

Evaluation and Benchmarking of LLM Agents: A Survey AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.530025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.530025Z digest=sha256:19cfc408dbd77538640c1fe97080d00446ef9f4ac5142dfa420853a359252e41

Observation ae999c9c-d449-48de-a343-0099d907a1c3 · outbound

This paper cites The BrowserGym Ecosystem for Web Agent Research.

Evaluation and Benchmarking of LLM Agents: A Survey The BrowserGym Ecosystem for Web Agent Research

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.532807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.532807Z digest=sha256:4f52c3a7825e3507a26f15f166e219108e5098ae4c37d76b32f359f323caf0cd

Observation af3d1478-45bd-4d65-b7e9-7e7899f8da05 · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.

Evaluation and Benchmarking of LLM Agents: A Survey T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.526623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.526623Z digest=sha256:062dcc16a4aeed9c32fec4d8ffe69dfbb0f20d0c37abfa97e74dc6c496dc81ea

Observation e3010720-da73-41a1-9bf6-9021a3d30020 · outbound

This paper cites Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI.

Evaluation and Benchmarking of LLM Agents: A Survey Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.538622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.538622Z digest=sha256:0893ab5c21ebfd9b51028227e631821f66bf8e1c7d3b6a25f5e8e6c2b57ce726

Observation 25d1dbed-283a-4454-84b1-19d51b7eb010 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.541517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.541517Z digest=sha256:05f165f1792e49b610fbb977edc32a622a8b596fd178501db4584ee040f2c0b7

Observation fd5e71eb-3fd5-4fc8-bee5-f08a2d330abe · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.535787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.535787Z digest=sha256:50f084af9f4b1eecaeeddd19994b4b55d7a12f1668432646ee251d7a796d4b5e

Observation f830591c-6c6c-4755-b8b5-a1772cfe8ebc · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.547500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.547500Z digest=sha256:b0c1d0d6f362488ba37f77b3291e77faa8e8f91fc2a1b8742d1ebee194503128

Observation 93374fe1-760a-482a-98e8-9082b5f5cfd0 · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

Evaluation and Benchmarking of LLM Agents: A Survey Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.553434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.553434Z digest=sha256:e5dca8454deb6400c9e5f17089a42ba042379107b8a7189b5c868aec6acd0508

Observation 33514e0b-4b21-431b-b16a-68848f3dad5e · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.544451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.544451Z digest=sha256:0446aa61e867ac77ed754abfb98f8e13569761cff224224e9100b11dc1a574da

Observation 8c2992d1-6232-476b-b113-711a9cbf67ba · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

Evaluation and Benchmarking of LLM Agents: A Survey LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.562169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.562169Z digest=sha256:498cc970e6ca015788bbeb0488ef4395442325fd1bc79dceccf7954dc4bd9fe8

Observation dac87e5d-7d37-4502-a571-c8eec1d99f71 · outbound

This paper cites Re-ReST: Reflection-Reinforced Self-Training for Language Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Re-ReST: Reflection-Reinforced Self-Training for Language Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.550610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.550610Z digest=sha256:5e48711470913bf3b87e30e0844d7cdc081aa1d9bd64e850550783fb520fbb80

Observation 8f11e5aa-5671-45f3-8e40-c48c16c85aec · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.568290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.568290Z digest=sha256:4a18cbb9b16bdabb9b5fbdfdbe55957af0a76fe2026ac31930eaf7e81ef8ac8f

Observation 5f71b096-20f3-4f48-95de-22c4c408171b · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.556360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.556360Z digest=sha256:2eac565f8710adf983c40faea32a01c2850497b7c9ef4f050de9cd73530e037f

Observation 0dd12506-e3ec-4b8b-bf1a-21442e31dc7d · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

Evaluation and Benchmarking of LLM Agents: A Survey Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.559149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.559149Z digest=sha256:5fed02925900320b68a8e44059f6765d83ecea36b3a5f2e591504ddbd5171052

Observation 98fbabe5-a8ae-4a36-b3a1-d6b09064913c · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Evaluation and Benchmarking of LLM Agents: A Survey RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.577059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.577059Z digest=sha256:33dd08247efda600f635e518450fed5d8612983a981eed470f52a16d4d8decea

Observation 11fba66a-4d87-4501-bf86-cb4f7ade44bc · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

Evaluation and Benchmarking of LLM Agents: A Survey Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.565142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.565142Z digest=sha256:135f32f9a8ab6f279dace16653d1525d5181d8b2eb2994fb0e11a1d8068f436d

Observation f4164373-270c-472f-a5c0-c1411bfdc6db · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.583195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.583195Z digest=sha256:f42b0d29f1b0aaa33d1103bc19405d08b7dbd7442b0f03606f869c802ccece38

Observation 45062438-4da3-4cff-adb6-6ce4ec281027 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.570992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.570992Z digest=sha256:77c8ffdd7380580a1c22f0d4701e78c3680d10094d2768ef801682dee25f0e17

Observation 03ed874f-470c-43f2-a573-5a733ab8619a · outbound

This paper cites Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.574030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.574030Z digest=sha256:73bed710ec834a13beb5772a7593a9e15f4561638c5196e73349cff1569a5eea

Observation 3be0e896-3a10-443f-9f86-725f2db6b00d · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Evaluation and Benchmarking of LLM Agents: A Survey Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.591815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.591815Z digest=sha256:2ae6a68556c633c227deffa0d74e8a73f75d7b80225ebc36e6ded580b711edfc

Observation c0640dad-e5f3-42c0-9f55-e0a2bed5c23e · outbound

This paper cites AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.580085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.580085Z digest=sha256:ad15c54faa18c08750796a8dc0003f518e56dcb93a6047582cfa3f40bb1fbdac

Observation d720d498-29d4-46c4-8d38-e7aed3ab185d · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.597845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.597845Z digest=sha256:263815223c6a045393fd509bad14b0bdb8c77c94b9390a8ee67eaa9e0a63d055

Observation 67888622-73c4-45c2-9b2e-e5bc28d3bc5c · outbound

This paper cites A Survey on LLM-as-a-Judge.

Evaluation and Benchmarking of LLM Agents: A Survey A Survey on LLM-as-a-Judge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.585641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.585641Z digest=sha256:a5e7a2566896e06918360d1647663ae38aae2093daa1b2117a601ee39a3e7e9f

Observation 57d265d4-3b1b-4bb5-b9cd-4067450ccb32 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.588784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.588784Z digest=sha256:fb6d2674e4e3bf23e17f25fd0bd5f67f928ce725cfd974b085d6892dd8869c45

Observation 03971990-d79d-4f98-931d-b8d22cc98ff8 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

Evaluation and Benchmarking of LLM Agents: A Survey Understanding the planning of LLM agents: A survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.606949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.606949Z digest=sha256:acbbad4167bc162a27d62352af230d1df473a4550b2a3058ef0aa6f54d7cd540

Observation 1cf4dcc5-30b3-49f2-955f-a82ac905765a · outbound

This paper cites LLM Multi-Agent Systems: Challenges and Open Problems.

Evaluation and Benchmarking of LLM Agents: A Survey LLM Multi-Agent Systems: Challenges and Open Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.594720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.594720Z digest=sha256:16e3088b75854df88fe6441ceac582df42a9c3df99aea1faa92ae930954b5800

Observation a47616f2-fdfc-43ee-8e73-50016babc9c6 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.612824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.612824Z digest=sha256:007fd254a6b787550a5bf5b347756368ec741fd352218e99881fe8af630a47fc

Observation 7e4ffcf8-ac71-47ad-a8f1-b9f26e1d22c1 · outbound

This paper cites AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation.

Evaluation and Benchmarking of LLM Agents: A Survey AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.600583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.600583Z digest=sha256:961f6fa0ff76f79d5b144498f88c1476677f6a21928220b2fd8bb83996857510

Observation c0f909b1-34af-4a68-bfb1-59360d7225e9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Evaluation and Benchmarking of LLM Agents: A Survey Measuring Massive Multitask Language Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.603848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.603848Z digest=sha256:70b52c8090f203c86c01b7a1605b17bbfdc5b66ff130dd75457397135de2e0c0

Observation 38e49dfc-b559-4e10-8a8e-db2336f3cace · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.624201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.624201Z digest=sha256:bcf88d71055a99076c6e0bbcfd56df1d817c9b7fd9146d5b48f0fb61d8f91e88

Observation 1576480d-e593-4139-8f1a-353f60d92f52 · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.609937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.609937Z digest=sha256:9d9ae8d170b184a71cf07e9170a53327981385d7331bf89c635097f5325f7d35

Observation f5195e60-8380-49c0-be77-aced989195e1 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.633063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.633063Z digest=sha256:b45df83be45e75ed597371a4bd5efc8431242a43edfd62a8e88d7cceb3c78691

Observation 17afe007-792d-4036-81d8-4f6e86d4e29c · outbound

This paper cites LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments.

Evaluation and Benchmarking of LLM Agents: A Survey LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:44:22.818063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:44:21.615493Z digest=sha256:115ec8c6efdd0e43040023e3faf24277650866e60f25544c703208f0d913de9e

Observation 6fbcd1ef-99ba-44ae-9490-f3f8136fb2d4 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Evaluation and Benchmarking of LLM Agents: A Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.618485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.618485Z digest=sha256:165ac8a620c62255217d7a8958dbac076964926c9890ba5b11d65d1ee8ddbb02

Observation b75ef5d4-56c0-4b5e-a026-a442f6ea8500 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.621618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.621618Z digest=sha256:3bf4236de75408dfd665d5a8641d0b47b2bbc2af46eb02256608d78e855c6c43

Observation fae8e12c-84d0-4d3a-8cf8-5b793eede1c1 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.645335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.645335Z digest=sha256:210a2eacd1f53a27a3e6c753041bd3baf8e71f4f8b425c03ea56ca36baecc19b

Observation 95dbed4c-c632-46b2-847c-0137bac07bca · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Evaluation and Benchmarking of LLM Agents: A Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.627127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.627127Z digest=sha256:98cb20680fee7bd513bd323d33ac0b968878da417cad26d227803acd49c3614c

Observation 08dd9ff0-baad-4333-b7dc-9e5ef96fc258 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

Evaluation and Benchmarking of LLM Agents: A Survey LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.630164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.630164Z digest=sha256:f1e42fc82dab529858f75a978790ada12baf3a12664ca79788849bc44ccda345

Observation 988d73f9-f5c6-475f-830e-f6a5bc06d29f · outbound

This paper cites Manning, Christopher Ré, Diana Acosta-Navas, Drew A.

Evaluation and Benchmarking of LLM Agents: A Survey Manning, Christopher Ré, Diana Acosta-Navas, Drew A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.653693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.653693Z digest=sha256:549242f5968ecdb3d2c2d0d6a8f26f0f9121b6b73e7d7423384e880cc4118f0b

Observation a7a1f55d-5608-463f-ba76-40eaeb9259b8 · outbound

This paper cites IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems.

Evaluation and Benchmarking of LLM Agents: A Survey IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.636050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.636050Z digest=sha256:8802f681055fff3d8d415cc12f43fe09d9480d10f48da3fdf9ee729376efe07e

Observation 1ea9c14a-43b7-48a2-9a75-5b196197df94 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Evaluation and Benchmarking of LLM Agents: A Survey LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.639176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.639176Z digest=sha256:69644fd20bb241bd4438b392768d421b437f773344f592fb77a511d33cf20eb3

Observation 4f5134ab-59be-4e29-8c6c-d1a72f2f2df9 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Evaluation and Benchmarking of LLM Agents: A Survey API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.642196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.642196Z digest=sha256:a893daf8fba4f26827e237660a6d614dd10b92664b6ed05f6b6258967ced9d8f

Observation 32246f49-b35c-4ff5-9d98-25867ca6336b · outbound

This paper cites Autonomous Agents for Collaborative Task under Information Asymmetry.

Evaluation and Benchmarking of LLM Agents: A Survey Autonomous Agents for Collaborative Task under Information Asymmetry

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:44:22.763154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:44:21.671369Z digest=sha256:196a312973df396f54511deb561b426e13774ffb303352658ab1796427d83c87

Observation a16e74b8-bc20-4e6f-afc8-922a972f43e3 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

Evaluation and Benchmarking of LLM Agents: A Survey Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.648044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.648044Z digest=sha256:7834fe44cfd7539431e1a3c2182f777114231b6c9ddd89c8663da8671d308a77

Observation 87cae1f6-0021-42e7-abc1-8c40a2a3919c · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.651084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.651084Z digest=sha256:d6e0cc45bae374246fbf045858d965135ddbf88637b068dc8d62a6eeec6ffb07

Observation 5d0eda4b-5def-4aba-aad3-b03c122d49fe · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.679806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.679806Z digest=sha256:62357365586ea4b2e19081e992b84e65ec31d7760358355e3186b6a38ad0e3a3

Observation 4d83681e-34b9-4e2d-a9f5-7ea8018dcf15 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.682381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.682381Z digest=sha256:b888a75e740603a2132b13c82b83842191fb07b86c85e52b6efecc3da3523326

Observation 10e97fee-f5b5-4387-a5b8-1543f55ffa76 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.659528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.659528Z digest=sha256:65e733ea0a26e867fd561c211d50ff80e8b49e424bebdb64784371f42fa7979e

Observation acb862c6-b9b7-40ed-b0b5-1650e4f96031 · outbound

This paper cites AgentSims: An Open-Source Sandbox for Large Language Model Evaluation.

Evaluation and Benchmarking of LLM Agents: A Survey AgentSims: An Open-Source Sandbox for Large Language Model Evaluation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.662265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.662265Z digest=sha256:3163ef2e86ec57e574d665908ba331c56a707a3731c573190568cc1264a15c14

Observation a71eb74e-e7d4-4f48-a9fc-14f70eaf0f26 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 62

Resolution
verified exact
doi, observed 2026-08-06T12:44:22.136258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:44:21.665219Z digest=sha256:22bd9199db35ff3b53571d07901c98f4d340ad1c64e5c7c774ff6d63504043b0

Observation b4988fd1-00f9-411c-a20b-cd5f66f6298d · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Evaluation and Benchmarking of LLM Agents: A Survey Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.668218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.668218Z digest=sha256:d9f2ac510f8388f5cb5f501df6280f1551439389e52974e035cd6ccd40b43d7d

Observation ff4c10d5-7239-4858-9037-39477c88bd91 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.699335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.699335Z digest=sha256:9d50a652633fe83e6099c5de1629dfa9c39d88645790cccd3275a0bbdd89d943

Observation 4f007bf9-f524-4fe6-bddb-4a592f38fa06 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.674191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.674191Z digest=sha256:fd5bbf8a7e6894aa422067e8e70f7cf57e29f24212c519c3890ddc038bd6981f

Observation 8693fcb6-3f7d-4d62-ba10-250858c969b7 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.676884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.676884Z digest=sha256:a73d6f1e727b65288cf07d5d2224fba5b31d44ba8235dbaaf4594628b72d6ec7

Observation 5d6b6591-30e0-40d4-a330-824cbe460198 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.707889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.707889Z digest=sha256:f0cd1aaf0bab6889de9cf9af18d336926f1de6e3c02a2f4fa4d60db713e8c6e0

Observation a3bb5815-e184-4050-875e-aff77b02f5b8 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.710615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.710615Z digest=sha256:3c228e5d6d38b980bdaff8a299ff7cef921e70adf26d00cd56a26fab11bf9a8b

Observation e4833b80-cd87-46f3-804a-6d73e1bbc477 · outbound

This paper cites Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing.

Evaluation and Benchmarking of LLM Agents: A Survey Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.685043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.685043Z digest=sha256:4d731bbd7b822d9c2eca770b392f9636f364fdb09c67c9c2b254d4a9bc3e1eef

Observation d7f7fd7c-d72d-4031-b562-f6a09f14307b · outbound

This paper cites AAAR-1.0: Assessing AI's Potential to Assist Research.

Evaluation and Benchmarking of LLM Agents: A Survey AAAR-1.0: Assessing AI's Potential to Assist Research

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.687855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.687855Z digest=sha256:775b8e8d47f94b11b1248dc9b0c018d47310c6955f7fbe33544861c49aaf5d91

Observation 332611bb-3bbf-433e-ab58-d8aa519a7dda · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.691267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.691267Z digest=sha256:e035428c065a97ca9a500b30e27d850b3b06bdd57fd9c498f7e38ceabc8db392

Observation 51527181-da25-4aff-a12a-678c0932dc05 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Evaluation and Benchmarking of LLM Agents: A Survey The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.693952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.693952Z digest=sha256:aa004520d73597f749259d4f968c8de304fef002fcb4279375da12927699a0b7

Observation 2793d074-519f-4448-b9f2-cfc2d738c4fa · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.696779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.696779Z digest=sha256:6074aaba11ae5ef3bd40878bd14e8ff89415d4a14bd67e172f2d8bd540293ec2

Observation d82f2c72-baa2-4505-9ab0-5d77ac44e58e · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.730067Z digest=sha256:10150686dbc12bd300c1f2e065d7ad1b2aed09f7fb71f0b1746efb19f4dc5eb1

Observation ea8717b2-0ba8-4a3b-8d82-ba41f72e027e · outbound

This paper cites Evaluating Very Long-Term Conversational Memory of LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Very Long-Term Conversational Memory of LLM Agents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.702317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.702317Z digest=sha256:c74fc31b65d22e68aea316e80d244ae8ab661502bd6ff8689d77a44b22984d26

Observation 8d34f258-59af-46a8-8de8-a4132c3b0ccb · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.705264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.705264Z digest=sha256:ec83aaafbd19998a50e7157243db0bbee3d5023dceaa11f20c251cca4069766a

Observation c47bb156-6098-4617-81f1-d02a36e835ab · outbound

This paper cites Evaluating Cultural and Social Awareness of LLM Web Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Cultural and Social Awareness of LLM Web Agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.741655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.741655Z digest=sha256:a6b8c254d1c1c0d4c422519e6b6e36af7cdec70987cc1091c8bfca94cc938dff

Observation f34adac2-6809-4442-b63e-0ee4f1d601d0 · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Evaluation and Benchmarking of LLM Agents: A Survey Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.744477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.744477Z digest=sha256:4b7042827ced9eea440357008818d46fd0414159737396e670c6cee6fb3d6c8e

Observation a63d1cf0-e7dd-4d36-8075-c300f73d5f27 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.713486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.713486Z digest=sha256:b1ef56ea996b14cce4423e3536b50b5ded657e8d3deca14f4a25bcfe0bda1b70

Observation f021281d-6cf7-46de-9b3c-47488b8685eb · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.716073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.716073Z digest=sha256:5c0a73be37dc4c9d7d6ad59baf5e37b57b3cddad5f00582772b01ee30e4bab28

Observation e4b2c4c2-a9cd-420d-b5c8-9c8bceee7f3f · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.718892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.718892Z digest=sha256:615b2c0d5599345639704b6f45de72b5ee693227735b807539251441d8a8e7cc

Observation a32272ee-a3ea-4d59-bd10-7c943f08e3e2 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.721508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.721508Z digest=sha256:1c68bd7d9e8031cf97b31a8e39e6073db597f02b00cfee1c8ef0b2672101b539

Observation 40861379-0377-47af-9bf4-7b6d5b2f7118 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Evaluation and Benchmarking of LLM Agents: A Survey BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.724101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.724101Z digest=sha256:f0b164bdfeab7dd6c07c362b7af852a939520d70f9db6f3b60bf8d020fcd0bb2

Observation 56cfd1f3-cdee-4b97-9412-567365435198 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

Evaluation and Benchmarking of LLM Agents: A Survey WebCanvas: Benchmarking Web Agents in Online Environments

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.727061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.727061Z digest=sha256:b800c1bd1047f8fb388de66ddf82cb77bbea3d042a0fff87da42aa27ce30934f

Observation 670d85e0-0722-4c56-8bcb-46e51cbff8b5 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.768017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.768017Z digest=sha256:6a9361c1286bb6075efdfed153dbbb662ed65204d4a1126a9556123954b480e9

Observation 67ee1029-8780-44f8-8fee-756590eea613 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Evaluation and Benchmarking of LLM Agents: A Survey Gorilla: Large Language Model Connected with Massive APIs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.732658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.732658Z digest=sha256:1fc47bf6c63490522046d0b84d65c64b5ccdba7daec7ca769397a99baa8bac92

Observation d30536c9-4172-4044-8786-5528383e4f5b · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.736007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.736007Z digest=sha256:389ee806e6fd4dc4172d93c6b77da9c85da24f34f5f43307b833af4230412d27

Observation b42bbf4c-8db5-45a2-910d-5e0cbcef3c63 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Evaluation and Benchmarking of LLM Agents: A Survey ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.738675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.738675Z digest=sha256:959dde8eb514a4aea3db4277ec9dbf172abae519e875443536877a3c08d5f715

Observation 3df9fabe-2935-42bb-bc9e-917387192ed9 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.782168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.782168Z digest=sha256:42ed58a3642557056e7341201f966f541d50d985d27e0894514bdf3143df5abb

Observation 78653e92-0386-4bef-a808-53abdf6e81e3 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.785167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.785167Z digest=sha256:e1fb7193f36e96b635b811b6c5c494835cca2af2a6ceebcf671916a6a41bd3d8

Observation 9e607f34-0109-49b2-9994-8cd5f0bcc300 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.747431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.747431Z digest=sha256:a17cc29991977f9775ea45260783741c251307c9664935082b4a7d942d9e0dc5

Observation e8dab5f9-2e2a-445e-8ec8-310bff789f07 · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

Evaluation and Benchmarking of LLM Agents: A Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.793466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.793466Z digest=sha256:d910d1cb68a42a4540c8699ccc363ebc6d765831f47699a5dfa7552c24228d33

Observation a939c8f0-8b53-4922-aa76-00d617573ea3 · outbound

This paper cites Reimann, Catharine Oertel, Florian A.

Evaluation and Benchmarking of LLM Agents: A Survey Reimann, Catharine Oertel, Florian A

Reference 93

Resolution
verified exact
doi, observed 2026-08-06T12:44:22.059811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:44:21.752880Z digest=sha256:1c63b593bbe1e954629151bcdcefe470cdc7f495b59f08a3a46524e18f55901a

Observation ff967955-d427-497b-a773-035c266ccb71 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.755622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.755622Z digest=sha256:aff5880a4dad242a823b611f0619a945176808d28dcdffe897b57828b05cac09

Observation 78cc4a51-1f84-4584-8b9f-0ba4a884a3f5 · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

Evaluation and Benchmarking of LLM Agents: A Survey Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.759303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.759303Z digest=sha256:a893f796720968d6f436a6312c3e3e9c766e83d0cca396d9763798bfab01b6b0

Observation 8fd32775-a877-4f2f-8338-4d5d8e800dc2 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

Evaluation and Benchmarking of LLM Agents: A Survey TaskBench: Benchmarking Large Language Models for Task Automation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.762114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.762114Z digest=sha256:1c13f56f40d8ba498eccfa2b430e6eff54d2e75650e117b393f4c7134cd3e89e

Observation ab70d92b-935f-40c1-8002-2c449028ec71 · outbound

This paper cites Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework.

Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.765199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.765199Z digest=sha256:0b741a1369ea0d07eb31143f61f7c3ab971dbbe7de233169a7bd846314e6926f

Observation 85876bdb-f874-4a30-8cf2-1ece2333c195 · outbound

This paper cites an unresolved cited work.

Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work

Reference 98

Resolution
verified exact
doi, observed 2026-08-06T12:44:21.999884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:44:21.810777Z digest=sha256:39ddf08375df6e69cee46116039126f81fb745302f64ce87238c597ff4f4e55e

Observation 474df2fa-d0e8-460b-b8bf-1b720f13b87d · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

Evaluation and Benchmarking of LLM Agents: A Survey MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.813422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.813422Z digest=sha256:b1e449cc5ffd6d586227cef9af6b65a550cd80fb49f3e8e3e14e5f7e9810be87

Observation 13f5ad70-d8d5-4c0b-9765-0743eb507494 · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

Evaluation and Benchmarking of LLM Agents: A Survey CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.773326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.773326Z digest=sha256:33d67c62f958ebe73df092258a463d5c07c770a129e1b00636bc43df3ef2f299

Pith citing papers

Observation 502c01ef-29cc-448f-98a6-0fec0e76d9b2 · inbound

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams cites this paper.

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams Evaluation and Benchmarking of LLM Agents: A Survey

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:14:08.422394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T11:12:46.626382Z digest=sha256:0fd315010367749fb8c92fe5873380c0ce12df4aafdad61e9727b2def57895b0

Observation b79c7602-7f7c-4d27-a357-45e108c435fa · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Evaluation and Benchmarking of LLM Agents: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.680558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b8efa42150cc28f61630dacfed0cfa4cba0c39a12ed66e7b5459bb07c25bd674

Observation 83459fa1-372a-4927-952a-000830629ab0 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:13:37.607891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:96c9fa86f6b3586b9776c87778b2ccb1e459752be14be833f1aab64a6d416d70

Observation 9e9ad036-2431-4c7d-a36a-e36b80b5c2d7 · inbound

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems cites this paper.

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.744554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:53:37.616447Z digest=sha256:cbb2e224ecebbfc8b2c8a43814432942e4208a11f7969b57c7b4a8a4544880bc

Observation b0fd8271-1371-4f80-9afb-b8ec1bb1dc4e · inbound

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness cites this paper.

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness Evaluation and Benchmarking of LLM Agents: A Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.024402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:04:29.791558Z digest=sha256:59591e7273e577d7d8ad8f0178bd75b8884d45b71a256fb9c1bf37a6c3f4128f

Observation 6e863949-2b1e-4d52-90bb-120f43c84ea3 · inbound

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI cites this paper.

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI Evaluation and Benchmarking of LLM Agents: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T01:59:52.218232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:59:52.218232Z digest=sha256:074e3ccfd4e2fa7e1b0e344a8f275d4749b90667bc248451eddb6146625f1101