Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.813422Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 6 inbound Pith citation observations for arXiv:2507.21504.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.813422Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T01:59:52.218232Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:27:56.022361Z
100 of 143 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 01cdcaa9-3927-49bf-beb2-97395950eaba · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99349306-a237-44b4-b97b-f48a3ee2428a · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b75862-6f1c-4a82-b403-cf6c41ea3fb6 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey 2024.Inspect AI: Framework for Large Language Model Evaluations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acd0a5f-50e6-42ab-b9c6-5501b546d41d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0183b9e-8c90-4431-93d2-6f8bc626cc76 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e084299-b3f2-4353-8702-1ac4c1ff8a31 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f046b40a-e071-4cab-b409-b5d75cd66385 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b32d450-e913-47cb-8d89-384fcf71762e · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ce98c8-6a51-450f-a48f-dfd9539403e7 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5fc4b80-d188-49ad-920f-98e5b6681c0a · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2094d729-a247-4398-969c-1bfd1da0a94a · outbound
Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724229a0-0d2e-477b-bc0a-f9fcc9135313 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33451500-2922-4100-8d9f-25563c283b73 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae999c9c-d449-48de-a343-0099d907a1c3 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey The BrowserGym Ecosystem for Web Agent Research
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3d1478-45bd-4d65-b7e9-7e7899f8da05 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3010720-da73-41a1-9bf6-9021a3d30020 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d1dbed-283a-4454-84b1-19d51b7eb010 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5e71eb-3fd5-4fc8-bee5-f08a2d330abe · outbound
Evaluation and Benchmarking of LLM Agents: A Survey GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f830591c-6c6c-4755-b8b5-a1772cfe8ebc · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93374fe1-760a-482a-98e8-9082b5f5cfd0 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Agent AI: Surveying the Horizons of Multimodal Interaction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33514e0b-4b21-431b-b16a-68848f3dad5e · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2992d1-6232-476b-b113-711a9cbf67ba · outbound
Evaluation and Benchmarking of LLM Agents: A Survey LLM Agents can Autonomously Exploit One-day Vulnerabilities
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac87e5d-7d37-4502-a571-c8eec1d99f71 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Re-ReST: Reflection-Reinforced Self-Training for Language Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f11e5aa-5671-45f3-8e40-c48c16c85aec · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f71b096-20f3-4f48-95de-22c4c408171b · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd12506-e3ec-4b8b-bf1a-21442e31dc7d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fbabe5-a8ae-4a36-b3a1-d6b09064913c · outbound
Evaluation and Benchmarking of LLM Agents: A Survey RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11fba66a-4d87-4501-bf86-cb4f7ade44bc · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4164373-270c-472f-a5c0-c1411bfdc6db · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45062438-4da3-4cff-adb6-6ce4ec281027 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ed874f-470c-43f2-a573-5a733ab8619a · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be0e896-3a10-443f-9f86-725f2db6b00d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Large Language Model based Multi-Agents: A Survey of Progress and Challenges
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0640dad-e5f3-42c0-9f55-e0a2bed5c23e · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d720d498-29d4-46c4-8d38-e7aed3ab185d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67888622-73c4-45c2-9b2e-e5bc28d3bc5c · outbound
Evaluation and Benchmarking of LLM Agents: A Survey A Survey on LLM-as-a-Judge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d265d4-3b1b-4bb5-b9cd-4067450ccb32 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03971990-d79d-4f98-931d-b8d22cc98ff8 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Understanding the planning of LLM agents: A survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf4dcc5-30b3-49f2-955f-a82ac905765a · outbound
Evaluation and Benchmarking of LLM Agents: A Survey LLM Multi-Agent Systems: Challenges and Open Problems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47616f2-fdfc-43ee-8e73-50016babc9c6 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4ffcf8-ac71-47ad-a8f1-b9f26e1d22c1 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f909b1-34af-4a68-bfb1-59360d7225e9 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Measuring Massive Multitask Language Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e49dfc-b559-4e10-8a8e-db2336f3cace · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1576480d-e593-4139-8f1a-353f60d92f52 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5195e60-8380-49c0-be77-aced989195e1 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17afe007-792d-4036-81d8-4f6e86d4e29c · outbound
Evaluation and Benchmarking of LLM Agents: A Survey LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fbcd1ef-99ba-44ae-9490-f3f8136fb2d4 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75ef5d4-56c0-4b5e-a026-a442f6ea8500 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae8e12c-84d0-4d3a-8cf8-5b793eede1c1 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dbed4c-c632-46b2-847c-0137bac07bca · outbound
Evaluation and Benchmarking of LLM Agents: A Survey VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08dd9ff0-baad-4333-b7dc-9e5ef96fc258 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988d73f9-f5c6-475f-830e-f6a5bc06d29f · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a1f55d-5608-463f-ba76-40eaeb9259b8 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ea9c14a-43b7-48a2-9a75-5b196197df94 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5134ab-59be-4e29-8c6c-d1a72f2f2df9 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32246f49-b35c-4ff5-9d98-25867ca6336b · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Autonomous Agents for Collaborative Task under Information Asymmetry
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a16e74b8-bc20-4e6f-afc8-922a972f43e3 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87cae1f6-0021-42e7-abc1-8c40a2a3919c · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0eda4b-5def-4aba-aad3-b03c122d49fe · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d83681e-34b9-4e2d-a9f5-7ea8018dcf15 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e97fee-f5b5-4387-a5b8-1543f55ffa76 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb862c6-b9b7-40ed-b0b5-1650e4f96031 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentSims: An Open-Source Sandbox for Large Language Model Evaluation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71eb74e-e7d4-4f48-a9fc-14f70eaf0f26 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4988fd1-00f9-411c-a20b-cd5f66f6298d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4c10d5-7239-4858-9037-39477c88bd91 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f007bf9-f524-4fe6-bddb-4a592f38fa06 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8693fcb6-3f7d-4d62-ba10-250858c969b7 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AgentBench: Evaluating LLMs as Agents
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6b6591-30e0-40d4-a330-824cbe460198 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bb5815-e184-4050-875e-aff77b02f5b8 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4833b80-cd87-46f3-804a-6d73e1bbc477 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f7fd7c-d72d-4031-b562-f6a09f14307b · outbound
Evaluation and Benchmarking of LLM Agents: A Survey AAAR-1.0: Assessing AI's Potential to Assist Research
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332611bb-3bbf-433e-ab58-d8aa519a7dda · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51527181-da25-4aff-a12a-678c0932dc05 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2793d074-519f-4448-b9f2-cfc2d738c4fa · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82f2c72-baa2-4505-9ab0-5d77ac44e58e · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea8717b2-0ba8-4a3b-8d82-ba41f72e027e · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Very Long-Term Conversational Memory of LLM Agents
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d34f258-59af-46a8-8de8-a4132c3b0ccb · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47bb156-6098-4617-81f1-d02a36e835ab · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Evaluating Cultural and Social Awareness of LLM Web Agents
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34adac2-6809-4442-b63e-0ee4f1d601d0 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Know What You Don't Know: Unanswerable Questions for SQuAD
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a63d1cf0-e7dd-4d36-8075-c300f73d5f27 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f021281d-6cf7-46de-9b3c-47488b8685eb · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b2c4c2-a9cd-420d-b5c8-9c8bceee7f3f · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a32272ee-a3ea-4d59-bd10-7c943f08e3e2 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40861379-0377-47af-9bf4-7b6d5b2f7118 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cfd1f3-cdee-4b97-9412-567365435198 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey WebCanvas: Benchmarking Web Agents in Online Environments
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670d85e0-0722-4c56-8bcb-46e51cbff8b5 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ee1029-8780-44f8-8fee-756590eea613 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Gorilla: Large Language Model Connected with Massive APIs
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30536c9-4172-4044-8786-5528383e4f5b · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b42bbf4c-8db5-45a2-910d-5e0cbcef3c63 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3df9fabe-2935-42bb-bc9e-917387192ed9 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78653e92-0386-4bef-a808-53abdf6e81e3 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e607f34-0109-49b2-9994-8cd5f0bcc300 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8dab5f9-2e2a-445e-8ec8-310bff789f07 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a939c8f0-8b53-4922-aa76-00d617573ea3 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Reimann, Catharine Oertel, Florian A
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ff967955-d427-497b-a773-035c266ccb71 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78cc4a51-1f84-4584-8b9f-0ba4a884a3f5 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd32775-a877-4f2f-8338-4d5d8e800dc2 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey TaskBench: Benchmarking Large Language Models for Task Automation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab70d92b-935f-40c1-8002-2c449028ec71 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85876bdb-f874-4a30-8cf2-1ece2333c195 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 474df2fa-d0e8-460b-b8bf-1b720f13b87d · outbound
Evaluation and Benchmarking of LLM Agents: A Survey MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f5ad70-d8d5-4c0b-9765-0743eb507494 · outbound
Evaluation and Benchmarking of LLM Agents: A Survey CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 502c01ef-29cc-448f-98a6-0fec0e76d9b2 · inbound
Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams Evaluation and Benchmarking of LLM Agents: A Survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b79c7602-7f7c-4d27-a357-45e108c435fa · inbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Evaluation and Benchmarking of LLM Agents: A Survey
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83459fa1-372a-4927-952a-000830629ab0 · inbound
The Scaling Laws of Skills in LLM Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e9ad036-2431-4c7d-a36a-e36b80b5c2d7 · inbound
Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Evaluation and Benchmarking of LLM Agents: A Survey
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0fd8271-1371-4f80-9afb-b8ec1bb1dc4e · inbound
Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness Evaluation and Benchmarking of LLM Agents: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e863949-2b1e-4d52-90bb-120f43c84ea3 · inbound
CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI Evaluation and Benchmarking of LLM Agents: A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.