Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:33.820587Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 42 inbound Pith citation observations for arXiv:2508.20453.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:33.820587Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:48.154137Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
37 of 37 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9bc495f9-5096-438d-ac3a-b4a7869698ea · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 517b534e-b9a3-442d-9b7a-6b2678e54196 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1ef0164f-c67e-4be0-9f6f-229c2c10f21a · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 59250d59-a07f-4ea4-84b4-cc81ad11d68a · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 528d9cbd-5314-4cc2-b65c-10b96f729e9e · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers https://api.example.com
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78d07faa-19d6-47fc-9670-1a83c847ac40 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 23bd3417-2427-4b70-8a10-eaf8279c7e66 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers user-provided parameters
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8cd65c99-0706-4587-be93-f3f1d03249d5 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers analyze heat exchanger with inlet temp 80°C, outlet 60°C, flow rate 0.5 kg/s
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 57b15ac7-158d-4368-8e3f-5c3f96e6ae9c · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1790129b-fa6a-4642-9e2b-80053fca5f07 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6adb6cf-595f-4170-88db-99f0b4433311 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1372f17d-f36b-460f-94e0-dffd1152f029 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb21d547-d497-40db-b354-29512c256d14 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8ebca78d-3568-4bdd-a5d6-6ee5d3f6cb02 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9c19ae7-dd75-464c-bd82-dbdd836a5dc6 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 191de8d9-b0e6-4fbe-b5bb-440e6ffc30c6 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers task_id":
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 54bdbb50-28ba-422f-9e73-11c8ca716b70 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 80416348-174b-4867-9ac2-43595f127855 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers solvability_score
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c46a3000-b3b4-412d-bc96-7764805e3115 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: Perfectly completes 40–60% of requirements
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88ca05a6-0557-4962-a34b-2eabfc03f57c · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of claims are perfectly grounded in tool outputs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a66308a0-12ab-4856-a3e9-3f1a49ffae1e · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of tools were perfectly selected for their subtasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1cf167a0-13e8-4e19-a005-3ab9b6300076 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of tool calls have perfectly accurate and complete parameters
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7956d996-f281-451a-8ce2-676f165273e9 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of dependency chains are perfectly executed
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b8149200-9703-4a4e-b572-54232f00cb9b · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% redundant calls OR 40–60% of parallelizable tasks were executed in parallel
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc2b8a19-d205-4df6-beeb-68292ce0707b · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers perfectly executed
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 45fca52c-9699-4f1b-b492-78f37f9aaa52 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Perfectly
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cc0e7a3e-0e15-484d-a062-a620d4617d8d · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6a8a78ee-8a7d-4de5-bd29-a6f43119e7cd · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 62dc9ffd-2480-4f9b-94ff-0758673e90ab · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0fd39ded-4035-42fb-bba7-a078754f890f · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9f6d5ffa-002d-4447-bfe3-64748952fde0 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 195aeaee-cea7-4511-83eb-78d8fc57b6d6 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers perfectly executed
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 644d4000-bae1-4221-b0ee-a9bc24436196 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6efecd99-9187-48f0-b933-6f8b95435b28 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1178d659-fa6c-4fad-9be6-b90c058b8786 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 23352d2a-0dcb-4455-9a67-dd15728e7418 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3b7dc23e-ba46-4d77-b98d-9351064a0318 · outbound
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers task_fulfillment_reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2d741c2-0fa2-4e5c-a3c4-086ab17086ed · inbound
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff51307-cc50-425d-a2df-e6a34fd40fbb · inbound
Toward Efficient Agents: Memory, Tool learning, and Planning MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46baf4f9-3bca-48b0-bbeb-ba76b3c5ed63 · inbound
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2aa53c48-4fa2-402d-8543-40934d1def6a · inbound
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 286c9eac-fabc-4c37-843f-12908d1aa990 · inbound
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69ff1207-abb2-48e7-8605-fd250ae49fdf · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406d4e06-13de-4e13-9608-e034324e1fc1 · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94fcd77-d42b-4ef9-9bc4-a0a62f496324 · inbound
PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb05c5c6-eb5d-4cb3-8073-2ffe7d1d8e00 · inbound
TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e640e963-d1be-458b-a080-81ddf2b6f775 · inbound
ClawBench: Can AI Agents Complete Everyday Online Tasks? MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab85f772-8a3a-4cf3-8cbb-2c66a856ed4c · inbound
ClawBench: Can AI Agents Complete Everyday Online Tasks? MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfdf2bd0-6e26-4321-9851-8e8e7855c463 · inbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 75d019a0-c510-465d-b6e3-a24a3bd23997 · inbound
Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 60a979af-a4ee-4817-95a8-ca09a4324b91 · inbound
From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3330379a-2427-464b-8fe1-0de32ec37c99 · inbound
An AI Agent Execution Environment to Safeguard User Data MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation df550c02-dd52-459a-aed6-c25951f65292 · inbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8abc306b-5faa-49d8-82b3-e66dafa5c7c1 · inbound
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40abeb55-a060-4076-aac7-4e3cd301c835 · inbound
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0f1fecb5-43f2-4630-8640-01b7f72fc847 · inbound
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 903350c3-bb02-499d-a1a1-783dacabfeeb · inbound
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9aed00ff-1add-4a6e-b59b-6d6afda7ab50 · inbound
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7fbcf87d-e43d-4bf5-b616-d1fef4bc1c42 · inbound
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cdb5c22d-e16d-4a8d-9190-ad25c21b077d · inbound
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bea7fd16-a9b7-4cbc-8c34-207c32d856fa · inbound
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 270
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2e42a51-9fd7-4a7e-99ab-9d6477bc2f49 · inbound
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8c021193-60c9-4b97-bd8c-a1ba5860f75c · inbound
Learning Agent-Compatible Context Management for Long-Horizon Tasks MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 25a4c94b-170b-4375-bda1-acf03394c30e · inbound
TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8badc2b7-429d-4466-b384-c0f303ee85e3 · inbound
Understanding How Enterprises Adopt the Model Context Protocol for LLM-Driven Software Engineering MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2464d5ee-ddc1-41fd-b7df-40499aa65895 · inbound
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a203ab98-fb02-4f59-80ac-3d462739b0ae · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 788aed7a-ae90-4f74-9b33-0b1d6029b8cf · inbound
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d231a8d5-b126-446b-b7d3-f621bb37afd0 · inbound
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f479b467-f481-42fa-96cf-e5840189294d · inbound
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a807c3dd-2474-4f86-9153-3be678c00937 · inbound
A Framework for Evaluating Agentic Skills at Scale MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93786827-d764-451e-89db-70b0e90cf7e0 · inbound
MetaPS: Adaptive Programmatic Strategy Selection for Market Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1ae63b55-b590-48dc-99f1-b2d020685d3d · inbound
Schema-Bound LLM Control of Scientific Instrumentation through Model Context Protocol Skills MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320fd98c-1383-4460-8d03-d7f05237fe76 · inbound
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5adf52ae-058d-4dfc-93c6-95c17fc42961 · inbound
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a9827f-7b4c-4ecd-8acb-40f1444fed33 · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98547dcf-eb3a-4191-aa38-35cf4adbfef4 · inbound
GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0dafa6-69d1-4a9c-a8be-acef5616beac · inbound
The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb54e9e3-c2c9-4565-9de7-17caffeff54a · inbound
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.