Pith. sign in

Paper Citation Record · LEDGER

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 42 inbound Pith citation observations for arXiv:2508.20453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20453 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:33.820587Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:48.154137Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9bc495f9-5096-438d-ac3a-b4a7869698ea · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.701232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.615257Z digest=sha256:fdb48b911bec9efae86e0a49fc14da8f0bd22820d5565c52bcb37725cdb706fd

Observation 517b534e-b9a3-442d-9b7a-6b2678e54196 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.683309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.621206Z digest=sha256:050d033fa67ec5fcbb4360d0f496bc8dff9ec4ce7e225b9e9966e4e00d849539

Observation 1ef0164f-c67e-4be0-9f6f-229c2c10f21a · outbound

This paper cites reasoning.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.665049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.626807Z digest=sha256:6b1bc3b528e219a08f473ff7d4829bcdcb80b5919b2cadb39fed24c5f87fc866

Observation 59250d59-a07f-4ea4-84b4-cc81ad11d68a · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.642526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.632885Z digest=sha256:c9559d85ff8e6cc35009708c63170d04abe64c89c5b6885ee914dc7e6372a518

Observation 528d9cbd-5314-4cc2-b65c-10b96f729e9e · outbound

This paper cites https://api.example.com.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers https://api.example.com

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.622811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.641177Z digest=sha256:fdcb335bac77a932d3237dabb1651d99009627220f56deb87db5464d04b779cf

Observation 78d07faa-19d6-47fc-9670-1a83c847ac40 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.601838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.647229Z digest=sha256:291d6eae542b7e96b0c0868583272b73be1f2a26b4619b10b11415160a0f39ac

Observation 23bd3417-2427-4b70-8a10-eaf8279c7e66 · outbound

This paper cites user-provided parameters.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers user-provided parameters

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.581366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.653173Z digest=sha256:014730a1c34b79e0638ac5e4cc1ce5736ab5c932770e53ecea8d82a64ab696f5

Observation 8cd65c99-0706-4587-be93-f3f1d03249d5 · outbound

This paper cites analyze heat exchanger with inlet temp 80°C, outlet 60°C, flow rate 0.5 kg/s.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers analyze heat exchanger with inlet temp 80°C, outlet 60°C, flow rate 0.5 kg/s

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.555841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.659112Z digest=sha256:278b7cb5ac17dcc52a603d48c8d15edfef1fe4a8fed993ec46dce28faa36db72

Observation 57b15ac7-158d-4368-8e3f-5c3f96e6ae9c · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.533889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.665761Z digest=sha256:8b004f004bfb289a5b35bb5983b27e73f53e8348d2ff9c5f980fb676bdc4bb5e

Observation 1790129b-fa6a-4642-9e2b-80053fca5f07 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.515146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.671267Z digest=sha256:033f66c7feb6c03f658da1e6b4f7ab6ad85c5f2c11a1bf761275ddcd16a59894

Observation d6adb6cf-595f-4170-88db-99f0b4433311 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.492456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.676661Z digest=sha256:ce4ef05dda5cfc6909660811f64e4a967c47a937bbf606620d21a47dae5080ce

Observation 1372f17d-f36b-460f-94e0-dffd1152f029 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.470651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.681825Z digest=sha256:7c6e32969d10a457b48c581c7fe57043ca8a925c84ec3d580bcbd7f08581e1a6

Observation bb21d547-d497-40db-b354-29512c256d14 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.448645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.687741Z digest=sha256:b330c371399e0abba57ecfc9b54dd0d6ff682fe8b4d5a42db69c9041bb3d6e18

Observation 8ebca78d-3568-4bdd-a5d6-6ee5d3f6cb02 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.423036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.692862Z digest=sha256:67c1971517208784f85c4d0bba7450f45a40e441e37f06a97b633c1f8a0b12cf

Observation e9c19ae7-dd75-464c-bd82-dbdd836a5dc6 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.398304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.697758Z digest=sha256:94930cd13c6c7f26c5114698b8e47eb7d1216828ea6a28f5e3732e32e47a4f1c

Observation 191de8d9-b0e6-4fbe-b5bb-440e6ffc30c6 · outbound

This paper cites task_id":.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers task_id":

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.365396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.703063Z digest=sha256:563017dc9aa1012cda0b222f918c557863a08918a082762f6a8d91df33ad06a3

Observation 54bdbb50-28ba-422f-9e73-11c8ca716b70 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.335491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.708300Z digest=sha256:b35e2a063b8613594ffa0e73c417a71105fb55b6a587d2ca983e29775cb381c0

Observation 80416348-174b-4867-9ac2-43595f127855 · outbound

This paper cites solvability_score.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers solvability_score

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.308068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.713531Z digest=sha256:59a07f0f7fc55238a1136718fabc3a127c5ca0428c6c16c493988e8436ba336e

Observation c46a3000-b3b4-412d-bc96-7764805e3115 · outbound

This paper cites • 4–6: Perfectly completes 40–60% of requirements.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: Perfectly completes 40–60% of requirements

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.286957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.719000Z digest=sha256:ece96c3a2167353bc01bb9e7046df6a70d041056b61f5b39a8d4e988d15d9a2c

Observation 88ca05a6-0557-4962-a34b-2eabfc03f57c · outbound

This paper cites • 4–6: 40–60% of claims are perfectly grounded in tool outputs.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of claims are perfectly grounded in tool outputs

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.265634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.723985Z digest=sha256:6555fee541fdf5c1c4a6520c0fdd2a189372a8d66955932bbf6418073277d9b4

Observation a66308a0-12ab-4856-a3e9-3f1a49ffae1e · outbound

This paper cites • 4–6: 40–60% of tools were perfectly selected for their subtasks.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of tools were perfectly selected for their subtasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.241211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.730077Z digest=sha256:9987c7a24f0a2e1330166169ca03bcd07ace8bdd6dbec1d103d2abbecd60a1d1

Observation 1cf167a0-13e8-4e19-a005-3ab9b6300076 · outbound

This paper cites • 4–6: 40–60% of tool calls have perfectly accurate and complete parameters.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of tool calls have perfectly accurate and complete parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.213497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.735441Z digest=sha256:9bc55abf4fe021296b089781e69478c1dfbd6a60395ea0d91810fd69fc204372

Observation 7956d996-f281-451a-8ce2-676f165273e9 · outbound

This paper cites • 4–6: 40–60% of dependency chains are perfectly executed.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% of dependency chains are perfectly executed

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.190093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.740724Z digest=sha256:0711c41ea726ea42be3db6e3137d44e6e3544352d28f6be490f0aa869c2966fa

Observation b8149200-9703-4a4e-b572-54232f00cb9b · outbound

This paper cites • 4–6: 40–60% redundant calls OR 40–60% of parallelizable tasks were executed in parallel.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers • 4–6: 40–60% redundant calls OR 40–60% of parallelizable tasks were executed in parallel

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.170666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.745864Z digest=sha256:1abda30b178a96811e147c88c0ef76a8e265ea6541cfe098f56018c8882a2aca

Observation bc2b8a19-d205-4df6-beeb-68292ce0707b · outbound

This paper cites perfectly executed.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers perfectly executed

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.150979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.751663Z digest=sha256:51f03f65ac3ab8825af309ceb79d666459359bd7843d8d61be3c617b0d3e2330

Observation 45fca52c-9699-4f1b-b492-78f37f9aaa52 · outbound

This paper cites Perfectly.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Perfectly

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:34.126997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.757556Z digest=sha256:3ad9fc48e157f24cf47b85f76b99a0a22b94fa5d65c2903a55eda564b6f9ae6f

Observation cc0e7a3e-0e15-484d-a062-a620d4617d8d · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.105165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.764098Z digest=sha256:92f9b51d12c4318d631b508723285723d55128a94438b20e1515467ee0cf677d

Observation 6a8a78ee-8a7d-4de5-bd29-a6f43119e7cd · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.084014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.770170Z digest=sha256:1f5743ea215b47a8f2b515ed732ad684ef0d41c818c41455155cef2351809d01

Observation 62dc9ffd-2480-4f9b-94ff-0758673e90ab · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.061251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.776835Z digest=sha256:ab371ad1897d07fae6f90f0c95d5e29b7f2da36841963f4e91e0e1297ef41718

Observation 0fd39ded-4035-42fb-bba7-a078754f890f · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.036228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.782065Z digest=sha256:84fdd314304c53ba1442ace7a232fdd3324cd18092809f0725f9de9f8293b2f1

Observation 9f6d5ffa-002d-4447-bfe3-64748952fde0 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:34.015983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.788104Z digest=sha256:85698e833c58e7ce84d4b6675ff3dd5358ff98f4da780caf9b5d2e2321f4786c

Observation 195aeaee-cea7-4511-83eb-78d8fc57b6d6 · outbound

This paper cites perfectly executed.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers perfectly executed

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:33.989558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.793553Z digest=sha256:a76c18f472fed5e3dca6c9addbcf1cae0fc779850affb8226897b4a205fd133d

Observation 644d4000-bae1-4221-b0ee-a9bc24436196 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:33.967077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.799020Z digest=sha256:f926e9e2787309bfe33904ae0aed813eecae8f6559513e5817f3fdc246990b93

Observation 6efecd99-9187-48f0-b933-6f8b95435b28 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:33.945415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.804486Z digest=sha256:6e31c91ff49a54f6fbce6a2500188a06cce081af6726dee8f0b43a02ec3e9345

Observation 1178d659-fa6c-4fad-9be6-b90c058b8786 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:33.922676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.809625Z digest=sha256:f1709e465228a0af679a7d5307777bc4c391d466d851a5c28f6b393b4c933d0e

Observation 23352d2a-0dcb-4455-9a67-dd15728e7418 · outbound

This paper cites an unresolved cited work.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:33.893169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.814724Z digest=sha256:f01d176a235c16201d0f147dfa27a5bc3c2f47258b8af68376b4d2c5ecfbb19c

Observation 3b7dc23e-ba46-4d77-b98d-9351064a0318 · outbound

This paper cites task_fulfillment_reasoning.

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers task_fulfillment_reasoning

Reference 37

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:10:33.872765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:10:33.820587Z digest=sha256:5e3b8988d6b97243b1ee5e0ea7e12fa73e664a74d958f835faabb138a9f8b00c

Pith citing papers

Observation d2d741c2-0fa2-4e5c-a3c4-086ab17086ed · inbound

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents cites this paper.

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:31:24.596899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:31:24.596899Z digest=sha256:f70cfff94539333c8d851de1195f7492bc84878561544154e16d7efe9ca25476

Observation eff51307-cc50-425d-a2df-e6a34fd40fbb · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:44.317366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:44.317366Z digest=sha256:314e81c9543820406ce71838f96bcd00d71c6960d27e605a496c923ad592bafe

Observation 46baf4f9-3bca-48b0-bbeb-ba76b3c5ed63 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.652708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:ac80b3c1ada26f2d29124bf8df87add83960e68f704fd42c14ff5c4f460cfa42

Observation 2aa53c48-4fa2-402d-8543-40934d1def6a · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.430701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:fdb3fe1af8d686c929baeb466cc2ef712857d55698bb084e51572af10b8a9144

Observation 286c9eac-fabc-4c37-843f-12908d1aa990 · inbound

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation cites this paper.

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:07:11.549122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T03:04:17.755968Z digest=sha256:8295f40ca4370a44f11454e0ff4853b09797930bd676ff12fd27f7c33d68ad01

Observation 69ff1207-abb2-48e7-8605-fd250ae49fdf · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:705712745d2dcdb17003faa0b5a4b74d80a5f5e29ac29f350afeb8ff98b2719a

Observation 406d4e06-13de-4e13-9608-e034324e1fc1 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.165636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.165636Z digest=sha256:61bab393d9c161f8f425745e003c965b21cb9de5f7f3ce7b7898c24b31fc9fac

Observation c94fcd77-d42b-4ef9-9bc4-a0a62f496324 · inbound

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools cites this paper.

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:08:20.745113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T22:05:25.133535Z digest=sha256:f7ecab450c1dd26e6157ec1c447e14221cf0f603bac0b92a211b2f6b2e6cd7f9

Observation eb05c5c6-eb5d-4cb3-8073-2ffe7d1d8e00 · inbound

TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation cites this paper.

TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:10:59.596728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:16:35.875601Z digest=sha256:2d6628afe92ff29b659417d6e092b250419ae2854839bfd1338a8dab3d09e786

Observation e640e963-d1be-458b-a080-81ddf2b6f775 · inbound

ClawBench: Can AI Agents Complete Everyday Online Tasks? cites this paper.

ClawBench: Can AI Agents Complete Everyday Online Tasks? MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:36:01.212806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:34:49.746922Z digest=sha256:48a1608324aa4ea8ba89a00fdcdf2cc1d019148d06e234b4a3215b53784199ee

Observation ab85f772-8a3a-4cf3-8cbb-2c66a856ed4c · inbound

ClawBench: Can AI Agents Complete Everyday Online Tasks? cites this paper.

ClawBench: Can AI Agents Complete Everyday Online Tasks? MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T16:32:52.060485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:32:52.060485Z digest=sha256:c0e0eb9c15e7421d0da8d9d0a66d6fa03b183737653d7655615f811c821f3cca

Observation cfdf2bd0-6e26-4321-9851-8e8e7855c463 · inbound

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation cites this paper.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:20:58.062227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:ca35b0dc90bca4494b269236404e70e04871579efc8944c605d3bcd84d6aef9b

Observation 75d019a0-c510-465d-b6e3-a24a3bd23997 · inbound

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub cites this paper.

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:35:18.555855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T08:33:53.625397Z digest=sha256:ae27e83918c40c3a01efd82af0ccabfdd1cc478675a3856040ad6bf697dc73e3

Observation 60a979af-a4ee-4817-95a8-ca09a4324b91 · inbound

From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation cites this paper.

From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.782566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T06:20:49.799989Z digest=sha256:c40799312d2b43a7af9a68c53a9742150e3b0f8ba8d5e512699e7b70d732e3f1

Observation 3330379a-2427-464b-8fe1-0de32ec37c99 · inbound

An AI Agent Execution Environment to Safeguard User Data cites this paper.

An AI Agent Execution Environment to Safeguard User Data MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:05.854051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:14:40.639143Z digest=sha256:7d49359e2a254793d96bc6858aeae8fd4486e22cd39f024c431c2c7e4291d14b

Observation df550c02-dd52-459a-aed6-c25951f65292 · inbound

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents cites this paper.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:10.779830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:670e40a5950036538f344754940373b02be86561ab6e164cb4a3c8fed301fd7d

Observation 8abc306b-5faa-49d8-82b3-e66dafa5c7c1 · inbound

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows cites this paper.

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.155151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T05:52:13.100867Z digest=sha256:25c4df497aae995820ac0b031424fa0a07379673f20068511114130ccab3b9ba

Observation 40abeb55-a060-4076-aac7-4e3cd301c835 · inbound

PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments cites this paper.

PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:06.500219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T16:36:39.233362Z digest=sha256:6113a71b9083909514acdb454705690eb50c53cbfe1b584b259ec6f38b61a78a

Observation 0f1fecb5-43f2-4630-8640-01b7f72fc847 · inbound

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents cites this paper.

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:40:54.001965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T03:31:57.068492Z digest=sha256:e276b8336c8acb9b6849ed188e527fc5557332d45c45e98bbcd33dbc165402e2

Observation 903350c3-bb02-499d-a1a1-783dacabfeeb · inbound

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents cites this paper.

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:50.066435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T07:59:38.911667Z digest=sha256:bdf77ebd9661d29bd53734c6bdce3b8d5a1f43a75637ed683c67b541ea3fcf9c

Observation 9aed00ff-1add-4a6e-b59b-6d6afda7ab50 · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.539497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T04:42:47.496496Z digest=sha256:e08e3504c1b86a46a9fd195308d19288d21a3743a33e2884d297848f13560f49

Observation 7fbcf87d-e43d-4bf5-b616-d1fef4bc1c42 · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:09:51.241818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T08:08:04.544564Z digest=sha256:b8d1701550e2d89abfdabc0451613f40d5f89cb3d49889af5e21b6a9690266b2

Observation cdb5c22d-e16d-4a8d-9190-ad25c21b077d · inbound

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation cites this paper.

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:24.094059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:40:00.725327Z digest=sha256:6ee24ddedc2091e4580ce92aebaeb5420d627a68a625cf61997c3181e06b609f

Observation bea7fd16-a9b7-4cbc-8c34-207c32d856fa · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 270

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:49:53.688210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:c50cf6e24e86f216d3c8394b27e99f38b7286beca7bebaaa6ae4f5ecda029042

Observation d2e42a51-9fd7-4a7e-99ab-9d6477bc2f49 · inbound

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents cites this paper.

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:47:45.988091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T20:43:16.364125Z digest=sha256:914f8534b507710524bb17d462d96e9a195992d9530fcba8ce4a79fa3d7556f8

Observation 8c021193-60c9-4b97-bd8c-a1ba5860f75c · inbound

Learning Agent-Compatible Context Management for Long-Horizon Tasks cites this paper.

Learning Agent-Compatible Context Management for Long-Horizon Tasks MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:47.091491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:35:40.724644Z digest=sha256:14de2f5c143465ec824360ad57419d1356d94320daad1d90b81120017f36313d

Observation 25a4c94b-170b-4375-bda1-acf03394c30e · inbound

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories cites this paper.

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:36:08.584297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:22:54.473952Z digest=sha256:59093824a98f9c18a203b4f70e8ae739de582dcf6604f1b7f9168e8490f969b6

Observation 8badc2b7-429d-4466-b384-c0f303ee85e3 · inbound

Understanding How Enterprises Adopt the Model Context Protocol for LLM-Driven Software Engineering cites this paper.

Understanding How Enterprises Adopt the Model Context Protocol for LLM-Driven Software Engineering MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:37:34.622928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T15:48:39.837630Z digest=sha256:c6857690e174ffe1a459440004f2256874de7f4400156486fbf93fed1e522afa

Observation 2464d5ee-ddc1-41fd-b7df-40499aa65895 · inbound

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents cites this paper.

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.608185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:10:25.191306Z digest=sha256:33022966b4d72e827fd32f3c8b47146e280270ddd0769cae6a914aceaa69ba32

Observation a203ab98-fb02-4f59-80ac-3d462739b0ae · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.561027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:353db3facb1b87c1d595b385a784823837db7d991f7fce171a70b8eb5f8d97b9

Observation 788aed7a-ae90-4f74-9b33-0b1d6029b8cf · inbound

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents cites this paper.

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.180694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T09:46:59.954720Z digest=sha256:447f88b477fb57190028583874109a2d8646266986a1faa93b8f2ff05b75a36c

Observation d231a8d5-b126-446b-b7d3-f621bb37afd0 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:33.183784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:0933f5a7126446e906c4b6b01aabb5e58c13a042a6ec0a6b69785dd1cd421dcc

Observation f479b467-f481-42fa-96cf-e5840189294d · inbound

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents cites this paper.

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:43.983786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T03:57:49.824752Z digest=sha256:a7e8a00f58507254792442297ebd539489fb8e4ac61b0fff5fd9b59fa43a1c50

Observation a807c3dd-2474-4f86-9153-3be678c00937 · inbound

A Framework for Evaluating Agentic Skills at Scale cites this paper.

A Framework for Evaluating Agentic Skills at Scale MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:59.871322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T23:42:37.325511Z digest=sha256:4f1c24122cbe3cbaf04a8a098170be71205fbea8c9a9d705a1f90da2fbe91aaf

Observation 93786827-d764-451e-89db-70b0e90cf7e0 · inbound

MetaPS: Adaptive Programmatic Strategy Selection for Market Agents cites this paper.

MetaPS: Adaptive Programmatic Strategy Selection for Market Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:42.675663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T11:06:28.690956Z digest=sha256:a5f6e61f21ea555386aa6c0b9745b6806e18a3c1b25219f7c9247f4f546c5a4e

Observation 1ae63b55-b590-48dc-99f1-b2d020685d3d · inbound

Schema-Bound LLM Control of Scientific Instrumentation through Model Context Protocol Skills cites this paper.

Schema-Bound LLM Control of Scientific Instrumentation through Model Context Protocol Skills MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T19:19:04.089748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:19:04.089748Z digest=sha256:8b59f307934ff1703b6d3c3fa25f579784eb3b021334da5d1cb9ea48f1dd461a

Observation 320fd98c-1383-4460-8d03-d7f05237fe76 · inbound

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers cites this paper.

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T07:42:13.260641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:42:13.260641Z digest=sha256:6ebd87798db45443a9aacf9f29d2285f0967961faef9004dc2ef67626eb06d6b

Observation 5adf52ae-058d-4dfc-93c6-95c17fc42961 · inbound

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving cites this paper.

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T23:34:44.796574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:34:44.796574Z digest=sha256:7a45e2e22affbf341329ff35d89c5e7cda656947dd532dd70b6737510d567408

Observation b5a9827f-7b4c-4ecd-8acb-40f1444fed33 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:41.424525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:41.424525Z digest=sha256:ab527fdec733d79f246afec89e317f3367f91b2a6df02f0d143da732b6946cbd

Observation 98547dcf-eb3a-4191-aa38-35cf4adbfef4 · inbound

GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks cites this paper.

GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:33.992939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:33.992939Z digest=sha256:fd4e6caee977c08fe007e62cb889ee7ed9d49ade869d53cc81306915eda9485d

Observation 0a0dafa6-69d1-4a9c-a8be-acef5616beac · inbound

The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior cites this paper.

The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:35.510996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:35.510996Z digest=sha256:0361dee7affc7d011964e181e2881c92febccac916138be53380a5f4e7a4b0e0

Observation eb54e9e3-c2c9-4565-9de7-17caffeff54a · inbound

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies cites this paper.

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:48.154137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:48.154137Z digest=sha256:066bd59b9b0605426407596acfe5e520ee3f33669e2ccd9583ea1c0bfaf63674