Pith. sign in

Paper Citation Record · LEDGER

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2507.12806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12806 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:46:02.635196Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:10:04.253250Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:10:16.398252Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965e8d88-5920-4fa0-997d-18671108562e · outbound

This paper cites online" 'onlinestring :=.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.805851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.805851Z digest=sha256:e63ec051a2505c79d9b3b1d32df647808a2d8137cadb39ab0ca70e263cb74919

Observation f48c8dec-5f67-4714-8de1-2e2eb14d62c1 · outbound

This paper cites write newline.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.871577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.871577Z digest=sha256:0321a6571da2f86ced63ccf8c66b6c78b08235999f93f936f7fae22aecf10e3e

Observation 54a68f72-8b63-4ba7-9b90-4afa6c213840 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.424251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:58.971979Z digest=sha256:f6612689679ccc7fe58c41f9f14cca22deea1934e9ee4a281a3f5b118166eee8

Observation 0ac6f1bd-aad2-4886-b45b-2511928efe4b · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.326853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.050121Z digest=sha256:721472c2db4c77b9752929be710c9321e280d4e92f1e3a3637d5d1da0b63f9d3

Observation 7f2ccccd-f0fe-44b7-975a-a4917b8daabc · outbound

This paper cites On extensions of the Jacobson-Morozov theorem to even characteristic.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models On extensions of the Jacobson-Morozov theorem to even characteristic

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.463822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.124683Z digest=sha256:4d8b4efe1276491b1a9b28b9f294352179b54a1e51a43e65e3688124b437c59e

Observation 6ddcdea5-a3c5-4978-b407-ccd568d67cb5 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.218192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.215761Z digest=sha256:3ec93cf7a6d63ff66d5aaab67bd847538679e93a6953959c2c069ee49e6a1590

Observation ff70678c-e4b2-44c1-80ea-76173553b1ad · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.117258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.308374Z digest=sha256:7900ebc6a83811705e45da79274145c3dbde53e6aea55df88ae5f6efd63601c3

Observation 382aea7f-4f3e-4d31-8b83-8ccda811beb7 · outbound

This paper cites High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.280548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.402697Z digest=sha256:119af50d3631acf59bd454ac106d1f181c4d384f021b17ab72ed120c7066aa1d

Observation 9963f502-26c4-4230-be21-15f6fc5160b3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.503158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.503158Z digest=sha256:3d319d99d089f783980e8e06947dd757930d16387893771835f463a37af716f5

Observation 8a58901e-2fd1-4e76-b9f1-3f512b7b46ab · outbound

This paper cites REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.600438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.600438Z digest=sha256:6f8279add7dd07a72e6fb4097993116c3864d1a48fb4bd9055cfa05579421de6

Observation 52fde6de-9293-4f89-9276-73de583e86ea · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.689261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.689261Z digest=sha256:5b172edf5b683654f5e67857a9accdf458f5f58d766d83be6293bbbc768fa4d6

Observation a05bf665-2307-4f98-84fe-e0c110d1fff4 · outbound

This paper cites LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.002763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.767986Z digest=sha256:2a008634647f52f2865cbf8617f055d74fe9351bbf2625ea38f1c4c7fa2d8178

Observation b413d103-a494-4f61-978e-3a676697e357 · outbound

This paper cites Recurrent neural chemical reaction networks that approximate arbitrary dynamics.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Recurrent neural chemical reaction networks that approximate arbitrary dynamics

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.888765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.835670Z digest=sha256:8cae1feb750717ae21668816fa001bd36126f7ff467074a710379dcc7e67a384

Observation 0c2a0c8e-840a-49e8-a759-c6788b9e8af9 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.972064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.913761Z digest=sha256:c14e84b40e7a9bb2551e8034225931c606c85d62fefce982c76ace575c6b6d19

Observation d5e270d4-5040-4774-9917-b0b1f1891b6b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.003696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.003696Z digest=sha256:0dae2c24ab53059482a6707fafe857c0f205cfa4d40a28f4e3915861cae385e2

Observation da1ec4f5-2cea-4d02-9714-852c90562a85 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.105921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.105921Z digest=sha256:bd02984c0e8c6834a5bd0f86de561b674077a34a18b9ae17f22ba2f2e39f8d6a

Observation 0e3145b2-54c9-4e9e-b271-36a5d95ca7b4 · outbound

This paper cites ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.165311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.165311Z digest=sha256:6bd6e6242e5d1aa0c98813e8ad1b31016da4e6af2a4da4102a7b48bcfc0b5932

Observation f16da7f4-79b5-4715-9121-e8d03a7cfd06 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.829666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.236217Z digest=sha256:99fed923f3f5b9850f51e496bb91f84bf42fc99f73ff64dc75d124b90ff56e98

Observation 39716159-e2c8-490a-8800-1e82ea79b651 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.317956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.317956Z digest=sha256:ce79526a75e12cf3e8d9fce9801ff5458701cf9038613ebb353dd7a101061d24

Observation 10d30c61-c805-483a-aca3-821be6d5c8e2 · outbound

This paper cites PRACT: Optimizing Principled Reasoning and Acting of LLM Agent.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PRACT: Optimizing Principled Reasoning and Acting of LLM Agent

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.694873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.387306Z digest=sha256:b1467cc0cac16b5930ebc7385e7f4de4c15baec1260ee32358774a6586dea306

Observation f1500648-c6dc-4fb5-bec1-bc9d37bc47dd · outbound

This paper cites BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.503290Z digest=sha256:fcd8d552edc4b0b59cbd2def45acb8950e6bcf4bc5f9775362f627c7130af777

Observation 3b85cbb9-9ee2-4265-8954-2dbfe0457389 · outbound

This paper cites AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.571687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.571687Z digest=sha256:bd12db6c899cd921b81a5e6185f778a0a1a599f164a674e644b96933caf01ee2

Observation 3533abeb-e1d2-4a2d-a2b4-fe04b0e9cc98 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.691587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.643162Z digest=sha256:1e7cab38dce1d7a2f3b93a18806ca9cee428a07935befb6c2434032f8caedf6c

Observation 5bacc67b-035f-41f7-9ede-cda04be96e11 · outbound

This paper cites ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.702329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.702329Z digest=sha256:74fe066fd56d81ee07665b9bcca576f7d47c9d0f747befac9d95c069de6491c0

Observation 59d6be4b-64f8-49cf-977e-956a9ec79f4d · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.781635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.781635Z digest=sha256:40f7f207b1bcf6c5fd0c26406f67eef371e75a3672db14eade27ce24324e5e45

Observation 601a5aac-4ffb-4790-836f-c4219bfdb3fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.878106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.878106Z digest=sha256:f10e11ad58254ce699fd4f21e96806b4cd038b3f85c2bbb6d3ac478f81a7eae4

Observation 4b9de82d-5851-4695-a7c3-af6f7b0f7fc1 · outbound

This paper cites GPT-4 Technical Report.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.956599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.956599Z digest=sha256:b150f5ed6c59c5f833a0ff05608beed29431fa18c5128e75c2129f5f05f7bbe0

Observation 1f61e1be-832c-47c5-ab38-ffe3057f52e3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.487334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.032125Z digest=sha256:e9c1d8b7e661e1771aad9a5f6f429ea224fd06b052d4104532f07092b5c936c5

Observation afb76a43-b3a7-42b7-aa5e-5c8f1f69623b · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.118506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.118506Z digest=sha256:6b55b9dd16e00cd99cee70f85d825b6892a0d672da9ac884e4abe72faf9f4776

Observation 98a46a96-655c-4937-a520-f84420cf7735 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.319496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.176261Z digest=sha256:7025908fa1390d27cb9696495f2a2d40866c5d1187f79ee248b4e19769e38222

Observation 277c3b40-f282-435e-8481-dc53cbe330fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.147837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.251698Z digest=sha256:004fbd6e1d7d1c5141022a7d2450920398a1d10aaee8c8a5389fae4dbf5a0631

Observation cea0ddfc-24f6-4de2-b2ae-41327a4f6009 · outbound

This paper cites PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.340983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.340983Z digest=sha256:c0180e459811491a4114cdba75454d1b5d49abfd527fc91810b451e5a4444314

Observation c31e5581-2880-462d-af0b-097ac6237e88 · outbound

This paper cites Peak Age of Information under Tandem of Queues.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Peak Age of Information under Tandem of Queues

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.339657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.420089Z digest=sha256:628d1d074f3a148f0609920300ad7e52c1d58df5f5225113537e4b766dcea1ff

Observation 91de0b95-21ad-4e1d-ad2b-f137266aff4e · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.019555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.505456Z digest=sha256:d2e66bf08df2e35e8e7ecbe221c0a3d15c0862c2c0e978e241b9cd37d4e060c7

Observation f13d3aeb-25cb-4548-b755-eafdb31c03dd · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.909834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.570593Z digest=sha256:6a2362238cb55dff57bbd07a23ae3d045721965e1f9cedf0dcffde999fb65cde

Observation 0fe92b4f-1fe3-41e0-8586-4fe378d8f71b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.623365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.623365Z digest=sha256:024eda43e9f775634bef5afc6d876be45559ab3e72fcf4ccd1dc3be41572c3ec

Observation b196f587-3af0-4831-97d9-a95dceff3733 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.750655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.676364Z digest=sha256:40b42d82135c705e586cb5203dd74fff189bc3204b3860be1e7d8bc3aa54c0dd

Observation 6aaf830b-3ec9-4809-b1b8-55ab9db3030f · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.738182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.738182Z digest=sha256:6d7df59d1a24b54149ccf84102abfea137c1a73acbdad0d39194c69efd07d4d3

Observation 24645806-899e-471a-9159-5064ce9a8a91 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.826594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.826594Z digest=sha256:da192e5449789d15077ba96ca06a594e7808ba2e857b5a549ae97809fb95f297

Observation 1eb775e8-6a66-43bf-8819-92406bbd28fe · outbound

This paper cites MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.921739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.921739Z digest=sha256:76855faa3f80773f77e2b19c1b78d5320e4baf464e6c5358f4355b63618c1bec

Observation 8cdc92fc-d07f-4d04-9a2a-05d6c3adb33f · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.613287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.974886Z digest=sha256:10add69217b3768d059130a7e2c96d29e7ec9d0abedc04b3192c211b18482b33

Observation a3b0ae8a-5fc8-465d-bbed-990c8b4dfcc3 · outbound

This paper cites A Survey on Large Language Model based Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models A Survey on Large Language Model based Autonomous Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.056791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.056791Z digest=sha256:52047ed54078d276b2fd2caa69ed54a7dc4a2a5b320397a6af87c63674f4b6dc

Observation a242dc43-64ca-404a-a3e3-fadb42c1a74e · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.127936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.127936Z digest=sha256:0afbbb740d74008c099517455cdbde418c4f84f73926628b1208f53bb6f3993d

Observation c607c54f-dc16-4cb5-bf78-d021f4fff877 · outbound

This paper cites ActionStudio: A Lightweight Framework for Data and Training of Large Action Models.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.088040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.199468Z digest=sha256:1a8f15cbe273281dcdb86ee69d089ff2057906caceb567a912add2619def7368

Observation 380140d5-000a-4b40-ba1e-8f10e9c4d317 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.276368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.276368Z digest=sha256:eb24937b3a4533b04a61ca609b64e5a3f04d64b57d21e899534b6bda5abbc1f1

Observation dbf4f137-366a-4b92-904e-d90b7f7e9806 · outbound

This paper cites DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.956374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.333562Z digest=sha256:88d9f4873a4ea5c2e935d7a0b5c21b27ca7c011f9fde707933d139b4fd5486b3

Observation 247dbebb-307f-41a6-a3ad-76cf635dbe64 · outbound

This paper cites Laser Printing of Silver and Silver Oxide.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Laser Printing of Silver and Silver Oxide

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.814754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.418806Z digest=sha256:602a28598e6e0670ebfb1e254737c66d3dcf66bf5f3c162bcdb8aec16067f695

Observation bdbe1501-a723-46ee-a294-bc95b9e765c1 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.488733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.488733Z digest=sha256:85112d39fd8df1d0751976360321c57492442349fed55884d4d48f9073aabbc7

Observation ccbb6752-8574-4506-9e62-d15cdad40e39 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.574065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.574065Z digest=sha256:b093885cc35be379a4ece3f15dc720c3a9b2bd4f509e36c51f280c6f890ce342

Observation f51c148d-fa70-4e16-babb-41243255782c · outbound

This paper cites Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.635196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.635196Z digest=sha256:05da38330cda1d6c312cfb6f45a9175085394452d360884fdfbf228d5c9e2062

Pith citing papers

Observation f50b43b1-2286-4179-9d51-536a3c89175b · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.621785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:edbf0df5bcd97989f48cc99802e27813763d294aaa0101d53fde03a9c5a7c755

Observation 639cab6f-2318-46d1-b552-f2ad2c5d1ae4 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.401136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:d0e4362dd7d98660a6fbd2ca3460e9200349fd4a246aee309d2a1342de5a0fda