REVIEW 3 cited by
Evaluation Report on MCP Servers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the rise of LLMs, a large number of Model Context Protocol (MCP) services have emerged since the end of 2024. However, the effectiveness and efficiency of MCP servers have not been well studied. To study these questions, we propose an evaluation framework, called MCPBench. We selected several widely used MCP server and conducted an experimental evaluation on their accuracy, time, and token usage. Our experiments showed that the most effective MCP, Bing Web Search, achieved an accuracy of 64%. Importantly, we found that the accuracy of MCP servers can be substantially enhanced by involving declarative interface. This research paves the way for further investigations into optimized MCP implementations, ultimately leading to better AI-driven applications and data retrieval solutions.
Forward citations
Cited by 3 Pith papers
-
Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models
A new MCP benchmark across six LLMs finds that proactive tool use is rare on first prompts, instructed tool use mainly improves in two-turn dialogues, MCP context degrades accuracy by about 9.5%, and input-token overh...
-
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
An LLM agent framework where the model actively emits structured server/tool requests, retrieved through hierarchical semantic routing, reducing context overhead while maintaining tool-selection accuracy.
-
Adapting Embedding Models for Agent Capability Retrieval
Fine-tuning three off-the-shelf retrieval models on AgentSelect improved query-to-agent ranking on two unseen marketplace catalogs, MuleRun and ClawHub, across all three model families.
Discussion (0). Continue with ORCID to comment.