Pith. sign in

REVIEW 3 cited by

Evaluation Report on MCP Servers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11094 v2 pith:GKOIFX4Q submitted 2025-04-15 cs.IR cs.DB

classification cs.IRcs.DB
keywords accuracyevaluationserversachievedai-drivenapplicationsbeenbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rise of LLMs, a large number of Model Context Protocol (MCP) services have emerged since the end of 2024. However, the effectiveness and efficiency of MCP servers have not been well studied. To study these questions, we propose an evaluation framework, called MCPBench. We selected several widely used MCP server and conducted an experimental evaluation on their accuracy, time, and token usage. Our experiments showed that the most effective MCP, Bing Web Search, achieved an accuracy of 64%. Importantly, we found that the accuracy of MCP servers can be substantially enhanced by involving declarative interface. This research paves the way for further investigations into optimized MCP implementations, ultimately leading to better AI-driven applications and data retrieval solutions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models

    cs.AI 2025-08 reject novelty 6.0 of 10

    A new MCP benchmark across six LLMs finds that proactive tool use is rare on first prompts, instructed tool use mainly improves in two-turn dialogues, MCP context degrades accuracy by about 9.5%, and input-token overh...

  2. MCP-Zero: Active Tool Discovery for Autonomous LLM Agents

    cs.AI 2025-06 conditional novelty 6.0 of 10

    An LLM agent framework where the model actively emits structured server/tool requests, retrieved through hierarchical semantic routing, reducing context overhead while maintaining tool-selection accuracy.

  3. Adapting Embedding Models for Agent Capability Retrieval

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Fine-tuning three off-the-shelf retrieval models on AgentSelect improved query-to-agent ranking on two unseen marketplace catalogs, MuleRun and ClawHub, across all three model families.

Pith tools