Pith. sign in

Paper Citation Record · LEDGER

MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2508.14704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.14704 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:55:44.720516Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69125550-6611-4428-b14e-e6d3a9f12807 · inbound

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models cites this paper.

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:26.757169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:05:26.667750Z digest=sha256:ca67e831e4178189338f49bc58382b2a3226c93c4ba0cf80c63c70e3d0994ae0

Observation 9d92ba12-d334-4ef6-82be-45e464a741d0 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.635713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:a68b4138313e66c79029b0d0778d69bc40eb84ca759ea40151fd5c04d7626ab4

Observation dfa94d0d-b9c8-43b9-ae2d-edaf6eb15fa6 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.406044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:afe9b0cc79c1c8f5c8e4d8770acf2541b33862f2e132ad8e012dcbbbff7e7e82

Observation 593d568d-b3df-4616-9c2b-7c9f65680ce3 · inbound

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning cites this paper.

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:21:07.423753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:21:07.423753Z digest=sha256:2252e3bcca19e060a85cd7f653765c00f6f0c9bca768efac90f7768da9274958

Observation ce949294-6236-4f31-a63f-0081bf149ec0 · inbound

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation cites this paper.

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:07:11.521829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T03:04:17.755968Z digest=sha256:9ebffe7d9cd7d0ce74028292a6f12aea51d7bedc056bff6715a8ebbb0655bbfb

Observation 8e779399-18b9-4c3a-9840-75d953e16790 · inbound

Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions cites this paper.

Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T23:04:32.736348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:04:32.736348Z digest=sha256:8fb399bba6ef13a4ad60a0909227d0cf021ee59aa83f0e50c7d5c513e0067cd6

Observation 22552c04-e76f-4e2f-8c91-db610cfafe7e · inbound

Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy cites this paper.

Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-04T05:55:44.720516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:55:44.720516Z digest=sha256:c825e1368c49ed2fc0ce6e4765e4a12d2d4a88e951961e1aaf88603cd483761c

Observation a6a86117-703e-4b9c-8590-b5f47a97040b · inbound

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools cites this paper.

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:08:20.734278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:05:25.133535Z digest=sha256:7c4c7952fa5f1cf6a4fad0aa5df5ab421713cfa93e4475e96e17a2ac81cee970

Observation 2e7f9f2f-92a2-47ba-8be8-9724141b9a85 · inbound

ANX: Protocol-First Design for AI Agent Interaction with a Supporting 3EX Decoupled Architecture cites this paper.

ANX: Protocol-First Design for AI Agent Interaction with a Supporting 3EX Decoupled Architecture MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:49.940609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T18:42:40.250417Z digest=sha256:f2f9ecca88e3003f29145d18f4d0390a9f0cf9520e6a51d2f4d29a5b3a7df21d

Observation e42fd763-d473-4809-ba00-dbce8ed4a7b7 · inbound

From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation cites this paper.

From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.755658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:20:49.799989Z digest=sha256:452acf19e0dcba20c5952987bb4385fc0cca16300ff5cb2943ebc0880e3bbdac

Observation 21447ab9-9ac1-43d2-b5e5-bbb40ea66f89 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:25:54.737759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:ea0b2625e65d7ae95fbc64c39975c79a23dc6989748f9d30eacdbf3d4073274e

Observation 655a4ae8-3cce-4af4-8345-63cbf2fcafaa · inbound

Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization cites this paper.

Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:01:07.105970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T23:42:00.548132Z digest=sha256:29d7af768ab801f8636bebbd6da57abbf2c780dea799c981e94e9757f60ea1c7

Observation 1d10706c-2cf3-4669-9e20-90ea824c2b23 · inbound

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments cites this paper.

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:11:15.317013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:11:05.857658Z digest=sha256:10d39d245ffd9ec55008f0b835c2da8570401a215a32bd8fa7c4c1191ae0f789

Observation 71709c90-878f-40d4-b904-428615cab0ec · inbound

PREPING: Building Agent Memory without Tasks cites this paper.

PREPING: Building Agent Memory without Tasks MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.160195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:14:14.385586Z digest=sha256:5867074d9046dfae5bdc402e4fa6a26789025e50a377d57e2319bac2acb44840

Observation a3ed0ffe-3c32-4572-9ab3-7a7969a97fe5 · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 271

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:49:53.738775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:09dc4aeada02ef4d748391504e6c5e698d7b0d162e6891030fb4ada82b097321

Observation 59c9c8a8-3d89-4314-a560-1749b756683e · inbound

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents cites this paper.

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:47:45.985164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:43:16.364125Z digest=sha256:b1ff7bfea40a9dcb5756a16969a505d7d973615358c2b8e28dcaa68105576dcd

Observation 09363262-59e4-48f3-8ff4-47c02fbe7bcf · inbound

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems cites this paper.

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:13:16.038067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:13:03.046209Z digest=sha256:51a46de572e4063f803dc7d14d56df12cbf8de3a04a778ebb9ecef5d7bcff988

Observation 434f5c42-a5a1-4f85-aeb5-c896307475a6 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.301730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:a54322137b23870ed4841fe2f9c724fa4e290018a7d1bbf0533a55f86e367bad

Observation 64d89991-898c-4f9d-83ab-65751ace8828 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:33.191600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:fcb15e42dc1f71af44341eab5026d66a8b58285f7dd8643a493c886a8015e6f1

Observation b432a545-0c98-410b-9ef8-55c4d090ae68 · inbound

DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling cites this paper.

DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.727542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T04:04:23.415175Z digest=sha256:61ecd725a4462a9365ef00602c5585f06568b5e9a520d16eda0282831d1fb061

Observation 22815be2-c235-482a-8200-94c1d5d31fbe · inbound

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents cites this paper.

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:30.310286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:01:55.716141Z digest=sha256:4edb41e02145ce30225aef860994b6d2e6ccd33afc5045284fa8586d3b3a8f95

Observation 9c54a24a-1415-432f-aabe-98facabdc82b · inbound

Metis: Bridging Text and Code Memory for Self-Evolving Agents cites this paper.

Metis: Bridging Text and Code Memory for Self-Evolving Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:56.693386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:42:55.758578Z digest=sha256:decb0930a2db20ab886f62fb7c2173ca073e7eff9c1d8f412a7c05828223c740

Observation 64bf4e58-c472-490b-9086-b2cd84c4b7e1 · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:35:48.601770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:43df6760658557b4e32caecd340fab116922e5f30b739d70cb4b6c227bb49c42

Observation 9faf4ce9-fe33-44b6-8e92-0c9a08061033 · inbound

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios cites this paper.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.485800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.485800Z digest=sha256:3ed12b9cfecc061cc48218462522c2a4f3ccc2730595e23117d704fe4456e413

Observation 4e53b575-3a82-4944-b637-80a33b86ca9b · inbound

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting cites this paper.

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:20.572273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:43:20.572273Z digest=sha256:128a60fba5b42fb25870876bf9056567349bb80b1c0c092ce61d6ddb1e420c74