Pith. sign in

Paper Citation Record · LEDGER

MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2503.01935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01935 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:51:24.479668Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.774566Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a61c6093-c986-49db-9947-37c7a3aeb1ab · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.357124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:49a4ee74717c54fc1859beba45a9f3c622f0a4a4a82ae5ce107468d4dca92942

Observation 27510b94-8cf9-4669-8915-68833004d2e4 · inbound

Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design cites this paper.

Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.558074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:36.558074Z digest=sha256:bbf31353476ab03392a30cb9ac64d1a228f21ec9be1144d00916ba55e0bf696d

Observation 5f9e543e-bfbe-4e9e-87ff-8f1bac9195ca · inbound

MAEBE: Multi-Agent Emergent Behavior Framework cites this paper.

MAEBE: Multi-Agent Emergent Behavior Framework MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:47.480547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:14:47.480547Z digest=sha256:c560d5727b520c71095d01c8d622960b4f3134b7971f95ebd1cee2b66aa08d9c

Observation c0b45b05-2cb2-4c42-a962-c81f9586c150 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.420035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:a5caa0ef03eb328b08fb2356a0bc94a8072034ee68d694797e4ef4d699e6f68c

Observation 0c3d80f2-62f4-4044-bb31-68a94500eb03 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.713862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.713862Z digest=sha256:05152271632b238c0ad99f36a115d1f58ef0b69c8c05783b4a81193038a89866

Observation 1b7069b3-c9b9-4bfe-aff9-bdfb63238d12 · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:59.020753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:59.020753Z digest=sha256:e2853983e341eef7484a46e034dd018867232384077f1241daedf31a356d26db

Observation d02d8bf0-e065-4fe5-8a9a-ad4782b7ba78 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.663765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.663765Z digest=sha256:00b6498670e8bb5e675cbb7949174e1d70f201e02492549685c06d696a524472

Observation c99c668f-8563-40b6-b67b-6700452afdfd · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.694169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:4429556c504a132a45cba050a51921f1f1bb4feebaa87d5d2d12368f8cb6d85e

Observation 8170dce8-26b6-470c-ac13-5f6f89c498e4 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.597705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.597705Z digest=sha256:26f266e644bbd2fdb38dccdee10ec4ee8e08d6a8ce12689c3bcb38346bda656f

Observation 49b3dc0e-83f4-4622-9456-f88cb66c88e1 · inbound

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback cites this paper.

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:51:24.479668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:51:24.479668Z digest=sha256:0baca06a0d6726605b7149c41db5b99447eb3a276fbe7a1916afde21f968f70e

Observation bfef6379-2f38-4db2-b083-513148057539 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.314704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:831cc40e834873487da987acb16270d57d87fb559d354dcee2774fc40f4d36f5

Observation 8d35b3e5-4ab4-4ab1-a102-a5fa0cdcab1c · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.748174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:f5f8d6ec0ee3f5dae5900c688bc627a948745247e680e32240f31e33132f6757

Observation 18a3958e-d235-45ee-9b01-d126a6916f3e · inbound

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation cites this paper.

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T23:58:48.504758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:58:48.504758Z digest=sha256:0541f2f21d7fea5cacdca1779830e164694a29fb6f925dc05391e33c5c313a3a

Observation d15010e0-8901-4f6b-8194-cf257e3f1519 · inbound

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence cites this paper.

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T21:00:05.911327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:00:05.911327Z digest=sha256:35c293215b73cb06062603c6a5ce1f493ac646a2a0bc8ef13ee8d8709d277d59

Observation 2881e264-9ac3-46c6-abf5-762bfc09a607 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:8937dc8600ccdc5109f89eabb5d2044cb8a458805be8c411b6c7a5a5ffe94b01

Observation fbcb25a6-0a15-4e04-9a92-5f7ab06450ae · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.216490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.216490Z digest=sha256:378333880fe0ff5f234d602cc490f43b7ba602fdb61f38984307e3676cda5896

Observation c0955dcc-c49e-4551-abb5-d0876c56fec9 · inbound

AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks cites this paper.

AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.908180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T21:50:20.721358Z digest=sha256:f805e71089a8b7edeee9b302968cc1a6f2d8aeb03a661423f32ed36b1f446788

Observation 252e6b62-495a-41a3-9e60-04fd99f068d2 · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:06.904354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:98bb48434c0d3e9894703846a80e115dfa7d7c8d8f3aeebabbc210105caab3a4

Observation bfe52e0a-8231-4ace-9ecf-4359059f2713 · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.521476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:194bbd5160fa1af609171e5826d2435eef6be1a35d7e094d31e69d446338df2c

Observation 7372c09c-75cc-441f-be0a-3fb26c719e42 · inbound

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples cites this paper.

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:55.920340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T03:27:03.589336Z digest=sha256:ea6ba09053a58ed33ecfe188769a9ca6525db53b2a646ddde0b848ac324dfc9c

Observation 00aa1280-cc3c-445d-840c-a5a1a8a3157b · inbound

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows cites this paper.

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.855701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:41:40.370052Z digest=sha256:ece5332e98155ba9fd21eae7f50c24d3595953572b87cf2c5f49f2aebd136840

Observation 10c0f2f7-b6ee-4420-803c-3172a489afa9 · inbound

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation cites this paper.

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.733658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T12:39:00.460842Z digest=sha256:5ff0c234564f2483586992a5767ad81e1b7c1092667e1902870d5e4278482a42

Observation 7773379c-34cc-48ac-8804-23ed0c0d0fca · inbound

FASE: Fast Adaptive Semantic Entropy for Code Quality cites this paper.

FASE: Fast Adaptive Semantic Entropy for Code Quality MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.776192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T15:18:44.055533Z digest=sha256:7f490ac0c055deffbb1ca3a09993c3619f54400afd2da0ebb8f9451d659c9ece

Observation 50ca8191-5879-40c2-9331-230a41baadee · inbound

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams cites this paper.

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.061568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T23:53:23.807636Z digest=sha256:9203a3ac6fc5a8f97386fef8827b897bdcf45f043b27e9234698cbd0527518b4

Observation a7326bcc-83b5-4437-9451-6f157cfe7998 · inbound

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems cites this paper.

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.951375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T00:42:20.220668Z digest=sha256:bdfab478873f33aa1b51ad2227ac2896b3f38114dd93dde1fc8215d9fa6153d4

Observation 4fceb272-b46e-48c0-829e-7c6da2ea20d3 · inbound

MAS-Lab: A Specification-Driven Validation Framework for Reliable Multi-Agent Systems cites this paper.

MAS-Lab: A Specification-Driven Validation Framework for Reliable Multi-Agent Systems MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:24:12.544089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T03:14:38.619889Z digest=sha256:f2ffee87345bf969386c522d9d47dcfd31a77900d67085a81048d844bc580060

Observation a261d788-53f4-44ef-8fe5-4c0af4caab2d · inbound

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities cites this paper.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:31.310957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:31.310957Z digest=sha256:cbf30651cbeae7b79660c0396f6f68fe975184d2225a77e2ec0fd20d9d9d54d8

Observation ff51b4c0-9728-4b8b-a89c-5779826742ec · inbound

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards cites this paper.

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T23:11:43.198650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:11:43.198650Z digest=sha256:58a69a8c6b6d8159e0a083bffcab6908f196fb45dd607d0050c486deb2d0dc53

Observation d9d2bbf0-6bf3-475e-99a6-2a48752b9d5b · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.435557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.435557Z digest=sha256:58fd77a8f9d5c569a73280296e454e15b8b74998025adeee0df2df43df3f4491

Observation 9c402360-e95d-4532-8683-620a58084def · inbound

Fluid Structure, Rigid Record: A Layered Organizational Design Framework for Agent-Native Organizations cites this paper.

Fluid Structure, Rigid Record: A Layered Organizational Design Framework for Agent-Native Organizations MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T04:38:28.907978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:38:28.907978Z digest=sha256:02f56d83b29d69e38cd84cbff5e56983e9688a4f4ea977db82408ccad557e319

Observation 206253d0-070e-4ad4-a5df-689771b3092e · inbound

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA cites this paper.

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:20:58.914022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:20:58.914022Z digest=sha256:eb5e177f10a3dd0540110cf099fdc1d9495fb39d85ad8faf56c2d0bd58dd2d20