Pith. sign in

Paper Citation Record · LEDGER

MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2503.01935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01935 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:36.558074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.774566Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a61c6093-c986-49db-9947-37c7a3aeb1ab · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.357124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:b646d592d8022946f55973f20e2ffc37b4c516c67b8d2d559062c0381f53f9b7

Observation 27510b94-8cf9-4669-8915-68833004d2e4 · inbound

Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design cites this paper.

Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:36.558074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:36.558074Z digest=sha256:8b2daa9d217fec8beaa986c02ad6700332a6ddf1e17adb87186c64a02674b494

Observation 5f9e543e-bfbe-4e9e-87ff-8f1bac9195ca · inbound

MAEBE: Multi-Agent Emergent Behavior Framework cites this paper.

MAEBE: Multi-Agent Emergent Behavior Framework MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:47.480547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:14:47.480547Z digest=sha256:542f70090a97efe1bccf95e9d0c6a43858626253cc2453315e75666413467c7c

Observation c0b45b05-2cb2-4c42-a962-c81f9586c150 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.420035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:3faebec9a12cd4e24bcc94c269561563b9522e29c41a51b52115da72016163e7

Observation 0c3d80f2-62f4-4044-bb31-68a94500eb03 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.713862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.713862Z digest=sha256:297635eaf6f02562758e577c52cb58fa224327f24d7ec9794671396e8d76f45d

Observation 1b7069b3-c9b9-4bfe-aff9-bdfb63238d12 · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:59.020753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:59.020753Z digest=sha256:fc45a57b6a0901c9d7d98e5b4f9b1c51676b6dacb4747adee7e573d1b89006d0

Observation d02d8bf0-e065-4fe5-8a9a-ad4782b7ba78 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.663765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.663765Z digest=sha256:5ecb5500583516963270d37f53021bbcbb35516a30c1ce000690676f5b1c9ca5

Observation c99c668f-8563-40b6-b67b-6700452afdfd · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.694169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:d6e4d6ce245b0744f59ad3471e81e3fcc30972b83eb4597c059467f042af43d5

Observation 8170dce8-26b6-470c-ac13-5f6f89c498e4 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.597705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.597705Z digest=sha256:81ef0f504df78a9b775d7520aef451851726fd573cd1c6cf4808fb570a24dce9

Observation bfef6379-2f38-4db2-b083-513148057539 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.314704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:a6d8c567f75931117ee9855f0a8bb57b67fecd29d6ec7c27f15d67f489132f23

Observation 8d35b3e5-4ab4-4ab1-a102-a5fa0cdcab1c · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.748174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:dca559e5ca37c0cc1fbaa1c72012c9945ccc7475e8a2f1ae0cb77ebc7792f96c

Observation 18a3958e-d235-45ee-9b01-d126a6916f3e · inbound

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation cites this paper.

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T23:58:48.504758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:58:48.504758Z digest=sha256:c4bbcc261c296d7d942197f9f78798f83407afc63e914aabea0d300a5727ec59

Observation d15010e0-8901-4f6b-8194-cf257e3f1519 · inbound

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence cites this paper.

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T21:00:05.911327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:00:05.911327Z digest=sha256:bf8bc7392094f63010b800e97ded6f39cb7e53dea5f00cae482ba96c9b6b67e4

Observation 2881e264-9ac3-46c6-abf5-762bfc09a607 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:405ae905c05fa3619e63151245f0f3a91c2d9374753372940fae3bb122f46ef4

Observation fbcb25a6-0a15-4e04-9a92-5f7ab06450ae · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.216490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.216490Z digest=sha256:7634e927a9d16b0df6d33b7f1abfcc743af2868d93dc9acfccedf9f42a8bcf8a

Observation c0955dcc-c49e-4551-abb5-d0876c56fec9 · inbound

AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks cites this paper.

AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.908180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T21:50:20.721358Z digest=sha256:55a330d71544ecc932cc513ace51da355a734b6b5669112d18ae92760da7cf33

Observation 252e6b62-495a-41a3-9e60-04fd99f068d2 · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:06.904354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:8d3f6860f70659cc55657d271161550a0136041440b4df9aceeb61d6549e549c

Observation bfe52e0a-8231-4ace-9ecf-4359059f2713 · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.521476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:0bcead67769c99ac21507033ed850c9535d599499fb032cbfd538f76fb79330a

Observation 7372c09c-75cc-441f-be0a-3fb26c719e42 · inbound

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples cites this paper.

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:55.920340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T03:27:03.589336Z digest=sha256:5c30384e161f5f6e1d4bc86f21cb29f505515c056a140fbbc9537b98c1070c6c

Observation 00aa1280-cc3c-445d-840c-a5a1a8a3157b · inbound

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows cites this paper.

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.855701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:41:40.370052Z digest=sha256:451499196faac582eff4b44adc5a0dd854d90b7f4757d93e7c11196fada45263

Observation 10c0f2f7-b6ee-4420-803c-3172a489afa9 · inbound

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation cites this paper.

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.733658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:39:00.460842Z digest=sha256:322443143ae7e9a88b6588ceb882909f4e3052cb16a43207455ac66ce412cd34

Observation 7773379c-34cc-48ac-8804-23ed0c0d0fca · inbound

FASE: Fast Adaptive Semantic Entropy for Code Quality cites this paper.

FASE: Fast Adaptive Semantic Entropy for Code Quality MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.776192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T15:18:44.055533Z digest=sha256:f16485f84d98a095be7c7fa805a9a10f46519c0ff3777f30b36597f6b38c279b

Observation 50ca8191-5879-40c2-9331-230a41baadee · inbound

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams cites this paper.

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.061568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:53:23.807636Z digest=sha256:496586c967aeebb6a3245ee6b2a8dc8e8a710e713bfed2c410a198ccfe41e55e

Observation a7326bcc-83b5-4437-9451-6f157cfe7998 · inbound

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems cites this paper.

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.951375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T00:42:20.220668Z digest=sha256:74f00ef41eaec94a89aa55fef477d781b0b91f80fe5e1d3f0a15f934bc5af7ea

Observation 4fceb272-b46e-48c0-829e-7c6da2ea20d3 · inbound

MAS-Lab: A Specification-Driven Validation Framework for Reliable Multi-Agent Systems cites this paper.

MAS-Lab: A Specification-Driven Validation Framework for Reliable Multi-Agent Systems MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:24:12.544089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T03:14:38.619889Z digest=sha256:ffb5190b4b99c49ae6dfebd426a477021d99c85e203c4de2c7a169f3d6422367

Observation a261d788-53f4-44ef-8fe5-4c0af4caab2d · inbound

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities cites this paper.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:31.310957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:31.310957Z digest=sha256:ab0ae521a41c747cd8b032f98274ed33d0d7bfa668cf2d79b61ec51add3a9a2c

Observation ff51b4c0-9728-4b8b-a89c-5779826742ec · inbound

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards cites this paper.

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T23:11:43.198650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:11:43.198650Z digest=sha256:c41bbfea6b913851b6f6670dcaff9215ccf5bc5f055e12ae34bdeafb9940e6fd

Observation d9d2bbf0-6bf3-475e-99a6-2a48752b9d5b · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.435557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.435557Z digest=sha256:64ae03b8eead0076d720c1749312a8a328ad3c70f30cd40abbae525f872b3708