Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

As of 21 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2606.08340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08340 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T19:18:56.039587Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:25:39.314247Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch3

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation f7a5f70c-ba09-48b8-ac61-afd210f3f65b · outbound

This paper cites Melting Pot 2.0.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Melting Pot 2.0

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.837587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:56275c373be6d4d55274e163fb89e5de1015d2599b2794e0ce1e417137d87082

Observation 50d813f5-5692-4c5d-bd47-706b88e236c0 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.Advances in Neural Information Processing Systems, 2021.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Deep reinforcement learning at the edge of the statistical precipice.Advances in Neural Information Processing Systems, 2021

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:178ac94f133033facee9be7b167a3eba664014186e45e795524068790f188c09

Observation b59aa67e-4c3e-4ef6-9122-dcf7dc512583 · outbound

This paper cites Llm-coordination: evaluating and analyzing multi-agent coordination abilities in large language models.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Llm-coordination: evaluating and analyzing multi-agent coordination abilities in large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:21c46906b5da3338eab08baaa65c2aeff7ca8616ecf61f7b2203f82a0faa8fa4

Observation 8db17a47-ad0b-4236-b701-0bae4c2f5bc4 · outbound

This paper cites The hanabi challenge: A new frontier for ai research.Artificial Intelligence, 280:103216, 2020.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The hanabi challenge: A new frontier for ai research.Artificial Intelligence, 280:103216, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:da55bc24faffc6dcd02a63457d76ca29ade9f0470506d21df9302fd11a18fb97

Observation 803e569b-d1ab-4ba3-94b6-bc9fddcb0dcc · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279, 2013.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279, 2013

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:55ca97c8386060a44dc56954e793a93a99c197d08fe82f906936b9c8fc0e77e6

Observation c514a66b-27d8-4ed0-af8e-14ce846ac99e · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents JAX: composable transformations of Python+NumPy programs, 2018

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:37502770bf84f6b4953e9d2a8f8408455d7edcf42fff9cc550837b1053b78f4d

Observation 26c9b99c-9012-4622-a9d5-e36cf6f0ae91 · outbound

This paper cites Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.Science, 359(6374):418–424, 2018.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.Science, 359(6374):418–424, 2018

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:ee59a64e5b4bf2ab746bc84ac099834eb76906aae0ba433143f3c270fa578b22

Observation 2c1b7b02-577f-4e52-ae4d-07d4ed26a3f6 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:11cd294221513432d1346d17da99971567b84b8257c9b00da7b5a9ddeca8cfdc

Observation 3220a018-d7c7-4453-864a-b777d31e1d42 · outbound

This paper cites Deep blue.Artificial intelligence, 134(1-2):57–83, 2002.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Deep blue.Artificial intelligence, 134(1-2):57–83, 2002

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1c6912ea48763126c2171a7524678658deec090033a413ca67f23a745206a39a

Observation 847bc0f1-5d20-4c7d-bc1c-256753e091e3 · outbound

This paper cites On the utility of learning about humans for human-ai coordination.Advances in neural information processing systems, 32, 2019.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents On the utility of learning about humans for human-ai coordination.Advances in neural information processing systems, 32, 2019

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6949b8f6891301caa89ea90f5e00f408c5836ec08df05bf6afccd2057e577280

Observation eef40610-6b56-42f6-bcb7-caf2542005c9 · outbound

This paper cites MLE-bench: Evaluating machine learning agents on machine learning engineering.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents MLE-bench: Evaluating machine learning agents on machine learning engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:7da540089a785bc0794c83c1261859ce43d7ba4c7504fe0e5e03aa50de9133dc

Observation 25717b9d-3024-4e1f-be48-7c30eb504949 · outbound

This paper cites Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:07:25.850015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:c7ceaac115e732f6c455e76ec8355f16030918cc73421f0331a74df9aaef3398

Observation 022147d5-e39f-4f92-a6fe-d46d39ef1861 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:a4eda9e03eeece48d456937a8da4927755720ce510934c4527e10279f6bf10a5

Observation 6135fc64-bbc0-45cb-a220-e6507f5fa373 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Improv- ing factuality and reasoning in language models through multiagent debate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:88cf6b68ffbdd8381619848aff4eb6d21006250a24c84f18147437549c9e0ad8

Observation a1bdd995-3ba6-4820-9b00-385280459451 · outbound

This paper cites Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning.Advances in Neural Information Processing Systems, 36:37567–37593, 2023.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning.Advances in Neural Information Processing Systems, 36:37567–37593, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6ab821bc217dc2031e3a2d846ebd1458022d3aac5b339fc4497a1e178941b546

Observation 4b2b49d0-aace-4643-834d-4def46249717 · outbound

This paper cites Simplifying deep temporal difference learning.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Simplifying deep temporal difference learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:871282480126d92a7263d38efe40252311e90a12369eeb8b225824d9cb1c83ad

Observation f0b3de9c-05b8-4e14-8dce-40942b4e84fa · outbound

This paper cites Overcookedv2: Rethinking overcooked for zero-shot coordination.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Overcookedv2: Rethinking overcooked for zero-shot coordination

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:9b0d2682245f750c096373b8917e0fbe65d2fd7852fff51061b8595d149fcf3f

Observation 763b77cd-2070-4a74-88b2-85bf9692f2af · outbound

This paper cites Gemini 3.1 pro model card.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Gemini 3.1 pro model card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:8ba3cec437b2851ab22e7ad04412c60a47e8cb9dc16251d7c02ca761b7d992c5

Observation ac242237-0a68-44d5-881c-fcc2c989b991 · outbound

This paper cites Gemma 4 model card.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Gemma 4 model card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6baccec56b048c5c5cd0fa178562f09207d2e966baea8e72a75cc751c5c30ffa

Observation ed9eee9c-c55a-4f6d-8bad-2081d0459431 · outbound

This paper cites KellyBench: A Benchmark for Long-Horizon Sequential Decision Making.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents KellyBench: A Benchmark for Long-Horizon Sequential Decision Making

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:07:25.852225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6099cce162c17e3ee99f3350f9c87bc4943a18e88a3314a66be0e30f6fb55432

Observation 7a87bf2f-8fab-406b-878b-458a0116a3fd · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The Llama 3 Herd of Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:07:25.834809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:7ab28fc1854426c3f5092be6bb38c7012f386d2180fdc2a7436a2ab33e1cbe09

Observation 7fea0335-8c3f-4c5b-a5e3-6f1af02faf90 · outbound

This paper cites AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.843950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:f039190dd21b1e4273245a23fd81c30fb6cff6c74aab2f42e7cc913aeaf236cc

Observation cb811d90-d298-45cc-9c56-231bf212d622 · outbound

This paper cites Large language model based multi-agents: a survey of progress and challenges.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Large language model based multi-agents: a survey of progress and challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:bfb1578624ae9476a891ef854597da5394be2fa1fdc2ecc5cfc1a92b7a774bdd

Observation 2a159151-40fd-4730-b126-e7f4f75c37c3 · outbound

This paper cites Benchmarking the spectrum of agent capabilities.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Benchmarking the spectrum of agent capabilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:72ee36b7e13466bf6ef6b55639aba5f76bca9771814cf72cc5ba72bb208863cb

Observation 74079e33-86fe-41f0-82d1-e538efce29c1 · outbound

This paper cites Multi-Agent Risks from Advanced AI.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Multi-Agent Risks from Advanced AI

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.939633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:55c88f2035fd13d9e95796f2c32ae84cbe6b90f7397a3b2093b707c98a366c38

Observation 8103005f-0ba4-451f-b42e-564e88c0f201 · outbound

This paper cites Dynamic programming for partially observable stochastic games.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Dynamic programming for partially observable stochastic games

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:80dbdabeb768cdd076a993b76be3d66cdbb331b2215f1f5b2b2b62c6465b1d86

Observation eaccb156-1623-4f79-93e9-2a856c7576e0 · outbound

This paper cites YC-Bench : Benchmarking AI agents for long-term planning and consistent execution.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents YC-Bench : Benchmarking AI agents for long-term planning and consistent execution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.942643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:08038cd19247182b42bd5caefcbf48720ed92ebc3cd1f02e018f04b4acc16a9f

Observation d42b5e69-1652-4223-9b57-095a253cbf05 · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Metagpt: Meta programming for a multi-agent collaborative framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:2d83ace721949028325dc7fec76bf0e5db4e7d7e8c2e3b4ae2a568f4832ff309

Observation 2954ecfb-c452-4f81-beb9-d58126a60d2a · outbound

This paper cites Position: Open-endedness is essential for artificial superhuman intelligence.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Position: Open-endedness is essential for artificial superhuman intelligence

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:052fea57b2359c07665eea648e95ddd2b3c2656a94fcb91bcef7317ffcef99d0

Observation 00930267-5a84-40ed-864c-64ba1f10a56d · outbound

This paper cites GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.846756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:b380fdebcc4297789991e33e2dd175e74f09b8b0f2b048b6cf0765fc1f7019d7

Observation 11c58ef6-bf05-4360-8048-055cc3656be6 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Measuring AI Ability to Complete Long Software Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-14T01:19:52.953392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:65588f65498cb9bb8b0d6149b2cc3e1a90fc4f8688783e049c278b2d8220657f

Observation cb55698a-3642-4f69-a387-ebad98254270 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Efficient memory management for large language model serving with pagedattention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:da8bcab9ecb1e879b0f8433be5ce04480ddc893bdcdd373486c796a34c90574d

Observation 3231cd7d-1a90-49ba-a6ac-8b504e7f6514 · outbound

This paper cites Scalable evaluation of multi-agent reinforcement learning with melting pot.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Scalable evaluation of multi-agent reinforcement learning with melting pot

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:72df079c850683766dc549c18d045b5d33b54ccc9a5f5f89927e2cee735b17f7

Observation a0b1e157-6a7b-47d7-ab2f-72a9963c013c · outbound

This paper cites Stateful active facilitator: Co- ordination and environmental heterogeneity in cooperative multi-agent reinforcement learn- ing.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Stateful active facilitator: Co- ordination and environmental heterogeneity in cooperative multi-agent reinforcement learn- ing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:84f1a7b2b6136f7674d6d792edf52775592708225474c08b89ba6aa1abfcb7df

Observation ea4881bd-221b-4779-b1f3-0ead8873bd02 · outbound

This paper cites Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:6316e97c560d686707adc07696b3b02ca17e9d843b3668d98313861a55568ae6

Observation 1bbad8e0-454f-40cd-8511-8eabfabaaaea · outbound

This paper cites The interdisciplinary study of coordination.ACM Computing Surveys (CSUR), 26(1):87–119, 1994.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The interdisciplinary study of coordination.ACM Computing Surveys (CSUR), 26(1):87–119, 1994

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1b268fd4471cd6f099cd64f43741daef82aa9e6292187fe6c099c8bb9c0a9d0a

Observation 98dcb019-0380-425d-b8e9-03780327c2e2 · outbound

This paper cites Craftax: A lightning-fast benchmark for open-ended reinforcement learning.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Craftax: A lightning-fast benchmark for open-ended reinforcement learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:5c6ef526cf17ddc717215a479490eebf99e847dcbc240354e19af8fc213d1ee2

Observation 3f21675e-7647-4bf3-ae90-000a2509a9af · outbound

This paper cites The influence of scaffolds on coordination scaling laws in LLM agents.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The influence of scaffolds on coordination scaling laws in LLM agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1a7067638ce1d99e07c9fa155953efd426f10f761e9173e124ff69bc0fb80345

Observation 4ddd6292-8c34-4cfd-a64a-695761e9a01f · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:72f85cec05a58d46b41f9732c1ba8079bee0098ea0694e8ea568a9a5a796a6b4

Observation 6db07c23-3a1c-4dfc-9b1b-f2cbd4fbba87 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:57:26.474953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:3add132449e500be7d03998644d4c47ad5a559cad7307ea16f7116cf2082b00a

Observation 7045ef07-23ee-41f0-960c-2a54d2f7f4a1 · outbound

This paper cites Multi- agent craftax: Benchmarking open-ended multi-agent reinforcement learning at the hyperscale.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Multi- agent craftax: Benchmarking open-ended multi-agent reinforcement learning at the hyperscale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.486842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:4865dfcc420c6ea0c8467cef84270361a4544844493f09b974c6bc926b641060

Observation 91de4d21-d495-49df-9d57-95f0bcbf6899 · outbound

This paper cites Introducing gpt-5.4.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Introducing gpt-5.4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:7a0e6aa36bfbc9ffaf035320eab776f54bd9b693a023d843dd0cc1aff10b68c9

Observation 3ff00003-78e0-46bc-aac7-a6fe72926969 · outbound

This paper cites BALROG: Benchmarking agentic LLM and VLM reasoning on games.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents BALROG: Benchmarking agentic LLM and VLM reasoning on games

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1af34b8c5d0c6c5495eff37b2e592098c8ac38914967a83cf1b70eb7478e32e3

Observation fbfde19b-9397-4bd9-8672-af4f9cac20f3 · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:d18921aecfaf759a2a4853eea92ab8c2ea91ffdffd089a92f571061f61800a36

Observation fec06b1e-d3ae-47ea-b106-0d6b4559e6b5 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents, April 2026.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Qwen3.6-Plus: Towards real world agents, April 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:4bc6b6b8dc1a4b72c5307aaebdcdec5a7e121e9b0667a61092aa5b9043d1026c

Observation 887ce479-2103-414e-8e5e-2161f6e6315d · outbound

This paper cites Jaxmarl: Multi-agent rl environments and algorithms in jax.Advances in Neural Information Processing Systems, 37:50925–50951, 2024.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Jaxmarl: Multi-agent rl environments and algorithms in jax.Advances in Neural Information Processing Systems, 37:50925–50951, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:cb0b41480e1a0fd67ca34ffda8c8de512c9f7f6481ba42e22b21f544ed224be4

Observation b8f1c05c-71bf-4509-ac57-355352c23984 · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:3627ae7a38a5b261b24c53a92d837ac3aa266b28f1b863fe61315b2e77452677

Observation 5c34c920-297c-4d37-a20c-7b16900e83de · outbound

This paper cites Harvard university press, 1980.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Harvard university press, 1980

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:21457b6c44ac63aa4ce0669ac85725a403164a575f58aecddeba17dcdaab5b2d

Observation 6fee729b-d76e-4edc-8263-13d70f66793c · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.Nature, 588(7839):604–609, 2020.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Mastering atari, go, chess and shogi by planning with a learned model.Nature, 588(7839):604–609, 2020

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:e44d0937970abdb6328902ef7fe4ad96a0d4a0eff55a63f36e60cfbf9d045943

Observation 2f382997-0879-47c2-9126-2cbad57a8a90 · outbound

This paper cites The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:08b237078b19b29bf7f25a2286cdbf2b5ca07459650bd1c4398df1b77f4417d4

Observation bee96ae7-7e58-4c62-b5d5-68ff8da78311 · outbound

This paper cites The illusion of diminishing returns: Measuring long horizon execution in LLMs.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The illusion of diminishing returns: Measuring long horizon execution in LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:a3e530fe133a44476741c98ddfccd5bcc0eb1547f0572480215a11a4c711b50c

Observation f25956c3-aaa5-4824-9764-fbe7fcad84bc · outbound

This paper cites Cambridge University Press, 2004.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Cambridge University Press, 2004

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:e195ed890bd1102ba14ef82c9f2568b00e4f8a2a60233c771ec6576837c637ec

Observation fae28ccf-f4c9-4696-a22f-fa22e3edcdd3 · outbound

This paper cites Neural MMO: A Massively Multiagent Game Environment for Training and Evaluating Intelligent Agents.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Neural MMO: A Massively Multiagent Game Environment for Training and Evaluating Intelligent Agents

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:57:26.480037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:d74dade205814d44f4361802feba88d321340ecd835c7290cedc9fc24efddf7f

Observation 40972b0c-6593-462d-ba9b-1334d739ad61 · outbound

This paper cites Neural mmo 2.0: a massively multi-task addition to massively multi-agent learning.Advances in Neural Information Processing Systems, 36:50094–50104, 2023.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Neural mmo 2.0: a massively multi-task addition to massively multi-agent learning.Advances in Neural Information Processing Systems, 36:50094–50104, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:9aa876a8f902ea3d15930f4685510cb6e741a5faca6ec990efa7d66f03764fbf

Observation 318fe340-a567-4ca8-915e-c2fc3385273e · outbound

This paper cites Collab-overcooked: Benchmarking and evaluating large language models as collaborative agents.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Collab-overcooked: Benchmarking and evaluating large language models as collaborative agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:312fad08babdaf2a1692f96838c2a731485bdabd98d55c071fd93a0071b63846

Observation 1e7e5c25-6426-41dd-aed2-1e926467d02a · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Qwen3.5: Accelerating productivity with native multimodal agents, February

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:e101b0fee55400c7208a0a130d305d8a320808f3b98f9b2fc0af712882d75774

Observation 19d16a19-f1bd-4a6c-88c6-095781dc2e74 · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:cb958ce85be1757b09ae648c59f9f9568955dca423c04f990dae4890d6a48d3c

Observation 9e23311f-45df-46a4-af54-b6879e8a4fdf · outbound

This paper cites Hypermarl: Adaptive hypernetworks for multi-agent rl.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Hypermarl: Adaptive hypernetworks for multi-agent rl

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:ebb6eb0cd5e58c1edd3417c3d48cc385bc87b136cfcbede09313c6aa9ea2a5e4

Observation 99b49b6b-a53a-45ef-9b26-ab105b49b0a2 · outbound

This paper cites Probing dec-POMDP reasoning in cooperative MARL.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Probing dec-POMDP reasoning in cooperative MARL

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:3f99ded41759ec6b4a33119d24f6b22d379e7e1b00acd7ba0141c6bea893d906

Observation 3f88a188-cc22-4713-a1aa-dc06d4ebe76a · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:a73f021cd2d559b0e6aabceb069978ff9b4238647aebf75abe15c503cee16218

Observation 2e385151-8669-4fd3-8c19-063af38846e0 · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:fe26f8412cb13e6f98cf0dee5d762578dec132af964f904d3f0d8d4e6255720c

Observation 12aef7d2-b64b-44cb-858d-10f963bccb3c · outbound

This paper cites V oyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents V oyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:a24219307095074e4c1b600f0b099190c79995a6270ca3c79fde3ef3b463171a

Observation 54db5e46-b1ba-41bb-82a2-6991ed0d43bc · outbound

This paper cites BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.483468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:18334512178936e3ab4274a578e500aa7d6e21f0674083491f5f0bc88842eed3

Observation 2fb866c7-3830-4bb4-b3fb-a9ecb1dd1d5a · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.489813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:81a372942d59893167a9a6b557a3217efb5813a07f6c62472edfad2971d7e313

Observation cb41eca6-8e69-42cd-b10a-5246adfd5b41 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:57:26.492973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:056fecb9104bd695608068603a478da0fbcbfde1d944cd35762d3853f2771db0

Observation 268c12ed-7f4c-4f93-a7f8-a60d6bbb6580 · outbound

This paper cites LLM- powered decentralized generative agents with adaptive hierarchical knowledge graph for co- operative planning.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents LLM- powered decentralized generative agents with adaptive hierarchical knowledge graph for co- operative planning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:c344349aa67e083f3e2892bad8b04e9ca5d1ce4feba3d36458650a223d2cd00b

Observation 519dafca-3ff6-448c-a57a-f305f7f4144c · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents React: Synergizing reasoning and acting in language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:c6c3f880145bccdc901a112bcdbc5e1efc2c2a9af63a351e7d46df82d6ecccf0

Observation 043edb4e-f88b-4561-a60c-2f835456b35f · outbound

This paper cites An Efficient Open World Environment for Multi-Agent Social Learning.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents An Efficient Open World Environment for Multi-Agent Social Learning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.477512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:4e830790733209329b742b905cb2ec1df1ad278478cfe4ba1c370f77252a58d4

Observation 65f49f67-45dd-44b8-acdc-022a66b3e71c · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.Advances in neural information processing systems, 35:24611–24624, 2022.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The surprising effectiveness of ppo in cooperative multi-agent games.Advances in neural information processing systems, 35:24611–24624, 2022

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:857e38080c7e73956ed2aa9c36cfd9da10b552aedf5b9a3b96ad73a3d4e426be

Observation 6cc8f711-ba8f-4f87-abde-556685233499 · outbound

This paper cites Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu, Yang Su, Lianghao Deng, Xudong Guo, Chenxu Lv, and Junyang Lin.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu, Yang Su, Lianghao Deng, Xudong Guo, Chenxu Lv, and Junyang Lin

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:25.945763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:9d7862923ad0a0ef811dd35c007fa019924a4b1a0bd16727827ab75d0ff08153

Observation 23c1e0b0-0674-4e00-888f-1ae0fe52be74 · outbound

This paper cites AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:07:25.841029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:3f21e441809900741727e7da8ac5ffa7c5053318933545acb58eada6eea18013

Observation 5be93d94-2e6f-484e-8c63-3d311130b72f · outbound

This paper cites Multiagentbench: Evaluating the collaboration and competition of llm agents.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Multiagentbench: Evaluating the collaboration and competition of llm agents

Reference 72

Resolution
malformed identifier
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1d9f2cf1998ab3cd286ed116e35ff7a359bb736a70f025dc72332c27d76f167f

Observation c93ba9b5-378b-43f3-bcaf-26a3b6d9c128 · outbound

This paper cites By default it also optionally includes the full action catalogue with natural-language descriptions, game and coordination mechanics.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents By default it also optionally includes the full action catalogue with natural-language descriptions, game and coordination mechanics

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:04e2b212623a5a7025f4f00d57bac9f0e73006250dfa0f73019610460a6c5644

Observation 7aa8abb5-ca0c-4529-9124-7295b191e3d1 · outbound

This paper cites Observation from k step(s) ago.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Observation from k step(s) ago

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:d3cd36bc447703e186e4fb2d41ca29afeb888f45acbbcee7fe7197cb7a7bb408

Observation f96908de-3996-4cbb-abe3-a9e558bcc3d1 · outbound

This paper cites The agent is then prompted to respond with an action.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The agent is then prompted to respond with an action

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:9cfe68915f2c4de3a5220eb294e6b50d13d0311eb21dbba4e89c5076d0dde8e0

Observation 686764f3-2877-49a3-a676-07e2aae58d08 · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:189a3081bfa764b2e1565eb329a35b00043b7cfba6f22b03365ab8f9aaf799ec

Observation d25e8562-2e63-4b6b-bd42-07e8474d4c1e · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:d9803ef0401340af4f338106c6d34e4f93be4279927101e60678ecd1418e9dfd

Observation 9fb18d4c-5910-40fb-913e-01172986b31e · outbound

This paper cites The ladder only becomes usable after enough monsters on that level have been killed.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents The ladder only becomes usable after enough monsters on that level have been killed

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:a654f64ddc4008e3afa0d02f074e81b02ee81d25ed7f62b0c95503b6164d6aaa

Observation 222ddb51-b827-4e46-92d8-581257202220 · outbound

This paper cites an unresolved cited work.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:01a93ef8d7e9bdaf4c186c937b263b83030c3f2a32cff9d0e2637386ebb0962f

Observation 4b25005e-4955-47b8-af2b-951272fa3635 · outbound

This paper cites Teammates can only act on what you tell them.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents Teammates can only act on what you tell them

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:1e51cf43591b51d9b3130c9ee475fe58ad3c36d5e9ac18cd075f0e9a612581ea

Observation fe74fa9e-167e-482e-8370-86032dbbc114 · outbound

This paper cites TN: ” or “TurnN:.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents TN: ” or “TurnN:

Reference 81

Resolution
malformed identifier
no resolver link, observed 2026-06-27T19:18:56.039587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:88a00142943e890a3f2a8b8c2c12d5c2cf9da5c90edbce831b7a21c56fc959a7

Pith citing papers

Observation d040f47b-37ae-4a37-a8ee-81dee79bd2f7 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T15:28:46.299124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:25:39.314247Z digest=sha256:f3c1d6c75e26f5394d8842d663ecf3e84930d33f8616e121652e9c7b4449799f