Pith. sign in

Paper Citation Record · LEDGER

BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2408.15971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15971 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:49.681539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:57:26.233385Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e535dccd-c87d-439e-bf4c-2d038c92299b · inbound

Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game cites this paper.

Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T15:20:07.197962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:20:07.197962Z digest=sha256:d585509ec964f6a1a9fecac82d95e2af584cb25dcd3fdbd53361c72c19090478

Observation bdeb7504-dd83-401b-9027-98444f02d6a4 · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:57.912351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:d29e38519a4c94ac7d71b8ef9b89817a1b5dcd94619dc89b10e8a6c054fc6639

Observation 54ca6850-b338-4c94-b88e-acb57778a5e9 · inbound

Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant cites this paper.

Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:49.681539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:21:49.681539Z digest=sha256:ca7068f75d8bf90d63876bd4ccd3c416e4209221ab1933643d67c8d1b0d927b8

Observation ea404cf2-4c3f-4944-91f5-ceac3ea54560 · inbound

MAEBE: Multi-Agent Emergent Behavior Framework cites this paper.

MAEBE: Multi-Agent Emergent Behavior Framework BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:47.229043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:14:47.229043Z digest=sha256:8b90bd31212a486efb3fbc00de2d2e586d6beec9c16215034515461f884cc5a2

Observation 9583adb3-51b2-41e2-bde5-666e154bd097 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:58.050301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:58.050301Z digest=sha256:39947f901e70c13dba511c8f91a4b25e7b08964f5f9047b497a1918bdaba7989

Observation bdc7057b-65d2-454b-a014-bacf7035e2d2 · inbound

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems cites this paper.

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:58.873471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:58.873471Z digest=sha256:0c083938f486b46e8d6554bfb727b0fc7760b4678be6226a89543292f949b0e2

Observation b26edf21-e77e-4e9b-ba80-c3c40b23c3bc · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.709367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.709367Z digest=sha256:1a2d4cbb20b9517bdbbb3c7d6bc92d8afe81970f228ef00ebc1f027a2c1eac05

Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.475102Z digest=sha256:adf255a2f3337b77276b4ed70266bc10971a212f34a375eede2347264b26a3c2

Observation 5818bfc3-d269-4ac5-a028-690f96bb1ba7 · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:58.980429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:58.980429Z digest=sha256:03234f25256cfba9b5dc834b4d2a4a926ca6394eaa841bf7fe0c5e927e7e6453

Observation fea5e3e3-1c7c-4e94-8f15-652267ca4ce0 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.347355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:fc198999c47d80b2ea4fb0c24a3150989a32b105d0d1a5ae5208519f7fd1f008

Observation 8810323f-30ad-4717-974d-a0aca85c4436 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.817356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:ebcfb98c5a2bdfd9b67765c8de7ee65c26b72211380d27f3899519350be1f557

Observation 6c817dfd-babd-4460-a9de-0eb0b64aecfe · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.876999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:27fbaa19c27596f814bea3e40cb36ed6c6745ef806b02fcce4ee174e5a3be39a

Observation 1939d673-476e-47e9-9aff-740c74e61895 · inbound

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest cites this paper.

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.074202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T16:37:58.860183Z digest=sha256:a59b60cc3c6c3b28d5b68f5fc9a38f95c6bd8c6c65f35fe452b454989bc64aa8

Observation 878cc54b-7d8c-4338-a0f8-f425dfa61bde · inbound

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling cites this paper.

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:23:37.908608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T02:19:10.912487Z digest=sha256:c9f894533fec3373ad48263a3e143ef513cb2b4e2f7e5323ab4da01525b0be37

Observation 6f036a5a-a79e-4408-ba1e-50cc7b7eb24f · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:58.183231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:78ec86469be9ff8a899384315731d685ce5712bc267c93ccfc254e7abdfa1b83

Observation a0f51e4a-ddc6-42f5-bf98-e83312e239df · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:52:39.862066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:1bf4a0d1b22c4b20d1c6eb43711802364ef40a8f6a1c757c68e769b894309c15

Observation e417114f-0aa8-484c-b6db-7c7270a90df1 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.125624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:18ed3aee0528d8eae79ec97bff143452cbf6f5c209347a289b8ae5aade25e2a1

Observation f55197ef-28be-4ac5-9d2f-ff2cc0a035a1 · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.884053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:15942c2b2605421c83f4dfc20a3a978053e9c1b3ad0be76d140c14583ffbf670

Observation 54db5e46-b1ba-41bb-82a2-6991ed0d43bc · inbound

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents cites this paper.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.483468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:92daf92dda2d66325e502e1ab485d4ce74dadc5cc11080e4f10dda4f0c82fc77

Observation 034bce39-cc41-46be-96ed-0315da7095d4 · inbound

RAILS: Verification-Native Clearing For Agentic Commerce cites this paper.

RAILS: Verification-Native Clearing For Agentic Commerce BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.234866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T18:35:13.781966Z digest=sha256:06599adbdeeaa7e086b1364548620ad9f5b2d9e73838c446833365693cc26837