Pith. sign in

Paper Citation Record · LEDGER

Why Do Multi-Agent LLM Systems Fail?

As of 7 August 2026, this Paper Citation Record lists 100 of 123 outbound references and 100 inbound Pith citation observations for arXiv:2503.13657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.13657 v3

Coverage vector

measured 100 of 123 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:42:57.561468Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 163 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:12.357777Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 123 outbound references displayed

  • verified exact38
  • verified fuzzy40
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch15

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 61ff6413-ed8d-400f-b145-3fc267efab06 · outbound

This paper cites The Russian Messenger.

Why Do Multi-Agent LLM Systems Fail? The Russian Messenger

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.086478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:2bc7a43e53a1f11dbef5a43f5f531612eb5912f843762295e3cfc102441cd4a3

Observation 30d73cb7-8741-491c-8aad-02e3d9735fc5 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Why Do Multi-Agent LLM Systems Fail? Gorilla: Large Language Model Connected with Massive APIs

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:57.927341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:c90494a804941f86d8a23f424ad5907e3ac2237319161e9bbf5002e001f2e42e

Observation 739810f5-f2de-4d15-83a8-f8b5debf6b44 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Why Do Multi-Agent LLM Systems Fail? MemGPT: Towards LLMs as Operating Systems

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:57.953690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:48a16db39c487be258b2e9a51404f64c565a003553afc36a1228b032db6beef5

Observation 2be0b41e-9503-43cb-8f15-aa85fe0cdfe2 · outbound

This paper cites A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), March.

Why Do Multi-Agent LLM Systems Fail? A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6), March

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.222791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:c07c7631ff77236c79eb4f7e7887ab822dfa0148638132a5daaab827cc358bda

Observation d8fc0a52-fb7c-4315-8e44-d2d8b16cf673 · outbound

This paper cites Frontiers Comput.

Why Do Multi-Agent LLM Systems Fail? Frontiers Comput

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-12T05:42:57.789285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:2c6134ede40b3e6619f300793bb0889d11199b2c30521299246ea181d7227c9a

Observation 289507db-7e82-40aa-8f49-fd4c3ea96ca4 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Why Do Multi-Agent LLM Systems Fail? ChatDev: Communicative Agents for Software Development

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:32:47.766951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6156d9cff6fe47f10dc4de635d13404bf7e6c4a5d52f1f74d8ef07da7d3954e2

Observation e7792d97-0692-4dbc-a7ee-a2de045678d0 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Why Do Multi-Agent LLM Systems Fail? OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:57.822350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:83208bbb107d1b334f2563c4d3d8122817f1fc4d5645aac78f9e5253f039a319

Observation ff7b9938-c3a4-4e05-af23-08a973b8e3a6 · outbound

This paper cites Towards an AI co-scientist.

Why Do Multi-Agent LLM Systems Fail? Towards an AI co-scientist

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:57.840350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:3c8755ed00f1363c312f2317b091489307f5d8eb1b49c533faf4dbf1f962b478

Observation 6f61f12a-a9eb-4ac2-9852-6d115ab39dd8 · outbound

This paper cites Bulaong, John E.

Why Do Multi-Agent LLM Systems Fail? Bulaong, John E

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.139101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:049100c92b32ddadefa79876b0c07d3c29a40ccb8e6c307b6fbc88b745af62a3

Observation 68060de7-0d56-463e-af98-e41ec3c3486d · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Why Do Multi-Agent LLM Systems Fail? Generative Agents: Interactive Simulacra of Human Behavior

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:57.864354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:5be71690da23fae1dd88a5a49485644b05e8a80ac1cfa46e87f04eb138382889

Observation 1ab56c37-0cca-49d3-9e3f-6e1f88b44fa2 · outbound

This paper cites Openmanus: An open- source framework for building general ai agents.

Why Do Multi-Agent LLM Systems Fail? Openmanus: An open- source framework for building general ai agents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.958156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6a6d82141a7c8761dd934bacaadc8d373cae341e5a0bc0445c0b4a17dde437e2

Observation 197f52d8-f6e5-4328-a48e-8a9bbd61eb46 · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

Why Do Multi-Agent LLM Systems Fail? Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:49:38.916324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:1c254b5a626ff7ba44f14dc17bfa28827e0107e160ffdc81287d1b068a8714c2

Observation 5cb7e049-d3a9-4696-9a7a-5284781f0c53 · outbound

This paper cites LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead.

Why Do Multi-Agent LLM Systems Fail? LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.005346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:4d0b326ad4b0c39df16415f79b6ac6fbf004599bc1808392a045944f31575d43

Observation 6e3ded65-c98a-43eb-9c1d-9d77336f8bb6 · outbound

This paper cites RoCo: Dialectic Multi-Robot Collaboration with Large Language Models.

Why Do Multi-Agent LLM Systems Fail? RoCo: Dialectic Multi-Robot Collaboration with Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.031352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:456a4a5d6867c1cf727537c8cf58f43df03567013f19f0c5fe12386ad7765de0

Observation 8c030740-caa6-4091-8189-afef68b8e6ab · outbound

This paper cites Building Cooperative Embodied Agents Modularly with Large Language Models.

Why Do Multi-Agent LLM Systems Fail? Building Cooperative Embodied Agents Modularly with Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.054690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:11e9e4f8f76660fdb78712bdd78d06583a9f2a3918b93849f27a10234127b07f

Observation 22699795-244b-4224-a549-2d2c103e8fa8 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Why Do Multi-Agent LLM Systems Fail? Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:58.064340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6f98ba49294b7fa1309240bf9fbbbf178ef1e05b94645f8fc73947383ca38467

Observation 69f7bd8d-e1e6-4c88-bf55-74b2a0f09ac9 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Why Do Multi-Agent LLM Systems Fail? Generative agents: Interactive simulacra of human behavior

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.989078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:a7a1b1e55ebdd2da4eaea5fb1b19d80038c8199afd457508cb2764ddbf0a3524

Observation 819c299a-6a67-447a-80dc-cdd197ee4c6b · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Why Do Multi-Agent LLM Systems Fail? Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:58:59.611495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6d87718dea97bdca14db05fdbc0d4da495f9380e05b305e6683cba41fb798e28

Observation b9241fc1-66a4-4191-8246-69cbb53f4f58 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Why Do Multi-Agent LLM Systems Fail? Agentless: Demystifying LLM-based Software Engineering Agents

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.120345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:31c07f75a590ca151d620a81c3e9f1a102e35ccc5806cd7180cc199923404081

Observation f88d505d-a470-4893-bbeb-191600d0fc84 · outbound

This paper cites AI Agents That Matter.

Why Do Multi-Agent LLM Systems Fail? AI Agents That Matter

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.133915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:fdc607e20a04980086cacc7d1bef6ccfa5a438023ee133461e16d3ced914de8b

Observation 76aece85-d3b4-4ef2-9e67-abde5836d2f1 · outbound

This paper cites Glaser and Anselm L.

Why Do Multi-Agent LLM Systems Fail? Glaser and Anselm L

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.016164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:271c4cd829dd142db63608e0b0fab1f43009789ed0b76aee5fe0932a466c57fa

Observation 9315eda1-2d6a-4f23-9543-8901b8611b92 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Why Do Multi-Agent LLM Systems Fail? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:58.153412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:1e9bb5e0aac4f248b91595dfd9802d6746052e6cf68516c647b83cb5736ab219

Observation 43bad753-8a20-4036-a87a-196bb82a0957 · outbound

This paper cites Agent workflow memory.

Why Do Multi-Agent LLM Systems Fail? Agent workflow memory

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.028666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:2f6db5194e004b45191962e0f7f1e84685be353a1eaa09643b3a30e3a06303b9

Observation 1f4fec5a-73f5-46b0-ab73-9bf82acd8315 · outbound

This paper cites Agent Workflow Memory.

Why Do Multi-Agent LLM Systems Fail? Agent Workflow Memory

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:51:28.546876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:ddd20643d09d7c82fd133bc18e5bc3cec814e9bc1b1f327fce4445a53dc2403f

Observation 6bf491c2-5f10-49ad-adcc-6822fcc4f1e7 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

Why Do Multi-Agent LLM Systems Fail? DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:58.196027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:30a603b8379aa4ae3f43355ba284833f54e9791839f28f88841be45c1f14a3a9

Observation 02f7ae35-578a-49de-8868-30e0af2f459c · outbound

This paper cites Stateflow: Enhancing llm task-solving through state-driven workflows.

Why Do Multi-Agent LLM Systems Fail? Stateflow: Enhancing llm task-solving through state-driven workflows

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.042463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:165e4c9e3be609d561abcd5b748888936e74d5651a1a159debd3a21d50728cde

Observation ec8f5d8f-f827-4222-9d93-c188ce57ff1b · outbound

This paper cites LLM Multi-Agent Systems: Challenges and Open Problems.

Why Do Multi-Agent LLM Systems Fail? LLM Multi-Agent Systems: Challenges and Open Problems

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:23:05.283434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:af666faea183ff56e61e8cf96a2913367d38f2f27bc790167be0119f6b58e421

Observation fbbc47ce-ded6-422f-a4fd-2d44427b6b2e · outbound

This paper cites Multi-Agent Risks from Advanced AI.

Why Do Multi-Agent LLM Systems Fail? Multi-Agent Risks from Advanced AI

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.257297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:ec2d53af70d19590eafba38fb9672c4aafc3ea2d67bf16300a6232d043dcae14

Observation 0cb7e8d4-7b29-481a-848d-ce047df9db41 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations.

Why Do Multi-Agent LLM Systems Fail? SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.066110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:5b58b4db115ce8bc41e78e0da92610d4a0dad400723a3a6db5a4e182a92dbbba

Observation e4b2aea3-1b1e-48ea-96ca-70234cfbe3da · outbound

This paper cites A Survey of Useful LLM Evaluation.

Why Do Multi-Agent LLM Systems Fail? A Survey of Useful LLM Evaluation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.268435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:3d7731adc050c97aec51241f8978cb282115b83a4cec19d952a9223ff5e5fc52

Observation bdeb7504-dd83-401b-9027-98444f02d6a4 · outbound

This paper cites BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems.

Why Do Multi-Agent LLM Systems Fail? BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:57.912351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:31700d93ea5b80eaa829a878de35aeda46ffdad884edc5476f4fb529e8162e62

Observation 2bf167ad-24cb-45e8-a4e5-9a886f6d5e15 · outbound

This paper cites Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control.

Why Do Multi-Agent LLM Systems Fail? Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.301175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:dab6ff6f0b1055e6064e1d23c82901805d891ba1714a56b5a1464d6da2c966c2

Observation bfe1c119-9768-43f8-b818-46f5e005ea21 · outbound

This paper cites Benchmarl: Benchmarking multi-agent reinforcement learning.Journal of Machine Learning Research, 25(217):1–10.

Why Do Multi-Agent LLM Systems Fail? Benchmarl: Benchmarking multi-agent reinforcement learning.Journal of Machine Learning Research, 25(217):1–10

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.098835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:67ca7ea2dd5b7dc4e74647112a3518df703848b346e94a6a2cfb03f53a357aed

Observation d363ce66-e698-482b-be4a-0fa0f8c33924 · outbound

This paper cites TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft.

Why Do Multi-Agent LLM Systems Fail? TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.311345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:21261db75342537aa5adc4c42d57b4e7e9c446e7291dd4298a61c4cfbaaefc0a

Observation 1e092785-dddf-40bf-9b52-4e4d4e51c9c4 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Why Do Multi-Agent LLM Systems Fail? Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.329329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:69c8f379d2a2a989c8076fc87134a959b0c2b5122f8dee50a454062c0c1795ed

Observation ee79bb8e-fe6e-4747-ace5-a05f1fced8e6 · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.High-Confidence Computing, page 100211.

Why Do Multi-Agent LLM Systems Fail? A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.High-Confidence Computing, page 100211

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.119414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:ba5e28110795cecd1966ea3f0c8e63375525da011bf3ff5a61d0b1732f31ac18

Observation 534af70f-8c47-4fd0-9529-0f42c07a51b2 · outbound

This paper cites URL https://www.anthropic.com/research/ building-effective-agents.

Why Do Multi-Agent LLM Systems Fail? URL https://www.anthropic.com/research/ building-effective-agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.122634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:7a508803ac8cb7f1b274f25c09c644ee83be8a6ff37411231aa028dddac35e89

Observation f67b68d8-dae0-444b-bfcb-58605e6c90cc · outbound

This paper cites Specifications: The missing link to making the development of LLM systems an engineering discipline.

Why Do Multi-Agent LLM Systems Fail? Specifications: The missing link to making the development of LLM systems an engineering discipline

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.337875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:031d262d4f589c0d57980ffb4e454c207c75510fadc272c01ff5b6565ed0af8b

Observation a879c327-8242-4f7b-adf4-75bec7d01a31 · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:59.132387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:429dd798fb442c4016b2d82d59a20c392c9ef4aa2080e5df5a3f8c60f11d3ba5

Observation 488ea8cb-c884-4d73-b2a8-6a8dc9c70931 · outbound

This paper cites Knowledge-Centric Hallucination Detection.

Why Do Multi-Agent LLM Systems Fail? Knowledge-Centric Hallucination Detection

Reference 40

Resolution
metadata mismatch
doi, observed 2026-05-12T05:42:57.775599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:9e6a8d87922ad459b65483cc2b5d0057c6a580da8f28df3d313191b8d081e3c4

Observation c7000d0f-d0ec-403d-ac01-376e7a799e33 · outbound

This paper cites An empirical study of code generation errors made by large language models.

Why Do Multi-Agent LLM Systems Fail? An empirical study of code generation errors made by large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.146953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:deaa78ec773573d5bcd68f046fad4e01f60f174c3af3b6a576fdb521b551e1fb

Observation 7f3758a2-5fb9-45e3-b4fb-4cacc9f103ee · outbound

This paper cites Assessing and Verifying Task Utility in LLM-Powered Applications.

Why Do Multi-Agent LLM Systems Fail? Assessing and Verifying Task Utility in LLM-Powered Applications

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.371011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:503011dd78ffc7378a1eae1781996c2c83a5bf9345f03b9e2d57367147d02eb4

Observation 0481156e-a12a-4d0a-872e-0502e1affda4 · outbound

This paper cites Interactive Debugging and Steering of Multi-Agent AI Systems.

Why Do Multi-Agent LLM Systems Fail? Interactive Debugging and Steering of Multi-Agent AI Systems

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.388577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:191ddcc3d010f90bc041daad86025b97ef4112cbcde7344eb81bcb30ed3e5ee7

Observation 3d8cb982-91c4-47b9-a72c-c8dbaaee7738 · outbound

This paper cites Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems.

Why Do Multi-Agent LLM Systems Fail? Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.410764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:d623a3a36739b2df927b08e57991fb8041c73dade80ad0b388da7f023f183a4c

Observation 6672edff-68b5-4e42-ae70-e127df2829c5 · outbound

This paper cites Theoretical sampling and category development in grounded theory.Qualitative health research, 17(8): 1137–1148.

Why Do Multi-Agent LLM Systems Fail? Theoretical sampling and category development in grounded theory.Qualitative health research, 17(8): 1137–1148

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.175408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:ed1fa21eeb26b9d19218f1fba088d80164cb15bca8487752a69a74c58be27111

Observation 3822f4d1-a368-4adc-adf1-860a1b49dc64 · outbound

This paper cites Open coding.University of Calgary, 23(2009):2009.

Why Do Multi-Agent LLM Systems Fail? Open coding.University of Calgary, 23(2009):2009

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.180694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:27b70f79f85fce66a93407eff7ef6f90b41dcf96f06c789f87d7db891534f157

Observation 93dc3548-4106-4b5f-8a8d-aa4bf1ec12ba · outbound

This paper cites Manus.https://manus.im/.

Why Do Multi-Agent LLM Systems Fail? Manus.https://manus.im/

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.189479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:1fa9f2ad33d9ddde802d0c22dbdde4b9ad24a93068460de54350353e75da30b8

Observation cdf93fb5-9a18-4f3a-bc2c-5f97e5cb84ab · outbound

This paper cites Model context protocol: Introduction.

Why Do Multi-Agent LLM Systems Fail? Model context protocol: Introduction

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.194925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:f4ede52f724dd1d3c6a1228e4e0c4ba98a69eec4803fc92e1318c131f05358ef

Observation 46d1ea93-1e51-4569-8cba-909fad376b42 · outbound

This paper cites A2a: A new era of agent interoperability, April 2025.

Why Do Multi-Agent LLM Systems Fail? A2a: A new era of agent interoperability, April 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.199428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:9b09bd6694936dbb6a3ed123cb63bcae8a81848eecca689afd5c8b952e3d742d

Observation 5fd00d40-91c8-4e84-a229-a570db94bb0b · outbound

This paper cites LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models.

Why Do Multi-Agent LLM Systems Fail? LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.424334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:92b13366a480e79b74f8905879028d4aee309eda99aa30f3be4448f40236f590

Observation 127dfd1f-2b8b-47aa-a8d6-5a92e26188a7 · outbound

This paper cites Princeton University Press, Princeton, NJ.

Why Do Multi-Agent LLM Systems Fail? Princeton University Press, Princeton, NJ

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.216199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:af73ab3a2346902c16dc91b3a01c027c096f2556db14e1b1b61c1edcf519ebae

Observation 8e2aaae6-c585-4bfb-834b-3efccca5e045 · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 52

Resolution
verified exact
doi, observed 2026-05-12T05:42:57.760624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:7537ee3d8fd59df06e31ad40c553e47dbea6147af92bd8e7f9f66e61c3c28dfa

Observation b62c6a9a-f836-4065-a7be-645c0eaffc68 · outbound

This paper cites Reliable organizations: Present research and future directions.Journal of contingencies and crisis management., 4(2).

Why Do Multi-Agent LLM Systems Fail? Reliable organizations: Present research and future directions.Journal of contingencies and crisis management., 4(2)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.239516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:f6210ac55439bb2adad6425e3ead778dbb4b067e96e8eb616e0848234612d72f

Observation 3ce33145-a8c9-4b76-aa25-9824be35f6c3 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Why Do Multi-Agent LLM Systems Fail? MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.437347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:d05109044bd781cad4a3bcb6f7e218c6dcebae816a83d2c92fa45db0bcb52c89

Observation 2a7e4a7b-e169-439b-bfde-44c7809ac945 · outbound

This paper cites HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale.

Why Do Multi-Agent LLM Systems Fail? HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.484728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:03d8725fa1b3fe203185612bc7f26100da42bf75ce8bf266efac0123efb55214

Observation afc58cd6-4637-44a2-b700-8e99d19c00ed · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

Why Do Multi-Agent LLM Systems Fail? AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.505343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:a4f4ee143ed9ad5197ea551a37335c11ef09072380277bf77b13a2e0960878e7

Observation 90a7b087-48b1-40d9-955c-6c37fe9252d3 · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversations.

Why Do Multi-Agent LLM Systems Fail? Autogen: Enabling next-gen llm applications via multi-agent conversations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.785220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:e81e779801a3edb1a37459459caf6c4308884a888b153a881fd41b9d187c85fd

Observation aa313a61-6bed-4f7d-860a-f09f71a19a6a · outbound

This paper cites Chatdev: Communicative agents for software development.

Why Do Multi-Agent LLM Systems Fail? Chatdev: Communicative agents for software development

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.791055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:ae29077c838c9c98023aa106ebcc5bd783a58df2779bbfb77ef296ffa1db574a

Observation 392e6512-3e20-4ab3-951b-8c05998b2187 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Why Do Multi-Agent LLM Systems Fail? AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.520447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:88862cc04330a98832ec3bcb75c6f23a02dc92373febe0b557a1fe1b121384c2

Observation b982bc44-7e8b-4abe-a956-ac258d239819 · outbound

This paper cites Does Prompt Formatting Have Any Impact on LLM Performance?.

Why Do Multi-Agent LLM Systems Fail? Does Prompt Formatting Have Any Impact on LLM Performance?

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.542908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:dae5361f2dd29bc460f996c84eb55b9899f171f530941198121bc5643ebe0f2d

Observation bf6b9f6b-54f7-4675-8610-0ef902b0a674 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

Why Do Multi-Agent LLM Systems Fail? Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.571343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:c6166a9eff182743f14f84411d45a986202ffecbec317d8a8aed5084a753d98a

Observation 9b5fece2-e11c-4c30-b52f-ee29e96660ed · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Why Do Multi-Agent LLM Systems Fail? ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:03:18.868992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:22230e455bb692f0ddfc9f0ec819db4d26a21bcadce810e31c70f5893ad16d5e

Observation ab1bc6ca-0cc0-4f37-84df-077363b012d2 · outbound

This paper cites Large language models are better reasoners with self-verification.

Why Do Multi-Agent LLM Systems Fail? Large language models are better reasoners with self-verification

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.846353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:7e4bdb2dd61696049b3633bc5f34e453c6132d7a23670eb7473d24637289684c

Observation 2c3592b1-55c2-48d3-be7c-479ce76efd19 · outbound

This paper cites Langgraph.

Why Do Multi-Agent LLM Systems Fail? Langgraph

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.858079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:7e9319ff074e00c46a99f0b4ace4ed4fbd56070689e8e916aca49ab21cf76eaf

Observation d930b616-b69d-49c7-9207-6b23c7e93be3 · outbound

This paper cites Building effective agents.

Why Do Multi-Agent LLM Systems Fail? Building effective agents

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.869595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:e0465913929caf7341078cb4ef0231939a833f6beb2a6e4fd6d8ede5a10975bb

Observation c2bb0d4a-dc99-4e63-a4f8-3ed65d7d1f0c · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in Neural Information Processing Systems, 36.

Why Do Multi-Agent LLM Systems Fail? Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in Neural Information Processing Systems, 36

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.874105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:97562b3d2eb3f96e2f067d0ee90fb6cfc424a75fd18d25a8179646783dfd2129

Observation 23d7d688-411c-4e32-837b-b24e4a024ac6 · outbound

This paper cites Improving LLM Reasoning with Multi-Agent Tree-of-Thought Validator Agent.

Why Do Multi-Agent LLM Systems Fail? Improving LLM Reasoning with Multi-Agent Tree-of-Thought Validator Agent

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.611439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:2a9c8cb0e68a9d53592758883beef014adfed3a3622e62a78f29158ce9973b26

Observation 1c9cb407-270c-4f22-8b86-125102fb6029 · outbound

This paper cites Towards Reasoning in Large Language Models via Multi-Agent Peer Review Collaboration.

Why Do Multi-Agent LLM Systems Fail? Towards Reasoning in Large Language Models via Multi-Agent Peer Review Collaboration

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.623989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:41938852699c1330b87726e2eb52b92a59b9244111f22e0f8c90313b4c1c119a

Observation 8ea293d0-7f88-42ab-89f9-dceaf8face7a · outbound

This paper cites Inference scaling f laws: The limits of llm resampling with imperfect verifiers.

Why Do Multi-Agent LLM Systems Fail? Inference scaling f laws: The limits of llm resampling with imperfect verifiers

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.634717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:4459a09d307471a9a4a4efa47824d9493f54ef6385124b54d8fc2a2644103b77

Observation 6a29d052-7294-4fe4-9f35-4d8491d8d70f · outbound

This paper cites Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems.

Why Do Multi-Agent LLM Systems Fail? Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.663859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:e8e27ecbe95e8876fb8e157fc7352e276690515458c07d8e1c7175466850e211

Observation 8b3bcd12-303f-47a2-9af5-9897f9af9525 · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

Why Do Multi-Agent LLM Systems Fail? TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.684087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:574af2b8e80327fefc146d324ddf82c6fb38c49174e0fede76f0ecbc2b82933a

Observation a514a91a-489a-40c0-8adb-29ee3ebcd8f3 · outbound

This paper cites Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback.

Why Do Multi-Agent LLM Systems Fail? Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:02:52.531251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:2a257527d489d780763724eed5b87ee538cff011dd6ac70820605336dec51833

Observation 065f37fa-6838-478c-971b-b18d289bc565 · outbound

This paper cites Leveraging Abstract Meaning Representation for Knowledge Base Question Answering.

Why Do Multi-Agent LLM Systems Fail? Leveraging Abstract Meaning Representation for Knowledge Base Question Answering

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.712681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:566e407827e3507e043757f99c0307670d1928c4594c1e58ac1f0665320f64fe

Observation 425dc28b-ecb6-411e-b9de-30d066d44d2d · outbound

This paper cites A survey on llm-based multi-agent systems: workflow, infrastructure, and challenges.Vicinagearth, 1(1):9.

Why Do Multi-Agent LLM Systems Fail? A survey on llm-based multi-agent systems: workflow, infrastructure, and challenges.Vicinagearth, 1(1):9

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.917941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:4e335630b5a916c726abd954f74535a7a084391d511aaee1ccdf92a579928d33

Observation 8181f267-f790-4557-bb92-af927ece9610 · outbound

This paper cites Multi-agent graph-attention communica- tion and teaming.

Why Do Multi-Agent LLM Systems Fail? Multi-agent graph-attention communica- tion and teaming

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.922017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:b2625f2bf97083759ac4eb9e60bf1f8ef8815cbb57dd4c1da9c83dcf7dbf5941

Observation ff5d198e-6fa2-489a-b506-c8765befbdba · outbound

This paper cites Learning attentional communication for multi-agent coopera- tion.Advances in neural information processing systems, 31.

Why Do Multi-Agent LLM Systems Fail? Learning attentional communication for multi-agent coopera- tion.Advances in neural information processing systems, 31

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.927959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:98037de26c0bf24b173faf44cedcc6e5bff9d7ecad683379017b8a3bcfce60bc

Observation 1bfa97c9-0bc1-4a63-91a0-98b83f967c97 · outbound

This paper cites Learning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks.

Why Do Multi-Agent LLM Systems Fail? Learning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.717697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:5f0aa533887a0798b83ed2d63b8612d6f331992851eb4692654e0517e2a15577

Observation 7b5bb5fc-070c-4cc4-93e7-ad42e42c0134 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624.

Why Do Multi-Agent LLM Systems Fail? The surprising effectiveness of ppo in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.939798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:4e9d5a04f33f069041586023ad889f8a9552b29a3c95b83305bf7a4430013ecc

Observation 8e506054-92e5-4d67-8520-a2602e211a5f · outbound

This paper cites Heterogeneous Multi-Agent Reinforcement Learning for Zero-Shot Scalable Collaboration.

Why Do Multi-Agent LLM Systems Fail? Heterogeneous Multi-Agent Reinforcement Learning for Zero-Shot Scalable Collaboration

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.724310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:db662469f16c66afca36949f2cbae584f4187767b7c91d478ba5a157b525c380

Observation 08cbb6b2-20f1-4d51-8a9b-6f5bd8d8facb · outbound

This paper cites Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System.

Why Do Multi-Agent LLM Systems Fail? Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.737405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:13ae42458c1a9b4581147d4baad3692b0ce328c2911a5ab05eaf67576a5f8dbc

Observation 6729f59f-1ba2-4517-bf79-f55a94b980eb · outbound

This paper cites Uncertainty, action, and interaction: In pursuit of mixed-initiative computing.

Why Do Multi-Agent LLM Systems Fail? Uncertainty, action, and interaction: In pursuit of mixed-initiative computing

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.961798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:855633e4096ebd913130bd8a2499c20981766772c2630720fb31fed122712ff5

Observation 1ba11e89-ae6a-46d3-bf8b-0bc6279e827c · outbound

This paper cites Servicenow: From startup to world’s most innovative company.IUP Journal of Entrepreneurship Development, 20(1).

Why Do Multi-Agent LLM Systems Fail? Servicenow: From startup to world’s most innovative company.IUP Journal of Entrepreneurship Development, 20(1)

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:58.965040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:4f7c7a702f0b8811fb8b1005e4aa08614a78140379d0b63be90bedb20f01fd5e

Observation ac0af52a-9b08-4286-9cec-ebf9de1c221a · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

Why Do Multi-Agent LLM Systems Fail? GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:58.742230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:8748bceb300d6804279d828a37e4d4479e1816179ac3ea3ece352a2b4d518265

Observation 6720efe1-1d37-4156-bf89-bf0de63a5598 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Why Do Multi-Agent LLM Systems Fail? Training Verifiers to Solve Math Word Problems

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.748112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:139aee7f34f2a03de7017f7e5114a8ffbfd3ab043c1c4a4050367a0bb2bbe57a

Observation df0b3754-20fb-489f-9b5f-3a109aa88515 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Why Do Multi-Agent LLM Systems Fail? Qwen2.5-Coder Technical Report

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.754585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:b82bd4571b29349fc2ea3c5eed6b335f0c40e26035985be44548c9c5979920e3

Observation 38b5d91c-c422-41b9-8e68-de062576f625 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Why Do Multi-Agent LLM Systems Fail? Code Llama: Open Foundation Models for Code

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:42:58.758845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6946e5dedf7b603be10a9c7446419c1981f08de1da4957b5bbb6298e5cb38e2e

Observation 79e2c177-7a48-4813-8e21-9a45623028df · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:58.995357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:0225f6f544cfd54031f6f4b245131947c61b2b4d309128ce3d9eab786d21f89a

Observation 1249f12d-340e-4300-a13f-f1679988d943 · outbound

This paper cites Magentic-One.

Why Do Multi-Agent LLM Systems Fail? Magentic-One

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.000455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:a22fefc404f279be782e190a12af55d9a298bf98de8b34f02674b91be8c29dc5

Observation 5834ff44-db98-4b3f-909c-faed532e7802 · outbound

This paper cites OpenManus.

Why Do Multi-Agent LLM Systems Fail? OpenManus

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.021261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:fb136318f9c5db7e088e11cbd47d2e5718c1f98c717fb03d492146940f458492

Observation bfddfb73-cca6-4fc2-8953-62a9999b0b2f · outbound

This paper cites Communicative Dehallucination.

Why Do Multi-Agent LLM Systems Fail? Communicative Dehallucination

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.033912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:b2dd67b1447968eb48e21522ceb973d05fa81e8be9827809739dd9daceacc58c

Observation 2d3aaa6a-8b23-41b4-8e29-a4a72a94d32d · outbound

This paper cites Interagent Misalignment 3.

Why Do Multi-Agent LLM Systems Fail? Interagent Misalignment 3

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.038595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:6b88273ecc7c830589e4830521f743758d7b3fd7f155cb34c6e6a0e2cd0e0cdf

Observation 843d931c-b7bb-42a8-a625-da2ba3fa2e8f · outbound

This paper cites Write me a two-player chess game playable in the terminal.

Why Do Multi-Agent LLM Systems Fail? Write me a two-player chess game playable in the terminal

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.048118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:b63440cfae16e58ac8eb97a9b977334c74be33cd9b633bf6a12790f8c3deb4d6

Observation 55f9954a-e155-4d10-9a1f-ab758df925a4 · outbound

This paper cites Interagent Misalignment 3.

Why Do Multi-Agent LLM Systems Fail? Interagent Misalignment 3

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.059894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:63af64cb44465c20835d12da66169d8018c316b7c497a2b8ad701da05abf91c6

Observation aeeb615e-31ef-49d1-bf18-d172a61cb599 · outbound

This paper cites Interagent Misalignment 3.

Why Do Multi-Agent LLM Systems Fail? Interagent Misalignment 3

Reference 94

Resolution
malformed identifier
raw_fallback, observed 2026-05-12T05:42:59.079965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:406f373a836d284c463060b4c3a61c7afd1164a7081aab96ba91011da1cc6e71

Observation 7a4deb5a-8ced-4fd1-bfdd-63c1409c9e3c · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:59.094741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:d781264eb1415c83fec1ab07524293762a1851b35d9137ffb2eb088b48cabab4

Observation 9fad2d7a-ac04-4e00-bd21-2cf70284f342 · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:59.103748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:b150a9fc03b3344e646d484c4bc9cbcb049fc28d76e8b8d61e9edcf917915c95

Observation 6d84246a-b26f-480e-9dc8-6a96ec786220 · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:59.110337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:1e4d3922271ec04137eca5a80a18e91e6c6328550634ea53abf17226684bf574

Observation 5fda7fb9-e911-4eb4-bcf7-ba3e0c22616d · outbound

This paper cites If the result is invalid or unexpected, please correct your query or reasoning.

Why Do Multi-Agent LLM Systems Fail? If the result is invalid or unexpected, please correct your query or reasoning

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.126740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:917668c82dc7032563c74604bc380365babfcbca87c9513d4ecd3e7355237790

Observation 0ed24cf8-7680-4fb7-9c24-21090756eadb · outbound

This paper cites Use fractions or radical forms instead of decimal numbers.

Why Do Multi-Agent LLM Systems Fail? Use fractions or radical forms instead of decimal numbers

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T05:42:59.153749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:97f9a326b49dd28c685ecdc72baecaa79fd05f66a0a18b0f314d1a8c716f71ce

Observation c92d19d2-6cc9-4fe6-8493-13a9cd9ccc09 · outbound

This paper cites an unresolved cited work.

Why Do Multi-Agent LLM Systems Fail? Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-12T05:42:59.157736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:cd53d1bd374b514b8ca6e43fe5923c0dd5104d4e6e5df17fa1e41cde3ca46e6c

Pith citing papers

Observation 3fa6d92e-7307-45d5-99eb-f64291649600 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications Why Do Multi-Agent LLM Systems Fail?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:12.357777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:12.357777Z digest=sha256:e9cdff94a4ec5b97b44ca0aa7e22b30f9282306eb03fccbdc6e6b4c19ae06604

Observation 9d643e24-23f0-4a49-a6d2-bd19713a361e · inbound

AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need cites this paper.

AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need Why Do Multi-Agent LLM Systems Fail?

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:46.134886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:46.134886Z digest=sha256:8799db733d80bcbdb0114cdf6cfac45faf6545f5eaace41399d1213f00721229

Observation 948c32ed-b38e-48e7-a8a5-a4f0d22a5f33 · inbound

Initial Investigation of LLM-Assisted Development of Rule-Based Clinical NLP System cites this paper.

Initial Investigation of LLM-Assisted Development of Rule-Based Clinical NLP System Why Do Multi-Agent LLM Systems Fail?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:43:12.742792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:43:12.742792Z digest=sha256:e33f6df7700ff63a6086cacf21e6fa22d81049f9f86be7486d64322fe2b4728c

Observation be3b2584-87c4-4ca6-bc69-a67f78875611 · inbound

Kaleidoscopic Teaming in Multi Agent Simulations cites this paper.

Kaleidoscopic Teaming in Multi Agent Simulations Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:47.565189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:47.565189Z digest=sha256:0035a58f2d64ce360e060bdf9c328c5d1914840031ae8e1826cd35ed57347ea3

Observation 182dd31f-5b6b-4f5d-91ef-a376291f312a · inbound

A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection cites this paper.

A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection Why Do Multi-Agent LLM Systems Fail?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:15.100311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:15.100311Z digest=sha256:5d872cfb1a85c435737a64f7ea6da6b774f0a745766d608443fc4ac3a916bbac

Observation 31c9dfb6-da32-4c76-97b0-377d69240076 · inbound

MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation cites this paper.

MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation Why Do Multi-Agent LLM Systems Fail?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:23.099775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:47:23.099775Z digest=sha256:f2477e16cf3993031d939b08d88aba222556fdcd29a16ad9ba3035ffe38addda

Observation bcc1662b-5e3c-4f36-b36b-652713468c2c · inbound

Aime: Towards Fully-Autonomous Multi-Agent Framework cites this paper.

Aime: Towards Fully-Autonomous Multi-Agent Framework Why Do Multi-Agent LLM Systems Fail?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:59:39.665437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:59:39.665437Z digest=sha256:b713eb442a55ec63c82c05fadbbb2e74c505928d092d7a9c8f2615ee1a9edd50

Observation fa08179f-f065-4f53-8e6c-2b63b4fa8970 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:56.793374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:56.793374Z digest=sha256:bd42fe1d2abb1cbaa7dc05036ca2323dc1cf8b288418af548c7371383ead11a9

Observation e8afbcae-ec67-4b03-a0ef-2631e4b9e256 · inbound

FlowForge: Guiding the Creation of Multi-agent Workflows with Design Space Visualization as a Thinking Scaffold cites this paper.

FlowForge: Guiding the Creation of Multi-agent Workflows with Design Space Visualization as a Thinking Scaffold Why Do Multi-Agent LLM Systems Fail?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:14.226186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:14.226186Z digest=sha256:e259ba57ec8f4bded205f2f5f7ddd6634dfbed104f4fe097415b5695a8d27d9e

Observation 01792ce6-a308-4bdd-8adf-a405881fbde2 · inbound

Interaction as Intelligence: Deep Research With Human-AI Partnership cites this paper.

Interaction as Intelligence: Deep Research With Human-AI Partnership Why Do Multi-Agent LLM Systems Fail?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:37.747896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:30:37.747896Z digest=sha256:ec168339b1b3764685eaee61294f4b5965b815db0f7bac9843141e5abde96add

Observation f6eec7f7-0f70-4592-9f23-170733eb73d1 · inbound

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis cites this paper.

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis Why Do Multi-Agent LLM Systems Fail?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:40:51.106283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T00:37:11.945418Z digest=sha256:b12794a7f0c5327cb73f35cf47bbcec7246cb2a0f2e4491535ef7fb281116e84

Observation ad1845d6-d450-4b90-a854-b5fa0ffb049e · inbound

A Survey on Agent Workflow -- Status and Future cites this paper.

A Survey on Agent Workflow -- Status and Future Why Do Multi-Agent LLM Systems Fail?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T05:52:41.959776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:52:41.959776Z digest=sha256:458d0d09e3a37f9ed3f274bdb9adcd15876dc9376e0837b883e74ef1feba2fbe

Observation 47d17b23-1e8e-4061-be06-f01b691f33c6 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.657410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.657410Z digest=sha256:5063dfd1d40370cee2ef7dda714eb82775cd09af1dc305d87dea787144a7cd99

Observation bcf41024-2b76-4a90-a616-5ee54e0c683d · inbound

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation cites this paper.

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Why Do Multi-Agent LLM Systems Fail?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:26.108940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:32:26.108940Z digest=sha256:35badccb86e69c4ffcf0ed97a7da1475a9bd1534037897da4c1a39beac5e1e8c

Observation a1412a14-fc83-4d69-bc19-623e9459500e · inbound

Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics cites this paper.

Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:02.202587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:02.202587Z digest=sha256:c660a3bbbd447508e836f29cfa924402fe8545e1e8222af9003f038f32511008

Observation 0d0f6eb2-7bc6-4f09-b474-55ac03795e6c · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems Why Do Multi-Agent LLM Systems Fail?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:32.422233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:32.422233Z digest=sha256:be66518b4ce79548c51e0a1cbbf10e38222aae6229f98dff4ad295043a339b8c

Observation 781797b3-0d32-4f5b-bc1a-03f158563cb0 · inbound

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning cites this paper.

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Why Do Multi-Agent LLM Systems Fail?

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:51:08.519681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T08:50:53.486588Z digest=sha256:795d780786a811a0fe3db573ba5dcdabac6148e2f78e07461eb29b0de968997e

Observation f1d07e35-cbd4-426e-aa73-e21ffeee458b · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design Why Do Multi-Agent LLM Systems Fail?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:36.637742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:36.637742Z digest=sha256:d8c4b8147274813fc811b25d5bcf2bd9f49d40fb10b83ea15b84ddff1a9446f6

Observation 282b87a9-54d6-4635-b57d-a74281fbc048 · inbound

Auditing medical multi-agent AI reveals risks of false consensus cites this paper.

Auditing medical multi-agent AI reveals risks of false consensus Why Do Multi-Agent LLM Systems Fail?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:23:41.825975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:23:41.825975Z digest=sha256:78178a2e9adedcbacf195b71765b7fa323885c2c1c686744a9081cd0b1d2ccd4

Observation 01605d73-519a-46dd-84aa-334d78114c40 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Why Do Multi-Agent LLM Systems Fail?

Reference 134

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:42:22.253106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:e77a5417fd973231960d1900ed3ceba4c5cafad1cb7c7399a154c4d7cd4e9baf

Observation 01ea5f60-2550-4346-aff9-92343be391d9 · inbound

Latent Collaboration in Multi-Agent Systems cites this paper.

Latent Collaboration in Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:17:48.232736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:17:48.232736Z digest=sha256:9ccc1f579eb4c0949a9470de221c419dd53f0555f91b1690248ced119d01634b

Observation df95e243-8a70-4d65-8f37-2a92410713ca · inbound

Latent Collaboration in Multi-Agent Systems cites this paper.

Latent Collaboration in Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:52:30.129143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:52:30.129143Z digest=sha256:59e679d52b3c795d9ca26aec74f3d272df10192cdc172fceab2343b54f8e1a17

Observation c4f8126a-25b0-4136-8081-6c1db6ce9c02 · inbound

Process-Centric Analysis of Agentic Software Systems cites this paper.

Process-Centric Analysis of Agentic Software Systems Why Do Multi-Agent LLM Systems Fail?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:18:57.293757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T03:14:15.095330Z digest=sha256:d6330e288a2f95e8b5ddcc3dc294bd8f4a1c718e9a6a7f273fc02ad16679a890

Observation a05e70ab-67b2-46be-a306-ca8bbddd36ea · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Why Do Multi-Agent LLM Systems Fail?

Reference 250

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:31:24.836371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:0925cab020250a57681de64e9c65fda4c6871facd8f879470db4ecdaf96e0874

Observation 429ac015-9bbe-4dea-9f2a-afa89f80cca8 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Why Do Multi-Agent LLM Systems Fail?

Reference 250

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:33.589393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:33.589393Z digest=sha256:ec33cc8c2e8e9a26b065a189bc069c55bc2515c6bb93525193eaaf3dce9a6851

Observation 4fe0bdd6-279e-4cf5-8692-8e824bd271c3 · inbound

SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs cites this paper.

SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:28:16.011981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:28:16.011981Z digest=sha256:002ee57137b43b4f3e539180e5f3bacc641dd674dcdc2eac82ad25fbd1b0d72c

Observation 5c6de89c-c8ab-4207-9a86-6de045c0fc92 · inbound

Specification and Detection of LLM Code Smells cites this paper.

Specification and Detection of LLM Code Smells Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T15:10:13.595342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:10:13.595342Z digest=sha256:a1525836a67a7394087195936286c3fd610659949b2134efe6232e48355db71c

Observation 61dc81f6-4ea7-4dc4-a61c-10f3c4b96aa5 · inbound

Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models cites this paper.

Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models Why Do Multi-Agent LLM Systems Fail?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T11:01:43.609177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:01:43.609177Z digest=sha256:6005b0ada33816652fd94286e20a2282cea90cee483ddeef40ee102d704c3a17

Observation 40624c13-3963-4e1d-a7ee-736647bb0c6e · inbound

When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling cites this paper.

When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:57:50.116180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T11:54:53.273386Z digest=sha256:889c4237dfd7beac7654c30a0a0d78e6e1246e2a533daaf8effdad037797f25f

Observation 8b432749-dbb5-4730-ae95-7d80ae4644ad · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:50:49.392902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:c9f45daa87bbc03f70b51c821ab9c54c3f274ccd0b77250cadeb302be37c6331

Observation 33268900-46fb-47d3-b604-333e62fdf3e1 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:38.444941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:38.444941Z digest=sha256:4461f91a3869f4aa8fd4fcd4c3c1ff32a7285447bf457b0ec683c738ddf55589

Observation d817c6eb-0b10-4fd9-a738-2d332b064753 · inbound

Dual Latent Memory for Visual Multi-agent System cites this paper.

Dual Latent Memory for Visual Multi-agent System Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:04:50.491839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:04:50.491839Z digest=sha256:83cb28b305e5659a4b7c381d9c48c40c39d7c70cddb8764c4e1d70a769f0a604

Observation 0f84887d-09ba-408e-a202-20e5180dcda6 · inbound

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports cites this paper.

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:18:01.150313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:18:01.150313Z digest=sha256:2864a4a8a093551d943ec845fa19c69feed2e080d533ebcb381317f5a577e319

Observation 59e7c1cf-cd03-4b8d-9abf-c854dc97bb1c · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:04.493783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:04.493783Z digest=sha256:f161745247e64cded353f33e3cf42161ab714b0e292fe28f13017cb4f70beb2b

Observation f9ea31c3-60f0-4ae2-b87c-20f534d914c8 · inbound

Formal Policy Enforcement for Real-World Agentic Systems cites this paper.

Formal Policy Enforcement for Real-World Agentic Systems Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:10:18.572990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:08:06.259177Z digest=sha256:70a698d3dd9f6d50a9abe669fb0587d844be88787355e95312337b99e5235ede

Observation 9188d3c5-a3d3-4b2f-829d-c58554793e8a · inbound

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System cites this paper.

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System Why Do Multi-Agent LLM Systems Fail?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T21:59:14.368894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:59:14.368894Z digest=sha256:96267a12d2722a6ed8ec250ef73905af989e5fabe73cf3d24001c75625872f5a

Observation 9f120868-6378-437c-b9b1-e73b5ae20db6 · inbound

DIG to Heal: Scaling General-purpose Agent Collaboration via Explainable Dynamic Decision Paths cites this paper.

DIG to Heal: Scaling General-purpose Agent Collaboration via Explainable Dynamic Decision Paths Why Do Multi-Agent LLM Systems Fail?

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-02T19:59:45.979052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:59:45.979052Z digest=sha256:2c505499aa84c71dbc87dcd85b557df4faa0705e2b2e3760313311dfb4a318e2

Observation 4e2a098d-a09f-41a7-91fa-d72b408e99f3 · inbound

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes cites this paper.

Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:46:08.250157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:42:40.925632Z digest=sha256:f01ea3d6352d4b90d92d87c5396b16b70dd89d1d0c1ca775b561439d685edbf9

Observation 5349091b-0b94-41f2-b1e8-3ba8731965b7 · inbound

Security Considerations for Multi-agent Systems cites this paper.

Security Considerations for Multi-agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 199

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:15:55.849512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:12:14.160789Z digest=sha256:b47c0d6f414fed276c309db7ecc74eba352d4af5f41f83db7b2911e3097ddddf

Observation b688df51-36e3-46cd-b6b7-5aca1c125038 · inbound

Transition from Statistical to Hardware-Limited Scaling in Photonic Quantum State Reconstruction cites this paper.

Transition from Statistical to Hardware-Limited Scaling in Photonic Quantum State Reconstruction Why Do Multi-Agent LLM Systems Fail?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:15.261445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:15.261445Z digest=sha256:0f928c4f1ae54cc844cf8c2c17a70f24dd53fbb5886b6fb9ffc66012885546d3

Observation 222abec9-9f50-4afe-bb6d-3f8fb3256dea · inbound

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams cites this paper.

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:14:08.435357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T11:12:46.626382Z digest=sha256:628db068cc6961e2859ea8f37ab735f82363e5057676f1790b137de775357120

Observation 546e9adc-e5fc-40dc-8d47-b239e757bcb7 · inbound

Effective Strategies for Asynchronous Software Engineering Agents cites this paper.

Effective Strategies for Asynchronous Software Engineering Agents Why Do Multi-Agent LLM Systems Fail?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T20:49:09.477849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:49:09.477849Z digest=sha256:d60e5ebb4badb106a003d453e2dfb5fbae2094f5d21b626fe8bcfc87d59f47ba

Observation baacabd3-e3f1-4a9d-a53e-6593abbfc864 · inbound

Emergent Social Intelligence Risks in Generative Multi-Agent Systems cites this paper.

Emergent Social Intelligence Risks in Generative Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:48:00.820268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T21:45:04.625084Z digest=sha256:ad24f41e9dd23affc1454858771030ddb37893efd7c0a51bb682de78bc6a6aaf

Observation b166fdee-bb8c-4f49-95b5-eb7302c73baf · inbound

Do Agent Societies Develop Intellectual Elites? The Hidden Power Laws of Collective Cognition in LLM Multi-Agent Systems cites this paper.

Do Agent Societies Develop Intellectual Elites? The Hidden Power Laws of Collective Cognition in LLM Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:13:09.700478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:11:52.387623Z digest=sha256:352557fbb33c7d69bb2fd9b3f35b23f3b94bb3b18222eab65c65e8ad897bacd4

Observation 6c24b429-62d4-44f5-96e7-30e0f453318d · inbound

Improving Role Consistency in Multi-Agent Collaboration via Quantitative Role Clarity cites this paper.

Improving Role Consistency in Multi-Agent Collaboration via Quantitative Role Clarity Why Do Multi-Agent LLM Systems Fail?

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:13:12.977004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:12:35.823189Z digest=sha256:4151503e6f48bd7e719a38d2df6031330602db33db786025c18b07e7ef7c63a5

Observation 78fcf2a1-d005-4e23-838b-1426689ea456 · inbound

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation cites this paper.

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Why Do Multi-Agent LLM Systems Fail?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:57:34.386632Z digest=sha256:36dc4bae190056606898fe03d595565a901a78379a51c28287aeb30397041f18

Observation b3f955ad-ecab-4226-bf03-73196752c8a2 · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Why Do Multi-Agent LLM Systems Fail?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a733d7cb69f77609e0697d70a7fb25093ec62688f9ae7872067ea223e5cfc976

Observation 1dd9a4a2-d9e0-4100-9ecb-f4cbf1910467 · inbound

Qualixar OS: A Universal Operating System for AI Agent Orchestration cites this paper.

Qualixar OS: A Universal Operating System for AI Agent Orchestration Why Do Multi-Agent LLM Systems Fail?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:44:25.389675Z digest=sha256:5851db3145314fa8c39103f3e9dbc1b4e5a42aeb3f59c9f255b542a0df02b7d7

Observation 36eb84ed-8978-4293-a709-4879cc315d1e · inbound

Learning to Interrupt in Language-based Multi-agent Communication cites this paper.

Learning to Interrupt in Language-based Multi-agent Communication Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T19:08:45.818851Z digest=sha256:9db1e24d3e5ab5efedec6f48e3328542d1b1a18dc473313740374a3f1f63fa32

Observation 5c1c7a64-cc64-4f1a-b321-01e83e469adf · inbound

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout cites this paper.

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout Why Do Multi-Agent LLM Systems Fail?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:54:41.534985Z digest=sha256:f553d4bf6753feb0c8a54b6cb24d062e92d4deda087c5d346cfd58825bb65217

Observation e5688f27-c580-42c4-832f-712844961397 · inbound

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems cites this paper.

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:30:30.311285Z digest=sha256:c8fc25d14b7938ddfd55b0feb99da67c4c339ba560756b19eed9d8fe2022ce3a

Observation f737e688-7ada-4649-94d2-a5f437eb67c3 · inbound

Agentic Microphysics: A Manifesto for Generative AI Safety cites this paper.

Agentic Microphysics: A Manifesto for Generative AI Safety Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:41:35.898661Z digest=sha256:cc72dca8c12b451ad0d85b48034960d93a8347b6af3a0e790d10d3cf16f61b96

Observation aa3ed4e3-9d83-4fab-acbf-c8319e89d1bc · inbound

Semantic Consensus: Process-Aware Conflict Detection and Resolution for Enterprise Multi-Agent LLM Systems cites this paper.

Semantic Consensus: Process-Aware Conflict Detection and Resolution for Enterprise Multi-Agent LLM Systems Why Do Multi-Agent LLM Systems Fail?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:05:33.379668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T12:04:08.305246Z digest=sha256:579e36661a2f57ddb75041aaca9c2b0b25f79f1366252d74d9e3025902c2ae55

Observation 92aee467-b906-475d-8467-6cf880f3598c · inbound

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks cites this paper.

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:12:31.218055Z digest=sha256:0a0d87b4db613a76dc58c6fa462e822f9bc0681f1a7418017c58d577862422e1

Observation 74ded1c2-3233-481e-9fb0-37249382987b · inbound

Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots cites this paper.

Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:42:21.843629Z digest=sha256:49b6445f259d37ab1d6d4b7e6263ec8b72410317b3c3480994d111c02e5475fe

Observation 7206b364-ce66-4922-8a6f-d168a81d08de · inbound

Mesh Memory Protocol: Semantic Infrastructure for Multi-Agent LLM Systems cites this paper.

Mesh Memory Protocol: Semantic Infrastructure for Multi-Agent LLM Systems Why Do Multi-Agent LLM Systems Fail?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:59:44.763364Z digest=sha256:b0981de70eb4c942d07e4fec943fdc5e7fc6a257e9efb896e5ed4fe9738265a1

Observation 200b1a41-4106-42da-95cc-966381ca9590 · inbound

More Is Different: Toward a Theory of Emergence in AI-Native Software Ecosystems cites this paper.

More Is Different: Toward a Theory of Emergence in AI-Native Software Ecosystems Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:15:59.663339Z digest=sha256:82924c38da760fb64520126d8ddcc3faee3ac2d204d0c61b9057709fee145b8c

Observation 0129f0d0-8b46-4a47-926f-73f1edce8aa0 · inbound

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation cites this paper.

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation Why Do Multi-Agent LLM Systems Fail?

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:24:45.045405Z digest=sha256:c3936c85cdba99bb0bb785b3dffecff0d2bf18eb65be4d59143c420950402117

Observation aae665dc-2a9c-4466-938d-e8cff2af0ff2 · inbound

Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems cites this paper.

Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems Why Do Multi-Agent LLM Systems Fail?

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:50:57.149878Z digest=sha256:888b1ac5d4b4ce380a56b6564ad57e66d9f6f2b0c799370303c82e00fba12c24

Observation 4d8c387a-7794-48a7-a92c-faecf278644a · inbound

AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking cites this paper.

AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking Why Do Multi-Agent LLM Systems Fail?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T05:55:39.802529Z digest=sha256:74343e1f7d7f70277b025a1e08a941ed007e2c3c65034f3f97c643e3634c87cf

Observation a6f9599d-ef04-40a9-a690-31434d7f5960 · inbound

EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce cites this paper.

EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:49:17.063213Z digest=sha256:d94cd25f4241a7cad2e636358a89429b6ac9710044d5372edbfa3704050870f4

Observation da9328ed-8f95-4684-80ad-2c32f6f07462 · inbound

Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents cites this paper.

Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents Why Do Multi-Agent LLM Systems Fail?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:52:07.314962Z digest=sha256:bf62a02cbd67f73331915c98980a4b502d843c572b95543d1036cd557a008ca5

Observation 65850ca9-1be9-45eb-b8e8-830fad864ebb · inbound

TRUST: A Framework for Decentralized AI Service v.0.1 cites this paper.

TRUST: A Framework for Decentralized AI Service v.0.1 Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:01:28.673240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-07T08:22:14.239443Z digest=sha256:a706aed485d7955768eb88a30159979f89ddfa548daf46b0b1d86a60dc102310

Observation a57720f6-7544-4cfa-a3ae-05dd37d8c7d8 · inbound

Trace-Level Analysis of Information Contamination in Multi-Agent Systems cites this paper.

Trace-Level Analysis of Information Contamination in Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T09:41:27.088542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T09:41:15.185578Z digest=sha256:24aedd5ac377b437aa36b8b6a5c4e7c6e31dfd2f998edee3c4395ab5ebc5de2e

Observation 9a97d7ad-1d44-4630-81fd-6c59fc4237c3 · inbound

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems cites this paper.

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:43:19.764483Z digest=sha256:bc3a6ebcc54d691083946c92a94d00cfa7d17b6fd968dfbf861b2d9b1a05360d

Observation e9becf73-6381-4272-9ef1-c8aeb297e7fb · inbound

Inference-Time Budget Control for LLM Search Agents cites this paper.

Inference-Time Budget Control for LLM Search Agents Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:51:02.872129Z digest=sha256:c49d64fffbe7b71f5c4dfa9d49f0b2b28290ec69f529b5278df7a69ea61f2a76

Observation 5c27ca70-bfd7-46e1-86ee-9c53d18c5602 · inbound

Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs cites this paper.

Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:44:07.323690Z digest=sha256:f7e73ff814225804f9386ae36a126b235dd83836ae2dba9dbe08e9571801d387

Observation 760b94f6-a36d-4f36-8edf-27a0f0dbcd09 · inbound

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems cites this paper.

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems Why Do Multi-Agent LLM Systems Fail?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:21:15.692441Z digest=sha256:82031b1b3c4476680707f23fb67a6a97f499031fce921fd7e9e83410e070820e

Observation 11f9478c-4312-41f8-a716-42eb95085d82 · inbound

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems cites this paper.

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems Why Do Multi-Agent LLM Systems Fail?

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.802563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:25:03.492449Z digest=sha256:a45a6ae56c50c1947154f2d4c0475da950a06ad6fc9fe96c406d012450bfa9af

Observation add8ba5b-0276-46cd-8a3b-a97e8d39739f · inbound

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems cites this paper.

Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems Why Do Multi-Agent LLM Systems Fail?

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:02:32.067117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T07:57:44.235469Z digest=sha256:cd32a79c16f6196faa8709088d6bd01dd9ca6d7aeeeb34dfc662f38dd7dc61d0

Observation 5c5c8c89-6407-4dea-9078-f559d175688f · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:1f07ce100b5f231abbd6001a4ecc9ea299b721230813838d6b11cba40aadde43

Observation debfbd79-17c6-498d-b075-0c3ff0b44df4 · inbound

Is a team only as strong as its weakest link? Quantifying the short-board effect with AI Agents cites this paper.

Is a team only as strong as its weakest link? Quantifying the short-board effect with AI Agents Why Do Multi-Agent LLM Systems Fail?

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T02:28:49.522313Z digest=sha256:c8f782eb800f6eb679dd4e3b35ff6772a49793956175263ec4a5965e80bd5f35

Observation cafec856-56d2-4bc8-8b4d-161c9af60d7b · inbound

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples cites this paper.

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T03:27:03.589336Z digest=sha256:dbd0bc427f2493d51b2edcca364859aa89fe21f8b45ff0f3c1c7dfac7d6b45d7

Observation 72e3c4a6-ce69-493f-80bf-93e91d7ed3e3 · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:06:27.347306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:19:49.062330Z digest=sha256:f4c5d7c999fedbfee214d02a1a2774ee194b90f1448e469a339a5d19512b2e20

Observation 59a4ac8f-eb04-4d2e-85e9-5309776934d7 · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:25:03.756672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T05:24:54.265411Z digest=sha256:f9d63860c6415402f4bdc35ad9db010cda17edc1ce6a6efbd439a8e11bcdc274

Observation d0d71ce1-715c-4037-815f-5e25acfa5a8f · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks Why Do Multi-Agent LLM Systems Fail?

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:877c24d45bb01e90c509250acc7126cb2885da9bc6bff82ba29f1ba73b1efea8

Observation 505f7c51-9684-43a2-b42e-254146b51352 · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:16:31.006348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:29:33.497561Z digest=sha256:7e50bd7f5c1e7107bbdc0a4990f5204191002564e232b547af1f257f86509b3c

Observation 7103ba4f-9464-4d37-82f2-a7a1edf0f6bc · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.769897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:24:26.528532Z digest=sha256:84532a16fab14f6caf13d49c7934087dc85291101fce4be7daad56ffc1d484ea

Observation d75e29fe-489f-4a7d-a440-867245f2a15e · inbound

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement cites this paper.

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement Why Do Multi-Agent LLM Systems Fail?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:22:06.185933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:22:04.669760Z digest=sha256:f573dcb6d4366c7651722c116f1b423f453425ab86234bf96ec6d6555cdc0254

Observation eaae0ef6-4618-4ad6-8c0e-9cb930facb22 · inbound

Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance cites this paper.

Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:37:11.647491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:33:19.193036Z digest=sha256:ed78ec610a718e3d81ecffa4860bd76a801513c151fa6308f0ac0815354d2ffe

Observation 7ab4d192-face-462a-a426-aef955467ff8 · inbound

Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance cites this paper.

Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.463049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:07:05.514346Z digest=sha256:ee77f9739d4e95366c399d36a2af44ec9b841914fc018791c23a48e4297393b8

Observation 987a6c14-f684-4906-965b-107bfcbe25e4 · inbound

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation cites this paper.

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:52:35.525061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T18:51:06.379266Z digest=sha256:15673a1adf14909024f8e35d1beeb48ae16dc852a70a311e097bfa8a3c0dff59

Observation cf67d7fb-8240-4b25-b938-4a876ef53df5 · inbound

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy cites this paper.

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy Why Do Multi-Agent LLM Systems Fail?

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:07:53.856748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:04:57.638215Z digest=sha256:fffe80011b465d2e5d288d887f1cfd876a1303432aea1a439bb45354baa528f8

Observation b6131208-ee67-4dc7-abaf-d56f10cef1d5 · inbound

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy cites this paper.

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy Why Do Multi-Agent LLM Systems Fail?

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:33:46.597296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T21:30:30.384184Z digest=sha256:236af7eb5c11edaf89d74ec29387f283486a29a3d39d948dac8b57285c2b7c4c

Observation 0acde9b6-8974-44b9-87c6-d7fdf97d1b18 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle Why Do Multi-Agent LLM Systems Fail?

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:37:35.740135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:18d4453419d15f45389a3be94268e6feff80335fec18d4d885f488d1a2282720

Observation 2639d896-4d4d-43ca-ac22-3981adf1b8c4 · inbound

How to Interpret Agent Behavior cites this paper.

How to Interpret Agent Behavior Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:29:22.800113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:23:25.269217Z digest=sha256:64c7ad2e285d23391c5f94c1c69033d298ea317e8ed42312a8ff8ddc80f60679

Observation 015b8f16-31a5-4565-bee2-1d8bf13ff9b3 · inbound

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration cites this paper.

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration Why Do Multi-Agent LLM Systems Fail?

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:35:55.770773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T14:33:51.477306Z digest=sha256:b9f88a683207267fbd6bde5fb2d60a3d1468d9a5d60a6c7e28835359595f47a7

Observation ada85bc4-5a6c-4c9f-ba31-cf85c101a1dd · inbound

Holistic Evaluation and Failure Diagnosis of AI Agents cites this paper.

Holistic Evaluation and Failure Diagnosis of AI Agents Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T03:19:43.254626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T03:17:28.794622Z digest=sha256:32e2df746b10c5802dbae1640908151a3e013d1b18a6b2cd9d726caa613c5d56

Observation 86734d7a-0863-4ef4-b020-d075004c5ade · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T03:08:57.752402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:6928947a4caf07218246002c805118e1ebbf2c87fe08b178ef7ae3022aea40a2

Observation 54c43af5-ada6-4c12-9d59-7390a2c0b779 · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T16:52:39.937460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:c3f7888cae2a4c72088d8219e4667cd96fcab48d0e0b5c02aef7f3e86bb3fd1b

Observation 21980877-5a72-43af-8b5b-ce8c4d474fb4 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Why Do Multi-Agent LLM Systems Fail?

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:02:42.069950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:be6b28433121efcb290f0b5eabc6dfccb70f767e0c10fbd936a839d7f5d3c155

Observation 64a59955-792c-4894-8561-9c2230e39d4a · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Why Do Multi-Agent LLM Systems Fail?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.821503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:5615cd3b538534fe07f1360e7966d095bb996ae475c117fc413d89ab74f1894f

Observation cd817ec4-43ec-4f51-9782-a02f7c3ea265 · inbound

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems cites this paper.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.699454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:1d2b2e8d295939cd1501eb8b400fa7b8a601f8a2c392df95023dfcff0a4586f4

Observation 226a6da8-c4ed-4eea-991d-49c32abb8a16 · inbound

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment cites this paper.

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:23:12.024102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T10:21:11.725387Z digest=sha256:e97ecdfead0e0857a2ca6dab860b131dd1fd30665384933422a7ca6a9a9ae63c

Observation 2a6b07a9-19fb-42a2-880d-ef3a2ea16da6 · inbound

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On cites this paper.

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:33:12.356069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T10:31:16.368065Z digest=sha256:17e8d47b0f067a3ca09e0d26f98c6eabf404cc0cad89f134d2b53d9aef0b1862

Observation 66159a0c-0bd5-4879-96d0-eeda4a16ce45 · inbound

What Do Evolutionary Coding Agents Evolve? cites this paper.

What Do Evolutionary Coding Agents Evolve? Why Do Multi-Agent LLM Systems Fail?

Reference 72

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T03:48:03.044997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T03:44:18.658541Z digest=sha256:dce7fb9b304d1895d8eddb4802723b441aafa50c352019a37bb2f3021d848abf

Observation 686e88ac-5580-4ccd-b849-a84b66291f0f · inbound

A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents cites this paper.

A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents Why Do Multi-Agent LLM Systems Fail?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:53:04.205456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T04:52:24.576085Z digest=sha256:53b2297c4dbc731f4c7f2cc7d4bd9392279f0e219c44d50ce7164b24cf34f44e

Observation f1d6aca6-7ba8-4c81-9bb6-789c86250402 · inbound

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks cites this paper.

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks Why Do Multi-Agent LLM Systems Fail?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:13:56.440467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T02:10:13.030288Z digest=sha256:9c70b5d2afc027ca127ce72a3c8074ee91c1bec05bcdc5815fc464b117f20db7

Observation d6980f90-b8e6-4341-b0d1-89fd4f074a39 · inbound

Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables cites this paper.

Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T07:14:02.553228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:13:28.742008Z digest=sha256:7efe3460bc28649ff5dcefb5e06461491455560c22c80d7ca38d0625d4f2ccfb

Observation 4c0d6594-74a3-41f7-b1bb-31eefb3a3598 · inbound

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents cites this paper.

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents Why Do Multi-Agent LLM Systems Fail?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:39:44.120121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T06:34:43.686189Z digest=sha256:24cc0d3079c805dc2ea8ec99341782d3b8c5d610f8d4f8da437a23f099d9ce12