Pith. sign in

Paper Citation Record · LEDGER

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2605.17467.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17467 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:57:19.545442Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:11:42.030356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T21:11:42.215613Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact18
  • verified fuzzy15
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a66a28f9-8c8f-413a-9f82-1a9be409031e · outbound

This paper cites GPT-4 Technical Report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.714684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:b660a4db76c0da1f25ce2d42a5cb5d2e9468ef9c6db59808d1923fcfba9d1e8e

Observation 46cfc95f-1c38-4efd-89ae-b944e9ec67a5 · outbound

This paper cites Advani, Trajectory guard – a lightweight, sequence-aware model for real-time anomaly detection in agentic AI, arXiv preprint arXiv:2601.00516 (2026).

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Advani, Trajectory guard – a lightweight, sequence-aware model for real-time anomaly detection in agentic AI, arXiv preprint arXiv:2601.00516 (2026)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.691564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:f8e059747550ed1404aafec8fa4314025eba7e9ff3875cd4c269c9c051ee3dc1

Observation d35b6493-0fa6-40d8-bee6-760aaa9e0f68 · outbound

This paper cites Where did it all go wrong? a hierarchical look into multi-agent error attribution.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Where did it all go wrong? a hierarchical look into multi-agent error attribution

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.663409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:4a91abfcf17f0cc99594cf68e854599b6f0dea50d1757363b0ee62d6bd78c073

Observation cd817ec4-43ec-4f51-9782-a02f7c3ea265 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.699454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:bec9571728cd38373b8b0ce8ac2be818b43e5ac58af3316e146099a01baf2bc7

Observation b621e1c0-7fdc-4562-9515-9392607f9ad5 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.727464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:99352842cfa44331ab67bed36a1c8183695a425507452001a3da09e62ce0d212

Observation 3abf56db-dd39-4cff-b607-6c90950c3aa1 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Improv- ing factuality and reasoning in language models through multiagent debate

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.236298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:9b33a0dd2627f5eabbf6c6478ce532d5c7f2dae065a2a33afb1fbd125ece6dbd

Observation 16713f7c-eddf-4415-b861-f3363f5d3c9c · outbound

This paper cites arXiv preprint arXiv:2509.13782 , year=.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems arXiv preprint arXiv:2509.13782 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.723713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:f958776cb279242519f55e9c7885a2df9ea4003d9d99b9bb2b4b8380c9eb8b5f

Observation 6cb89088-f43d-4f8c-9e2d-44a003110ef3 · outbound

This paper cites an unresolved cited work.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-20T12:58:18.264577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:baf2cca40bc5320b59f2ab83d08ac2fc822e3c4676be8475f6b20fb961546ed3

Observation 2ad33ce6-0e1a-4c25-b0de-9770526b5834 · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.234446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:6e1801b833de748de8680e253a83e67676a7d79d4146c88b858ce55fe9e90d7b

Observation 7046a0b8-b99a-48e6-86ab-05a37cd927d1 · outbound

This paper cites GPT-4o System Card.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems GPT-4o System Card

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.640135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:a987fca372f199943faed2e61efb3a26cfe6ee29be80b929cbbcd9feaf8eb09c

Observation 17df799c-4202-45f3-af95-dcdcc32af8f0 · outbound

This paper cites Rethinking failure attribution in multi-agent systems: A multi-perspective benchmark and evaluation.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Rethinking failure attribution in multi-agent systems: A multi-perspective benchmark and evaluation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.686709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:dd635eab02f02af278c0cad229f54cf96349490cad01a3c64dc7c3f9dd985260

Observation 972fa3de-becc-4faf-b95f-87572778f632 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.649679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:7038f9e088395038601101b5a58b45330c27deaf1bfeb7e4f76e5748994c119a

Observation c05273a1-8276-4694-bd52-f464b975b65b · outbound

This paper cites Mas-fire: Fault injection and reliability evaluation for llm-based multi-agent systems.arXiv preprint arXiv:2602.19843.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Mas-fire: Fault injection and reliability evaluation for llm-based multi-agent systems.arXiv preprint arXiv:2602.19843

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.703277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:7b0e32123e0d726c0c4d6cb18753e7dd987c9a466e23a28eedcf7337225f0af6

Observation 88ad6f04-fd13-4933-9b48-2e1259c3ac0f · outbound

This paper cites Aegis: Automated Error Generation and Attribution for Multi-Agent Systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Aegis: Automated Error Generation and Attribution for Multi-Agent Systems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.676496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:7f607dd2d9e749c441bfa3fac361ce66f0e9feaed9e92a34258eab1e5014998e

Observation 4c0bc9fc-b6a9-4314-9eed-3b49a9be4771 · outbound

This paper cites Contractnli: A dataset for document-level natural language inference for contracts.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Contractnli: A dataset for document-level natural language inference for contracts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.260573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:9a1a1f3db6d4fc29232769ae5d579432632da3e0adbf56171b6eac60a2d96887

Observation dff3a778-a9bc-4254-8d30-305a020d06e8 · outbound

This paper cites Slm as guardian: Pioneering ai safety with small language model.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Slm as guardian: Pioneering ai safety with small language model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.262510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:73ff092289382b079712dd969903b0bc87fc3702fab54ef04748a5dcd9d26299

Observation 6636dfbc-2a7d-4d16-88fe-636d930ee506 · outbound

This paper cites AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.695469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:5ba1b3e2d86c82b12d75c124a0fd84a5af41cf9277496f1faeb1d92bcf54df9b

Observation f6f32d16-432b-4279-98f5-d73220035e79 · outbound

This paper cites MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.681267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:8b97fd16cb1848a98f9b5cf80109ef7698434dae7e510b130b0df415ce6bb794

Observation 1ca703d0-fa19-46f3-8ae7-8856b4441cef · outbound

This paper cites Explainable and fine-grained safeguarding of llm multi-agent systems via bi-level graph anomaly detection.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Explainable and fine-grained safeguarding of llm multi-agent systems via bi-level graph anomaly detection

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.731694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:df2130aeb6cb6ed4286a790d037936be156be2ca13d2995ca2a0a5cdec400d04

Observation 6ad536fa-9dd7-4afb-98c0-d439195567d6 · outbound

This paper cites Pathak, H.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Pathak, H

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.659835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:b52d0ebf123a982efbd4abe84a2603227670acc9c9966792e45b438e537fb5bc

Observation d59730c9-091b-46de-887c-ae0187476458 · outbound

This paper cites Deep graph anomaly detection: A survey and new perspectives.IEEE Transactions on Knowledge and Data Engineering.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Deep graph anomaly detection: A survey and new perspectives.IEEE Transactions on Knowledge and Data Engineering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.238172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:63ca6b24c0d57363a73f08d06049526f7cc9d02c79ceb8513971c2dc45ebe358

Observation 6050a654-dc83-409c-b9a8-667bec4e0a64 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Hybridflow: A flexible and efficient rlhf framework

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.258378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:c9c08cc69016b0bac74cd5f1811288a769802d3960ca0cd13eefde34ab1a2699

Observation 49aefb52-94a8-48b8-8350-f65c2acdbb24 · outbound

This paper cites Qwen3 technical report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Qwen3 technical report

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.255883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:59f53ec79bc5d75a525a2fa13742d9047c3905e38bc8a08696c092e710a3fc6a

Observation 82ff1cc6-df17-48a5-9c48-cc013c06305c · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.252091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:a54f294a01d773cb0d8d13f959d3ab2d34f071372cfffb95a17976bfdf7efb60

Observation 7ce76036-62eb-4f41-bc39-6fef4b4f8b19 · outbound

This paper cites Saferdialogues: Taking feedback gracefully after conversational safety failures.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Saferdialogues: Taking feedback gracefully after conversational safety failures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.254017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:cb19593c8c3d5724de6e80fdcb0242f279d3f228bb0c2c81d8adb37c06f91a3a

Observation 9cf5e510-6dd0-4748-9423-6321b4dd6352 · outbound

This paper cites G-safeguard: A topology-guided security lens and treatment on llm- based multi-agent systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems G-safeguard: A topology-guided security lens and treatment on llm- based multi-agent systems

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.248174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:91af3e27406e27e4b8116ee024b58e0bcf7f8a378bd613c28177302ccdd8148b

Observation 418c2c75-1c7e-462a-bf12-016df9170044 · outbound

This paper cites Guardagent: Safeguard llm agents via knowledge-enabled reasoning.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Guardagent: Safeguard llm agents via knowledge-enabled reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.246183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:986c1641c16099c5c0f6dec18b0a8df7da39589f08bbdc89b3a28dfda32f04f3

Observation c542d1c4-3bda-461e-82c0-ae0851ad711b · outbound

This paper cites Qwen2.5 Technical Report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Qwen2.5 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.646746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:61c7c22a2001428525d12f9787e65bb6340b45f001e9edb0cc55af3ae7fc8759

Observation 54c479ef-2d88-4c70-a759-921692385558 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.551920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:870f1a967ad6ddd3ae0a63c507862352c5bd5941afeb57cbc0879d2aba90d1e5

Observation 511f8513-bdee-42ed-9607-b809d9b600ef · outbound

This paper cites AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.707074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:4fe60d46ed74c95e1e345929d347f074e0d634a9fe877245c5d199d52cee648e

Observation f5eb13a8-0416-410c-91bc-0e67f88e1b64 · outbound

This paper cites G-designer: Architecting multi-agent communication 11 topologies via graph neural networks.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems G-designer: Architecting multi-agent communication 11 topologies via graph neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.250019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:ead6206764b0c70709926879fce7735b381f27b7946beb42b415c91eb6e99f01

Observation f7d126be-6a93-470f-9898-673ffb358b33 · outbound

This paper cites Graphtracer: Graph-guided failure tracing in llm agents for robust multi-turn deep search.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Graphtracer: Graph-guided failure tracing in llm agents for robust multi-turn deep search

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:0c72edc7d5b9a4b5e8c51aaf7e47a9160ceb0f212af065636c9f2771c66ad5ae

Observation 49530b9a-f2e0-4b7e-ad1e-4c5ae4685484 · outbound

This paper cites Which agent causes task failures and when? on automated failure attribution of LLM multi-agent systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Which agent causes task failures and when? on automated failure attribution of LLM multi-agent systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.241935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:5e39c079da0836a3523afd63b30568af812804af85491b17fcf1e241939fbc85

Observation 4b65a484-6b4b-4e20-8a1d-e7193e5bafdf · outbound

This paper cites Re- thinking the reliability of multi-agent system: A perspective from byzantine fault tolerance.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Re- thinking the reliability of multi-agent system: A perspective from byzantine fault tolerance

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.243877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:3a36578852741cffaadf9e0efccc70267ccfcda9bc85f62ddac32e5939ea54c8

Observation 69a201ba-5502-411c-a80d-a399982695c4 · outbound

This paper cites Guardian: Safeguarding llm multi-agent collabora- tions with temporal graph modeling.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Guardian: Safeguarding llm multi-agent collabora- tions with temporal graph modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.653002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:53b5d47b8e5863b2f172da029ee6431b0a4b3c113a5c641c8c760ba0345e6209

Observation f29ed659-61b5-4d5c-8007-d344914fa113 · outbound

This paper cites verbose database queries correlate with null results.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems verbose database queries correlate with null results

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.718757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:fd6ea3c24a552737213d5d1fe7992b6dc8cc328d5e30ab401fb34b1bf2261c00

Observation 1a1a34b6-5288-440e-b174-d64c1658b4f9 · outbound

This paper cites Agent-as-a-judge: Evaluate agents with agents.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Agent-as-a-judge: Evaluate agents with agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.239958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:d9189333979ab1084144ef36f585d040e2817a840331f8663669900ec7120fd1

Observation 21c7b0e7-c9c5-42bf-8b4b-c8c0b2f56a66 · outbound

This paper cites Latent Collaboration in Multi-Agent Systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Latent Collaboration in Multi-Agent Systems

Reference 38

Resolution
malformed identifier
arxiv_id, observed 2026-06-02T03:04:02.535394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:fe745100925007fac7b9747dc6fde874e64e93dd30d94d537a9d3c9ee14ce45c

Pith citing papers

Observation ae7fae6c-47c1-418d-9b2e-0f7887b6fedc · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:06.559706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:06.559706Z digest=sha256:b7cfa261120083e3b1b9fd9ea491077753c3bbf9f0c62b9aceb0f7d1ed4c954d

Observation 6d87de6c-8525-41f4-8553-b4df645464c5 · inbound

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence cites this paper.

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:11:42.221403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:11:42.030356Z digest=sha256:b6a4bfb337c8d0f3305e9d228d1e905ca273f2213105bf455c1e7120d1dd0937