Pith. sign in

Paper Citation Record · LEDGER

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2605.17467.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17467 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:57:19.545442Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:11:42.030356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T21:11:42.215613Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact18
  • verified fuzzy15
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a66a28f9-8c8f-413a-9f82-1a9be409031e · outbound

This paper cites GPT-4 Technical Report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.714684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:d1ddf5658ae161b3950d364b7125ab04721cb167aa6150db965f531e02ff4e0e

Observation 46cfc95f-1c38-4efd-89ae-b944e9ec67a5 · outbound

This paper cites Advani, Trajectory guard – a lightweight, sequence-aware model for real-time anomaly detection in agentic AI, arXiv preprint arXiv:2601.00516 (2026).

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Advani, Trajectory guard – a lightweight, sequence-aware model for real-time anomaly detection in agentic AI, arXiv preprint arXiv:2601.00516 (2026)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.691564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:31fffd2bc8996889b52ae3a5f681cab8db946ad169b95333c00a60b6f6ea7f5a

Observation d35b6493-0fa6-40d8-bee6-760aaa9e0f68 · outbound

This paper cites Where did it all go wrong? a hierarchical look into multi-agent error attribution.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Where did it all go wrong? a hierarchical look into multi-agent error attribution

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.663409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:dc3709c47812f16c7eefc2e0fe612118c01d6f9559c40a2e6dd413b8c4d8322d

Observation cd817ec4-43ec-4f51-9782-a02f7c3ea265 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.699454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:9adca1ff658b6a9bf882140b23539059b4b34c83eaebb706385a29bc4d62b3f0

Observation b621e1c0-7fdc-4562-9515-9392607f9ad5 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.727464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:00e15bc41c3a74f2386bd594fa8dbd90533f1fc56034f2ba047bd462853445a2

Observation 3abf56db-dd39-4cff-b607-6c90950c3aa1 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Improv- ing factuality and reasoning in language models through multiagent debate

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.236298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:dfce2916657b15a36b3aca44d9c6867a2d3aed5d9df74620808d2a82ab887e3f

Observation 16713f7c-eddf-4415-b861-f3363f5d3c9c · outbound

This paper cites arXiv preprint arXiv:2509.13782 , year=.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems arXiv preprint arXiv:2509.13782 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.723713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:1bf7449b059e1bbfdc1d503d786a30b1a0f9ed438e398f649c4308fa40365d7f

Observation 6cb89088-f43d-4f8c-9e2d-44a003110ef3 · outbound

This paper cites an unresolved cited work.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-20T12:58:18.264577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:67d37e7d7a63ae467c81b9064941165d7bf62a585b344ff6838b1f8e404cd4a7

Observation 2ad33ce6-0e1a-4c25-b0de-9770526b5834 · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.234446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:6bd8640f81f96866c32eac9d9e75dce5b580fcf6c7b8aca6faf531a1da3edb6c

Observation 7046a0b8-b99a-48e6-86ab-05a37cd927d1 · outbound

This paper cites GPT-4o System Card.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems GPT-4o System Card

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.640135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:e593632be530e5bd52667be3192f9f4767f05544411e97caed476326cd7d01fd

Observation 17df799c-4202-45f3-af95-dcdcc32af8f0 · outbound

This paper cites Rethinking failure attribution in multi-agent systems: A multi-perspective benchmark and evaluation.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Rethinking failure attribution in multi-agent systems: A multi-perspective benchmark and evaluation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.686709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:2ef5209671fa30c646fe455bc8bb71e9d8774663b31c732866772c1dae2b9206

Observation 972fa3de-becc-4faf-b95f-87572778f632 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.649679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:10f80061c16ccfdebf1cc65213da244a7a0930023d741ff416371b32b7d6300a

Observation c05273a1-8276-4694-bd52-f464b975b65b · outbound

This paper cites Mas-fire: Fault injection and reliability evaluation for llm-based multi-agent systems.arXiv preprint arXiv:2602.19843.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Mas-fire: Fault injection and reliability evaluation for llm-based multi-agent systems.arXiv preprint arXiv:2602.19843

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.703277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:a27e5ae10f4a6e1d9f4627c9fbaaf6f09af5f0647c259ea27dcb584c836ab702

Observation 88ad6f04-fd13-4933-9b48-2e1259c3ac0f · outbound

This paper cites Aegis: Automated Error Generation and Attribution for Multi-Agent Systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Aegis: Automated Error Generation and Attribution for Multi-Agent Systems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.676496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:7b91e365153496f173fa3398575845cca0bdd07beb5804ab8ae910fc25769801

Observation 4c0bc9fc-b6a9-4314-9eed-3b49a9be4771 · outbound

This paper cites Contractnli: A dataset for document-level natural language inference for contracts.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Contractnli: A dataset for document-level natural language inference for contracts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.260573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:401e4493901230a53a9b4a3a5eb12f1a1805fa8a047f1af1cfc5f746d80ddb5f

Observation dff3a778-a9bc-4254-8d30-305a020d06e8 · outbound

This paper cites Slm as guardian: Pioneering ai safety with small language model.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Slm as guardian: Pioneering ai safety with small language model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.262510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:b67ba3bff021c9054229c9b5d2c8b5c6c77be4a16abad572a50f5b4dcb2698aa

Observation 6636dfbc-2a7d-4d16-88fe-636d930ee506 · outbound

This paper cites AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.695469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:48714094b0c17af36c5391642c978f3420daf0357088169d1f3c843716be76d1

Observation f6f32d16-432b-4279-98f5-d73220035e79 · outbound

This paper cites MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.681267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:fbc81e6329f07fcdf3ae6aec759f163465e29939e04259371d74511fe3bcd4de

Observation 1ca703d0-fa19-46f3-8ae7-8856b4441cef · outbound

This paper cites Explainable and fine-grained safeguarding of llm multi-agent systems via bi-level graph anomaly detection.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Explainable and fine-grained safeguarding of llm multi-agent systems via bi-level graph anomaly detection

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.731694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:ef35162037e40f049f96ca996f11655c2e901583e0981574d4c5a81e24f573fb

Observation 6ad536fa-9dd7-4afb-98c0-d439195567d6 · outbound

This paper cites Pathak, H.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Pathak, H

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.659835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:db510888fcc70ff525c2907168487749667471c92614e9fe407d581feb834dc8

Observation d59730c9-091b-46de-887c-ae0187476458 · outbound

This paper cites Deep graph anomaly detection: A survey and new perspectives.IEEE Transactions on Knowledge and Data Engineering.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Deep graph anomaly detection: A survey and new perspectives.IEEE Transactions on Knowledge and Data Engineering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.238172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:629e2bbf75b11ac2ac73c718cbb1034585b9323d50a836fdaaa735ada70aa8da

Observation 6050a654-dc83-409c-b9a8-667bec4e0a64 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Hybridflow: A flexible and efficient rlhf framework

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.258378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:4c44b90fe8bc9bac48a462e57c139a06e6d0ddff0d208aa22cd4c3398b804347

Observation 49aefb52-94a8-48b8-8350-f65c2acdbb24 · outbound

This paper cites Qwen3 technical report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Qwen3 technical report

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.255883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:955e0b893d6d1e2a7df8938370d82a31d64b66e2709bbfab602ea1ccedfc4523

Observation 82ff1cc6-df17-48a5-9c48-cc013c06305c · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.252091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:a4a6b2336854ef9a6e64e3f89d23e41d92ea48046981a063ae6b7b9b77d9ccee

Observation 7ce76036-62eb-4f41-bc39-6fef4b4f8b19 · outbound

This paper cites Saferdialogues: Taking feedback gracefully after conversational safety failures.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Saferdialogues: Taking feedback gracefully after conversational safety failures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.254017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:73c790746ade924c0f005f0977b9bee9d7f24275989df13bbde14ac8531e2e55

Observation 9cf5e510-6dd0-4748-9423-6321b4dd6352 · outbound

This paper cites G-safeguard: A topology-guided security lens and treatment on llm- based multi-agent systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems G-safeguard: A topology-guided security lens and treatment on llm- based multi-agent systems

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.248174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:70e96be2f6762b7e9208199813c061cea07fca468204be33906074ea55c173e2

Observation 418c2c75-1c7e-462a-bf12-016df9170044 · outbound

This paper cites Guardagent: Safeguard llm agents via knowledge-enabled reasoning.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Guardagent: Safeguard llm agents via knowledge-enabled reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.246183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:bc06f691d739104eedce85807f6b81a04d17131521d57c37aa02c4221206f167

Observation c542d1c4-3bda-461e-82c0-ae0851ad711b · outbound

This paper cites Qwen2.5 Technical Report.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Qwen2.5 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:58:17.646746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:d0aec019961f0cda5427db437340f37d6119e38d2bcfae2274537d13c78b9d3e

Observation 54c479ef-2d88-4c70-a759-921692385558 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.551920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:e22f696b36779435681f23a31e217b6d9669a08336059978bd2e279413c7bbd4

Observation 511f8513-bdee-42ed-9607-b809d9b600ef · outbound

This paper cites AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.707074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:2bff7de989ca9e263a43035815227c4cb179796cd1d4cfbc2c654bac684208b7

Observation f5eb13a8-0416-410c-91bc-0e67f88e1b64 · outbound

This paper cites G-designer: Architecting multi-agent communication 11 topologies via graph neural networks.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems G-designer: Architecting multi-agent communication 11 topologies via graph neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.250019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:4029e8b92ecf616033ae0e37928bedfe93c77c5527abdd478d922450fe597ad5

Observation f7d126be-6a93-470f-9898-673ffb358b33 · outbound

This paper cites Graphtracer: Graph-guided failure tracing in llm agents for robust multi-turn deep search.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Graphtracer: Graph-guided failure tracing in llm agents for robust multi-turn deep search

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:61b983f3261d6da3d44e1fcf2c85d40e521c71824c44aedc45c400a531c97203

Observation 49530b9a-f2e0-4b7e-ad1e-4c5ae4685484 · outbound

This paper cites Which agent causes task failures and when? on automated failure attribution of LLM multi-agent systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Which agent causes task failures and when? on automated failure attribution of LLM multi-agent systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.241935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:e77e9be9e168bd66029988968ecae47596ab2f2812a34ed79d02ecfbadb78976

Observation 4b65a484-6b4b-4e20-8a1d-e7193e5bafdf · outbound

This paper cites Re- thinking the reliability of multi-agent system: A perspective from byzantine fault tolerance.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Re- thinking the reliability of multi-agent system: A perspective from byzantine fault tolerance

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.243877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:8bdc2ebdfeb91892ecab8e14e15d96ddd222a00ed41d55c34f674cb625940b3b

Observation 69a201ba-5502-411c-a80d-a399982695c4 · outbound

This paper cites Guardian: Safeguarding llm multi-agent collabora- tions with temporal graph modeling.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Guardian: Safeguarding llm multi-agent collabora- tions with temporal graph modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:58:17.653002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:21f748c4f98b420d26d794f69cefad406e9d2b7ff16d88aef7bad9b3ba3ed354

Observation f29ed659-61b5-4d5c-8007-d344914fa113 · outbound

This paper cites verbose database queries correlate with null results.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems verbose database queries correlate with null results

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:58:17.718757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:36239d2b5ee4fed241b54343ae7510a704b994af6557f94210f41120816a6fc1

Observation 1a1a34b6-5288-440e-b174-d64c1658b4f9 · outbound

This paper cites Agent-as-a-judge: Evaluate agents with agents.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Agent-as-a-judge: Evaluate agents with agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:58:18.239958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:e041c57cf691063d288014f23ab2f74e1c49df63bd1acd3d01d2a15d6d4d8fa5

Observation 21c7b0e7-c9c5-42bf-8b4b-c8c0b2f56a66 · outbound

This paper cites Latent Collaboration in Multi-Agent Systems.

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems Latent Collaboration in Multi-Agent Systems

Reference 38

Resolution
malformed identifier
arxiv_id, observed 2026-06-02T03:04:02.535394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:57:19.545442Z digest=sha256:5d18dddebedc5f606ab60533f9a8f56aed7c3fd5b53d85284e10b7512df778e7

Pith citing papers

Observation ae7fae6c-47c1-418d-9b2e-0f7887b6fedc · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:06.559706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:06.559706Z digest=sha256:b7cfa261120083e3b1b9fd9ea491077753c3bbf9f0c62b9aceb0f7d1ed4c954d

Observation 6d87de6c-8525-41f4-8553-b4df645464c5 · inbound

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence cites this paper.

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:11:42.221403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T21:11:42.030356Z digest=sha256:ebf6ae8602d23e8aacd0fdcec63c6fd7cb45ce41f3f3fd7dda435a281672e316