Pith. sign in

Paper Citation Record · LEDGER

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2605.08647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08647 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T00:54:36.868103Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T20:14:36.433937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:16:29.356741Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact23
  • verified fuzzy12
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ffc1c3a-fa5d-48f7-8975-f209b19218f4 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.713611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:3728324c72c6c03fc6acb563e20dc8e290db6a5b1ad46697acd6b6e50f3f4844

Observation af66539a-b9cb-4979-84d5-6c8afb2ac9bb · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.735680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:e9fe61dae3d713ca8aa4578358546ba52105ecdfb1305a064110b328d9cf3d3e

Observation 27d9129c-f892-413f-965c-2d97398a31f6 · outbound

This paper cites Chiang, J.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Chiang, J

Reference 4

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.388051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:b6e6c35c7f817e7e813d77fab58351ccf8a54a50a927f6f8acdeb7ac7c68f958

Observation 4ff60c3b-d11a-49a5-88fe-faf44a00eacc · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:5af72e72309dd012418b8319228c0df6ec6408e40d1767cc205245c418c8bfa4

Observation 34ca115f-8022-48cd-abc8-ea9f139a0168 · outbound

This paper cites Comanici, E.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Comanici, E

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.728610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:8436ed00eec7e30df6e8166533264874831c3d9528a8dc7c4fc78e24a2c8e3ae

Observation 60a41dec-7567-4ae5-a015-201d288f8f6d · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.746186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:768a5fe63e94946fdf4f7fb92ffe9ca925a1aef7c6fd7dd4f73418e489df9b06

Observation 39777817-2e85-4213-83d0-d80cc8ed4d43 · outbound

This paper cites Free-MAD: Consensus-Free Multi-Agent Debate.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Free-MAD: Consensus-Free Multi-Agent Debate

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.414011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:b449c7d581423ae1095d6e02e8d21d49f195ac59ee6757b872f597e7c8728081

Observation c376c9e1-7fb0-4a71-93fe-a3936a5efd43 · outbound

This paper cites Peijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li, Xiaowen Chu, and Bo Li.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Peijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li, Xiaowen Chu, and Bo Li

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.405204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:7b5582ea4dc1e331457ee7609d3eac5343653be669fe2749220ce5bc4ff82dc7

Observation b5bf4ff1-2282-4e6a-8a5b-990b8170755a · outbound

This paper cites Dubey, A.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Dubey, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.725012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:e90b47950eb5b2842421cea394e69b789e173fa9c55e209796080db5d2b3d35e

Observation a6a928f9-9981-45b6-80b5-6f22f480dd78 · outbound

This paper cites RAGA s: Automated evaluation of retrieval augmented generation.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators RAGA s: Automated evaluation of retrieval augmented generation

Reference 11

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.396123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:a934e6095ded652e1997a91759622e0b31985990bad6706db8fd4e9d2d14a47a

Observation 8b3ce02e-da6b-455b-9e38-3d390675d94d · outbound

This paper cites Field.Discovering Statistics Using Ibm Spss Statistics.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Field.Discovering Statistics Using Ibm Spss Statistics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.717948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:ba130a12c2c1f6c0248ef655080fb9150f673be953155ebd53dfbd07c87097fd

Observation 4dd1928b-3956-4d69-a0be-816505bb6cdd · outbound

This paper cites Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:49:38.916324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:736620265bdb6b22e6650b466df544a41e7b957e9f50466ff866897f8d4c5e3a

Observation d5a3844d-b417-4a6c-a3e1-818ec18d3911 · outbound

This paper cites AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.421672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:8817f74708cc8182bdb27778a6a19815a2db05e2f37f9c0bedc41c73444ea40e

Observation 2903c163-c8bb-4c9e-b49b-23a40e672c29 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 15

Resolution
malformed identifier
doi_truncated, observed 2026-05-12T00:56:13.380709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:f8273964c25b0519d79dbfbae63c8ad17d39bb4ddec720c0366836da1c1ac15b

Observation ce4aef87-5fd6-4e7b-8a5c-be504414e816 · outbound

This paper cites [Hansson(n.d.)] Sven Ove Hansson.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators [Hansson(n.d.)] Sven Ove Hansson

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.377330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:30c05a94b9e57a1969f405dae2402428d7f38949479ae6ecf43ddedb9cd2c34e

Observation 77cf9315-eb5a-44ea-a422-fb6e39845ca2 · outbound

This paper cites Agentboard: An analytical evaluation board of multi- turn llm agents.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Agentboard: An analytical evaluation board of multi- turn llm agents

Reference 17

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.384355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:0653ea76591a4b92de1dbefcb879d0386b12dc3abab016ef149c2ce8f83424b2

Observation 6b348ced-0525-429f-808e-8a52c8bd4f02 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.721374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:1878dbb2b559a362ec1906037dbd50f98dca5efcb5de7943c7f896a7f82a4f70

Observation b6e89cdf-6be9-4833-81ec-3d8c27767514 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:56:13.373033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:bb5ad85677e400827083a0791a7c9667988c64cd239d3ea192c40ff167de8c87

Observation 199c9dc5-1370-435f-a1b2-72f8960a3c50 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.732006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:55b77a0fc6b0c8fe083d892a9c40f7a9bda38eb889951fa74c1d841b4aae76b0

Observation 7fc50e3b-db4c-47e9-8b3c-01264ca4b72d · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.742766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:7118577ccc7ab13352f3a0db0b42a50c411ebc92c2b97a8cd7fe16d485c03d0e

Observation eb3fe079-c00a-429d-a6c3-67ef0c1cff11 · outbound

This paper cites LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.392973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:d405c8eb0c35d0ee6888dfd50cf8e1e817e8f1b1843443c425e333ec117791ec

Observation 2502ef56-b121-47b2-a70e-400ac9d114d4 · outbound

This paper cites Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:40.321049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:f72abf8acba90a5c4abac0d3d3ff0512afc8358a3aabcfc5943b5f954abfc591

Observation e1df6523-5c35-49b5-bbe4-e928fe24b1ef · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.739154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:c27f5d7ff5dfd37a38a1e9519a544e56eeef6fa05b1a95ac64ada569ceff43ac

Observation e720f72d-b237-4cb7-8033-07014092da4b · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 25

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.417926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:a692f881a17cefb369a17c3aa33bf5d828c1d18573bf5b5d12b2850e523d065a

Observation f5bd8392-6cff-4da2-83d9-8ac6130bbf6d · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators AgentBench: Evaluating LLMs as Agents

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:56:13.343595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:93a19f7f6d42b3579f400b5afcf2114c6dc3b2364a3be078f765fd796cbc3b23

Observation b630e215-0df7-4481-a65f-cf65af438978 · outbound

This paper cites Chatbot arena leaderboard.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Chatbot arena leaderboard

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.710001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:df1ac9894163175fefbac1d63adeaed3b2a1acf56cc218af6bb46a1914db3205

Observation 9cc47c4c-a9e0-4cf6-af2e-609cb5a2aadc · outbound

This paper cites Luo and Y.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Luo and Y

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.692338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:0e202a0cf3e688cb35164a04a3c8880afd6c29b987d3f6bdc33762f62bc996f5

Observation 0dbd1a47-b4a2-4fc1-b1ba-2f3b968377a5 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.696067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:3142d8a784dad7aefbbcd24bd27605941a7b7e862642eabd8e3911441a66a4dd

Observation 20fdc5d4-2e22-435e-abc1-a5fc45d53aab · outbound

This paper cites Mialon, C.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Mialon, C

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.699771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:271b33dca00051314e3a4ff694d088ac45b4b1fe3899ff7c7179f5d1ecce8fb4

Observation 40dd854c-d6a0-46ac-97d0-1cb8447975c3 · outbound

This paper cites Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.324459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:9fbea1c29d2f531facb4a3176374e9c4cc96670f22cde03feedb6c56e510e486

Observation b0f30d37-660c-4980-8c47-39bc16542b44 · outbound

This paper cites Mohammadi, Y.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Mohammadi, Y

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.685513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:918c26dd11e7b84179f4c4713382f9f0e18f0768c0aec39ad7ed79e4b37920fd

Observation 74ddf4d5-6388-4b4b-8389-4db7c816a090 · outbound

This paper cites Introducing gpt-4.1 in the api.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Introducing gpt-4.1 in the api

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.703039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:6e0f758740a768b80c768e99ddaab36c144a17a60fa40252ee394002453919c9

Observation c10910db-f523-4741-897e-cb294f56a830 · outbound

This paper cites O'Brien and Carrie Jun Cai and Meredith Ringel Morris and Percy Liang and Michael S.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators O'Brien and Carrie Jun Cai and Meredith Ringel Morris and Percy Liang and Michael S

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.335853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:9e6604d3514f4ad51521f91d24033cafa8a07efa1835e3a8ec1bdba281adec3a

Observation f2c276a2-cb09-48c6-bf14-16852d9ca888 · outbound

This paper cites HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.320509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:6a5dbd0cd5dd8dd884a4a9050d42e52837aa16f26410b97e8e8217db9a6a7a98

Observation d1ab13f7-36d5-456d-8456-217aa9090688 · outbound

This paper cites Pitre, N.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Pitre, N

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.706612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:f4ddf00e747d789671ad2c3a7a35e4464e6ae3385752421445383b863639e68d

Observation d3abf4b9-6f17-427c-b04a-780c43c540c2 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.673784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:e6a91d8b7d45f762d971e4855fccb9d5df8a9e3574d661598a1f8a4e70e917c8

Observation 3d839c9b-7b50-4e4b-90de-db88b2c2deb1 · outbound

This paper cites Qwen3.5: Towards native multimodal agents.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Qwen3.5: Towards native multimodal agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.678338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:68ee25359df0e4e51043f1de907e621ad29303ceafc7547adbc001d6721dd901

Observation b393f9d3-6e88-4a8e-bb58-0a7292be42af · outbound

This paper cites Sentence- BERT : Sentence Embeddings using S iamese BERT -Networks.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Sentence- BERT : Sentence Embeddings using S iamese BERT -Networks

Reference 39

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.327637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:f1dc16446aba76bca65b1246ee5cdd651b8d78b5353f499d2bae45a50ea6a8fb

Observation c18abb84-4e96-4492-8297-81bf74deb0e6 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 40

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.357676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:36e149a9d393a7b464f9de61999c5a2a8cc84a7387d0c1db9431d2ff2e89290e

Observation a5bc9716-9af0-4132-b2de-8cda982fe006 · outbound

This paper cites Sintes and A.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Sintes and A

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.681965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:11d45d72fd41cd669c32db4a5550e6b26cfb039322d65092f22a916eecff68b4

Observation 8f5e70eb-abc5-46f1-b2ba-cd4327c7a7bc · outbound

This paper cites arXiv:2508.03841 (2025).https://doi.org/10.48550/arXiv.2508.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators arXiv:2508.03841 (2025).https://doi.org/10.48550/arXiv.2508

Reference 42

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.368500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:621f091d13f2612c1c6db65e164d02c631797995a3a8808991dff1d34b0fa4ad

Observation 24e8664e-45e0-441c-ad71-ae14f5a25dc1 · outbound

This paper cites Anselm Strauss and Juliet Corbin.Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Anselm Strauss and Juliet Corbin.Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.339363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:530641826133b8c8dbbf8e358d212fd8fbe6f694ad03c43853400a2cceb2db1e

Observation 56e3e860-fabb-4f51-8a67-7cfa78be8274 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.688774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:a8927c8a8bc8306bc572a64c6095b90de7a23e17d9ec963cb61bc84af31be291

Observation e5acc373-1b07-4f75-9ccc-598d50690a32 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.655197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:22c54529bf4f0e63f666f43437a7af7c022f693bc471021a353d01af5c4fa28b

Observation d3c85c47-dc82-463b-9223-f9be948d6c7c · outbound

This paper cites Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.354565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:e9300494aa1c6674bceba3842b90766ae7428c005c16e084531ad5a5e874eda0

Observation 9c1887ed-f9ba-4c07-a64b-da978f065de4 · outbound

This paper cites Agentleak: A full-stack benchmark for privacy leakage in multi-agent llm systems.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Agentleak: A full-stack benchmark for privacy leakage in multi-agent llm systems

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.347194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:0e6c0f5e3d53b56e9b1dac5602dc0ab8240447b557691eed17c306d403418a61

Observation f5ced301-052b-40bf-bec4-eb2d5567751e · outbound

This paper cites arXiv:2509.23055 [cs].

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators arXiv:2509.23055 [cs]

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.350550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:d0ee86e6a7e25a4767377e63634b6dbb32a7cb5178f7f687708cf2d093aa20ce

Observation 690cd1eb-681f-4fe6-be67-cedf3c5e3c66 · outbound

This paper cites G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.331650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:c7951890beed9846cb5ace0c5ce0efa71e7bf465db0b6891dd8f7ddeaa6c7438

Observation aa6eae63-bc01-4907-9488-3601eaa2b437 · outbound

This paper cites Agentdam: Privacy leakage evaluation for autonomous web agents.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Agentdam: Privacy leakage evaluation for autonomous web agents

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:13.365464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:0b749683d598e67350cd1089c0fb2a40de0b80a680fe41d0aeb0e34b319a0f4b

Observation 198add57-647f-441f-9249-f0b4ad4f5f8c · outbound

This paper cites MultiAgentBench : Evaluating the collaboration and competition of LLM agents.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators MultiAgentBench : Evaluating the collaboration and competition of LLM agents

Reference 51

Resolution
metadata mismatch
doi, observed 2026-05-12T00:56:13.360886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:0143d6b4e77b545f8890ae7ac60ad6fab69963d9095c19a26c1b19980ab881f4

Observation cb15d020-21fe-4ca8-a592-e21f0d082c12 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.666940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:8fc298dfd5d514c9b8734edba63a9cceb532573f3c91ff0ba133644e1521a219

Observation 6e8392b8-c997-48f4-9ef8-45dd270c9863 · outbound

This paper cites an unresolved cited work.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:28:26.670174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:dccfe0494790218e1c6459e09614101fbe2826611bf5e48f712b93253aac6e22

Observation 58d2be13-0a99-4618-ab4b-b63a2fca7b49 · outbound

This paper cites Label NO only if the response remains non-adopting throughout.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Label NO only if the response remains non-adopting throughout

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.659271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:b10d72a27aee9b977bcc89151d638015316cf8b4ab236c5941de96a53a911a0e

Observation b1479334-f6f2-467a-b821-8731475c828a · outbound

This paper cites Critical Security Constraint.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Critical Security Constraint

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-05-14T09:28:26.663172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:d53ca37c494175e68fce9d135bd6352997cd0acc88eb2b70736500d58539238d

Pith citing papers

Observation 6ecf1f90-76c9-41b4-9b61-1be4aa347cb7 · inbound

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety cites this paper.

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:16:29.358011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T20:14:36.433937Z digest=sha256:3973d16775236dd626e6f432e5918f931f533ea00f256ed168e000cb33b3ca00