Pith. sign in

Paper Citation Record · LEDGER

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

As of 2 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2604.16706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16706 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:00:45.789649Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:04:24.398577Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact26
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b93512f3-f6a7-4429-96fc-8429d87f3c1f · outbound

This paper cites Krisztian Balog, Donald Metzler, and Zhen Qin.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Krisztian Balog, Donald Metzler, and Zhen Qin

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.369234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:e639ff3c7742b92eea24dd583cec7ba33468bc62449e87015dca076cb42e8ac0

Observation a4f6cfe1-bb76-4158-a342-97083b4e70f5 · outbound

This paper cites Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.979582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:2a6b3daf61a5c4eb064f10d70b49cea3d63868cd4c530235f7201edc6d9e6962

Observation c0443dad-030b-4911-9d9f-52715a0643d6 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.902869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:4229849a4153d866ad49f94515e40df935ba208068eba0861ab9e9b476dac495

Observation 894f4d4a-d92d-4f67-8db3-f8d7c631899c · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.964601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:6b31b28dfe103de76c9804311bc409dd6434f4254550dd205e105d02db4697c8

Observation 1982637d-2fad-4555-9868-189c85bf8fb4 · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-10T08:02:24.365360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:6ba04dd126fcc6a045c6a4282a9a8107222776c3d81aa5cac955d4eacacd9adf

Observation 347df548-b830-467d-aa4e-2a4d1785a44d · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:37:40.926471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:de479e9d39784563c598f4c1f3a31a7f48b7a06616a991064d74a82c91955c27

Observation 722986dc-69df-4417-9c05-8f4d29ab3bb6 · outbound

This paper cites Farquhar, J.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Farquhar, J

Reference 7

Resolution
metadata mismatch
doi, observed 2026-05-10T08:02:24.374673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:5fbe5405e3765ea3ca1d6784f572b9742739e2cdd005acf6ee5bbbfa623affa5

Observation e19a740d-89fd-4719-92c4-033f2953b5be · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.969672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:04170484cc447936eb4a8a6e011b7a411214ddb2a0c3b1e8f73966c103399679

Observation 96d1e742-97c9-4f62-8017-650134e2715c · outbound

This paper cites A Survey on LLM-as-a-Judge.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on LLM-as-a-Judge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:02:24.959501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:394eb1ad7a5dbb42a0c156cf8293ef555cd89283ac984eb2ab9c5a7edf8882a3

Observation 03c46076-2444-48de-b579-f21a78039509 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.888032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:059fd1ab161e166d9569623dabf64f7cfd368cdf71a00ce6292ca6629861c147

Observation d5830afb-2fe5-4599-9684-bf9b972c93f5 · outbound

This paper cites Richard Landis and Gary G.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Richard Landis and Gary G

Reference 11

Resolution
verified exact
doi, observed 2026-05-10T08:02:24.371924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:3f3f6cb594e80f06272c8f53397abb8ada99341bd2256c605948482fbf4f7efb

Observation 4bb8cc40-d144-4315-b4cd-0bd193dcff07 · outbound

This paper cites LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.977159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:5de2b55e8688cb9a0ea703d887bf506ad798ce61d234bcaa48ffbf357db53df0

Observation 747a18a9-7c70-4bdd-8932-3306f4902bb9 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentBench: Evaluating LLMs as Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:40:05.312349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:f723b39791fe7e3c8425bc06d22693b55040fcfca8e0aa0e403d373c4811541c

Observation e9f13570-68cd-49f0-851e-225019763238 · outbound

This paper cites AgentHallu: Benchmarking automated hallucination attribution of LLM-based agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentHallu: Benchmarking automated hallucination attribution of LLM-based agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.987100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:266fbb97c52772ce708a2da484487d535386b45f24d1c5bdf4b47d224a8408d0

Observation ce27df5c-9ea0-4776-841a-729af5d0bf7a · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:11:22.649439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:a7a10e8dcc17b316439316d0846428c82aacecf98c9bff437966e2de1edd3a13

Observation 7c9cb805-072f-4416-b1a3-8c3a330cc5b5 · outbound

This paper cites Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao, Prathyusha Jwalapuram, Amrita Saha, Doyen Sahoo, Silvio Savarese, Caiming Xiong, and Junnan Li.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao, Prathyusha Jwalapuram, Amrita Saha, Doyen Sahoo, Silvio Savarese, Caiming Xiong, and Junnan Li

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.972157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:98dd41c330c5da7c41b3af06f8ab6e3bf9e910df17de1bd02f3fa26d360233b2

Observation 01c3f8a3-5262-4b46-abde-c66aaaf1f6b1 · outbound

This paper cites RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.962158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:fa0181bf7f974b2babde0ca773fdcced58339fedb90027402de9fbbb443bb7f1

Observation f92234bf-83be-4c8a-a91d-296d73f7f277 · outbound

This paper cites In: Yang, G.H., Wang, H., Han, S., Hauff, C., Zuccon, G., Zhang, Y.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench In: Yang, G.H., Wang, H., Han, S., Hauff, C., Zuccon, G., Zhang, Y

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.354859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:9e45686c7fe49ae61f81c0f133d69f351446d2ac3ed90a58abbd01b54a3dc433

Observation 911c8154-7ed0-4cb5-ae30-e7a78005e804 · outbound

This paper cites Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.347506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:787ca3cfa38e02078c681a54cda3117e732b2dbe12bc39d3932e2b3dc9fe6857

Observation a506623b-98f4-4159-9a21-1e3775c73ae6 · outbound

This paper cites Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.999729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:e810865f35f9ce1ce3d2e4edd9b86fb80e12d3cb22cd559885601872aa487c40

Observation d962f3e3-e2d6-43fc-bc9d-f23beb0206c7 · outbound

This paper cites Don't Use LLMs to Make Relevance Judgments.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Don't Use LLMs to Make Relevance Judgments

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.994461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:5bd08ce6321f995efa96c6efa3a1b2182f1cc73975cf66a5d6b5208a80a3de43

Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.009796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:a214bced51b314ea3ca515c6516ac520be7ed2331738cb62398cb80b95d95a96

Observation 2cb13816-cab7-41f7-a857-1fdd14102449 · outbound

This paper cites ISBN 9798400704314.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench ISBN 9798400704314

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.358733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:a0b0f5617eb70b15e09b458c7858e56651b392416583319954b48203822731fb

Observation 7e54429f-3d8f-4a0b-8fe4-de845fd7585f · outbound

This paper cites AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:24:32.777586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:df26b83db1671add692d4926a8564809083b609314bf54d73e9f29d589c5a48b

Observation dd33ece9-5654-4f19-b8dd-fa39584858ba · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.004745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:f5e1e0a87aed63d53b798084642f3fd6a270e23d685e9553371ae84d5b8f5e7c

Observation 0692cdc7-69fd-4913-adea-25be4156338a · outbound

This paper cites A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.997039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:7afd51f1639ba093279a2942d3c3f87f49685ae92d32f48109d8164dfec3b3ce

Observation e4370c7d-19fe-45f6-bbb4-5d9bb82833a9 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.351219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:4f77da47480e4ab5454a644873f52f29d9f69cdc1023446d606c3d40b4299915

Observation 85ac5172-c564-4a96-a62e-0131fc66c5d6 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.987380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:b3907390ef550fa679236331098d1a271450e395c3487c9dfd137b813ec29bb2

Observation 0f1a841e-dc10-40d0-b40d-52b4d4e22e49 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Survey on Evaluation of LLM-based Agents

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:02:25.007134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:7e0e2a2c8da6b16cb4b59a8558b6267816257ce0b681368fd42ababbd6d5bca0

Observation dada0be9-080c-4b50-bfc1-8021b455ae3a · outbound

This paper cites Large language models for information retrieval: A survey.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Large language models for information retrieval: A survey

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.989481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:c4d0e96426785f19ed2df08a6a413587c272743d28ba52d1c311d7d7d13e667f

Observation 24978f4b-1c75-49a6-9540-898137b7599e · outbound

This paper cites A survey on the memory mechanism of large language model- based agents.ACM Transactions on Information Systems, 43(6):155:1–155:47.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A survey on the memory mechanism of large language model- based agents.ACM Transactions on Information Systems, 43(6):155:1–155:47

Reference 31

Resolution
verified exact
doi, observed 2026-05-10T08:02:24.361394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:1b82df826cc1c97f93f8a758f915814475677c0867ed4b748f32d23ac8229964

Observation 5b395675-7da1-40a3-9490-314c7bb298d0 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.240297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:793407f16333a541659dc46715ef9e2fa5f4e450074b80001c4a1de865ce3c90

Pith citing papers

Observation 33f66bf7-ec2b-46af-b59e-761d2af92727 · inbound

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP cites this paper.

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T07:06:37.817238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T07:06:37.817238Z digest=sha256:406618a98ed7a6c18c7b4a209d533d40fb4dc7eba1eefb62171ee8a56bbfb415

Observation b504a754-8419-45f3-a8c2-f4915f885fee · inbound

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP cites this paper.

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:04:24.398577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:04:24.398577Z digest=sha256:01c5991d98a1a155b10292739cb6a05de3bed3cc30c56a17e5fe89389bc0f653