Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T08:00:45.789649Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2604.16706.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T08:00:45.789649Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:04:24.398577Z
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b93512f3-f6a7-4429-96fc-8429d87f3c1f · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Krisztian Balog, Donald Metzler, and Zhen Qin
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a4f6cfe1-bb76-4158-a342-97083b4e70f5 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c0443dad-030b-4911-9d9f-52715a0643d6 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 894f4d4a-d92d-4f67-8db3-f8d7c631899c · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1982637d-2fad-4555-9868-189c85bf8fb4 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 347df548-b830-467d-aa4e-2a4d1785a44d · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 722986dc-69df-4417-9c05-8f4d29ab3bb6 · outbound
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e19a740d-89fd-4719-92c4-033f2953b5be · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 96d1e742-97c9-4f62-8017-650134e2715c · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on LLM-as-a-Judge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 03c46076-2444-48de-b579-f21a78039509 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d5830afb-2fe5-4599-9684-bf9b972c93f5 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Richard Landis and Gary G
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4bb8cc40-d144-4315-b4cd-0bd193dcff07 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 747a18a9-7c70-4bdd-8932-3306f4902bb9 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentBench: Evaluating LLMs as Agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e9f13570-68cd-49f0-851e-225019763238 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentHallu: Benchmarking automated hallucination attribution of LLM-based agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ce27df5c-9ea0-4776-841a-729af5d0bf7a · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7c9cb805-072f-4416-b1a3-8c3a330cc5b5 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao, Prathyusha Jwalapuram, Amrita Saha, Doyen Sahoo, Silvio Savarese, Caiming Xiong, and Junnan Li
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 01c3f8a3-5262-4b46-abde-c66aaaf1f6b1 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f92234bf-83be-4c8a-a91d-296d73f7f277 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench In: Yang, G.H., Wang, H., Han, S., Hauff, C., Zuccon, G., Zhang, Y
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 911c8154-7ed0-4cb5-ae30-e7a78005e804 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a506623b-98f4-4159-9a21-1e3775c73ae6 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d962f3e3-e2d6-43fc-bc9d-f23beb0206c7 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Don't Use LLMs to Make Relevance Judgments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2cb13816-cab7-41f7-a857-1fdd14102449 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench ISBN 9798400704314
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7e54429f-3d8f-4a0b-8fe4-de845fd7585f · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation dd33ece9-5654-4f19-b8dd-fa39584858ba · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0692cdc7-69fd-4913-adea-25be4156338a · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e4370c7d-19fe-45f6-bbb4-5d9bb82833a9 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 85ac5172-c564-4a96-a62e-0131fc66c5d6 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0f1a841e-dc10-40d0-b40d-52b4d4e22e49 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Survey on Evaluation of LLM-based Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation dada0be9-080c-4b50-bfc1-8021b455ae3a · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Large language models for information retrieval: A survey
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 24978f4b-1c75-49a6-9540-898137b7599e · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A survey on the memory mechanism of large language model- based agents.ACM Transactions on Information Systems, 43(6):155:1–155:47
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5b395675-7da1-40a3-9490-314c7bb298d0 · outbound
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 33f66bf7-ec2b-46af-b59e-761d2af92727 · inbound
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b504a754-8419-45f3-a8c2-f4915f885fee · inbound
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.