Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:01:51.733338Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 3 inbound Pith citation observations for arXiv:2606.09863.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:01:51.733338Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:04:24.401637Z
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b37cb120-f2a6-4463-8201-2a28f3921719 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents 2024 , url=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6d1927-6029-4ac4-b55d-7d4f635c9424 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents 2023 , eprint=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cb9204-69fc-46d7-957c-832fc5f1e5f9 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc417445-fbdc-4db9-a818-19ec4b53ba39 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4dfc4506-c6e8-43a0-b1fa-cc6eef5c5504 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents (2025) SABER: Small actions, big errors– safeguarding mutating steps in LLM agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d113c81c-5328-45f1-add1-82714ce5ac0d · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Appworld: A controllable world of apps and people for benchmarking interactive coding agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ea5bfbe4-e08b-4d3a-91e5-49811061bf0c · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Which Agent Causes Task Failures and When? On Automated Failure Attribution of
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf71e865-1a8d-4fb1-b134-cdd774c863e8 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c513626a-bae3-4f38-a8c0-9c020f48857d · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents verbose database queries correlate with null results
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dd14717c-c19e-41ab-8df4-2a77e01092e5 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c9d3a1-044f-47cc-9d0c-364d98d06f9c · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents 2024 , url=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff36383-f662-47f2-aa04-cc35edd2eca7 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Length-Controlled
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d782e89c-a885-49b1-8d37-50e1f4d55c0a · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Proceedings of the 42nd International Conference on Machine Learning , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d995c0df-3a9c-4478-8689-5ecbd65992c4 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Judging the Judges: Evaluating Alignment and Vulnerabilities in
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a716ce24-77b3-40f9-bebb-cdb82cbe4560 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Beyond task completion: Revealing corrupt success in LLM agents through procedure-aware evaluation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc1082c0-a23f-4d9b-bf4f-f510ec6b7f8f · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c10b3edf-9052-4f47-b7f9-4ca2e7c1544f · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bff2187-3801-4875-a303-7467ef04bc73 · outbound
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents Advances in Neural Information Processing Systems , year=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b1c7bc-af30-4d0e-802b-c9d11dafc8cc · inbound
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f849e8-ce67-4df3-a03b-c6d59193c50f · inbound
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba35ee0c-a4db-435c-9a67-67491996ea4a · inbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.