Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:04.819123Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2508.11027.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:04.819123Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:05:47.294068Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T09:55:40.559308Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e7be9ab1-4521-419e-877a-27faf18df70e · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b255ff-cd94-48b1-904a-8db69fe30c74 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Teaching Large Language Models to Self-Debug
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86078d18-7922-4534-a362-c275eb6ccb9e · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5c3286-c8f7-41fb-a2a6-e062187493a2 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7164496-35eb-4b23-bfad-11bc1ce8a902 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d9cb98-d90e-4767-bd61-7325d99c7da5 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeacd5d1-46f9-4b70-ada7-6bb6f5f6d577 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee430b8-f4d6-4960-bcf1-b0529eafa17d · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Measuring Mathematical Problem Solving With the MATH Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9158879b-5bb3-4b4e-b9e0-dbe9c8be432d · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Not All LLM Reasoners Are Created Equal
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f5f28c-1edc-420c-a0fd-2972f5e3973a · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Creativity in AI: Progresses and Challenges
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e97a6f0-fb64-4a58-abf3-4ed741cb1c8c · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Feedback friction: Llms struggle to fully incorporate external feedback, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741ba194-7da5-48fd-be8f-1366c203fa9b · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814c3afb-4b6c-4c02-934a-1bed1aaa7ccd · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Api-bank: A benchmark for tool-augmented llms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8254a3f-bd8b-4506-bc88-76b597e87b29 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bcfe983-083a-40b5-b0cf-ef6fc35cdb95 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250807b2-2fa2-4e0e-8ca5-cf06a0b6d7d3 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Benchmarking Language Model Creativity: A Case Study on Code Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 085a074b-e743-4ed1-ad98-aec6a4b0dba5 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Self-refine: Iterative refinement with self-feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 103481f0-5e4e-4254-ab21-8a132555ba2b · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures TALM: Tool Augmented Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7ba190-35e7-452c-8664-caede273682a · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Gorilla: Large Language Model Connected with Massive APIs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2708f164-0128-422e-971f-a02351dac6cf · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7455207-d292-4bce-a197-55269bc3314d · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Making Language Models Better Tool Learners with Execution Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5cb56fe-b445-4770-b59f-314d345d9320 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Tool Learning with Foundation Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18fdcfc-245c-47e3-a709-6c0ad197b9b4 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c7743f-3199-4ff2-a7eb-fafdf924c562 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Qwen2.5 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022af19a-0b06-47bb-8342-8da9b9302231 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Self-critiquing models for assisting human evaluators
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc8ebe8-cb12-4d80-8080-dd0908b53d68 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84308df0-181e-4cb6-80f7-551a42aec180 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f528b36d-0afc-4384-b1c0-811c75bf4ceb · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Tools fail: Detecting silent errors in faulty tools
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2c27f36c-6b43-464e-93d1-8c4bf7c1001d · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77161aab-ed0a-49f5-b8e1-51fb9c144aba · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures MacGyver: Are Large Language Models Creative Problem Solvers?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a7301e-ab49-4020-b33b-caf333fbfc26 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2c4c0f-7dbe-4066-a0c0-40a413774888 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af94b5a2-a022-406d-ab01-2f71680bf1b8 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Executable Code Actions Elicit Better LLM Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc266d72-cba0-415a-8b7c-d31c58da90c9 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cd00d879-5d49-41dd-a89f-20c9d2c3469d · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures TravelPlanner: A Benchmark for Real-World Planning with Language Agents
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02a91bc-171e-46bf-b8d6-73b96e307186 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 325dcc46-6bb2-40a4-9651-aa78fd779d91 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Narasimhan, and Yuan Cao
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 293a55cf-fdbe-48fc-b32b-0eaef57149e1 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10fb510f-3660-4539-9db2-1b87d62a0fa0 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures S pider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to- SQL task
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80662e47-dfbc-4851-ab83-11f7d49bc0e2 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Instruction-Following Evaluation for Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ac5ae7-4539-48fc-b1e4-a38fd99ade02 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Toolqa: A dataset for llm question answering with external tools
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 37e372b8-237a-4c2d-917b-61c9b90ac51a · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures @esa (Ref
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 697dbf48-d41e-4789-a822-439c811d3938 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6373f42-2c47-40e5-88fd-03d49e3b9023 · outbound
Hell or High Water: Evaluating Agentic Recovery from External Failures Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3c0eba-9da5-4d37-9924-5c6c4ef896d5 · inbound
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0adb1785-cff9-4e83-beb7-24b80ffce843 · inbound
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Hell or High Water: Evaluating Agentic Recovery from External Failures
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d82411c5-65f6-4a75-be0b-952622749dbd · inbound
When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue Hell or High Water: Evaluating Agentic Recovery from External Failures
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.