Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.758160Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2608.06663.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.758160Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 142 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4d91c039-ad49-4de0-8523-8e0ec040dd6f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Measuring AI Ability to Complete Long Software Tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b0131f-8e67-4138-8ced-bd0416dd7465 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Cannot Self-Correct Reasoning Yet
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2f02d2-5ee1-4303-83f9-b10ec55959bb · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c637a348-385d-4bdf-9277-48b0c9453c64 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e13dfe-c8ed-4052-a90e-8d56c9c6b9c3 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Understanding the planning of LLM agents: A survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b045725-a423-42aa-9b11-3b0314362d1d · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49569b95-0a7b-4b37-888f-b7de3ab610f8 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on the Memory Mechanism of Large Language Model based Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d276526-6a02-481d-8c49-59bec7991584 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fff376-c46b-40a1-ac08-6742aabd3304 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Model-Brained GUI Agents: A Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c4aaa9-5303-41b1-9df1-04d0f9fb8a74 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa6c2f5-88bd-4b5d-8583-c74358440bb0 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on (M)LLM-Based GUI Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79898d1e-c85f-4e7c-ac4c-01028f7fb264 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416228f4-611e-4317-816d-b6627376714b · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495abdf0-ba25-461b-9bfa-39dccef6eade · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7611fd6d-329b-4548-a39b-882f21ab0d74 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Sutton, Doina Precup, and Satinder Singh
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59fb5cb-4685-4af3-ad59-040a0e204946 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f23b146-909e-479d-ab6b-2f57f65a232f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc5a8c1-5005-47b7-a175-aabf23646121 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 805e7fd4-dd42-47f6-b2ea-372260a6d7cd · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1debe8c0-46e5-41b1-90a9-54954f4e4cb5 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e9ca7e-7693-4537-99a7-81361682208e · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3d99fc-a984-40f2-af62-86f7a7d69fcc · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a59d1c-9c40-4e89-9904-6c726f451504 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Harnesses for Inference-Time Alignment over Execution Trajectories
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b00e30-16ae-4bab-9c74-c69743b834ed · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32cb368-2985-4bdb-ba00-13b110e59b19 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ba2e0e-129a-4325-9afb-b97ab8098bce · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2076d76-067d-49b0-b556-169498d713be · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d9aaf94-2e8c-48c3-9c7b-37f5cb63231d · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1e3e64-304b-4287-bd15-f678d3f508d2 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b223cb23-1271-4f9b-a6a2-139b36034026 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ed3bc5-cb75-4177-b4a5-89dd8e52bbe8 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents W ALL-e: World alignment by rule learning improves world model-based LLM agents,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33743249-9f5b-4ae8-bcc4-d377b3b30ab2 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Laird and Corey Clark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8b4f10-3641-48e2-acb3-f6ea52dcf7db · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MobileDreamer: Generative sketch world model for GUI agent, 2026
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fdff7de-756d-46cb-b1a5-3f3f35638a9b · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304c4095-52cc-4303-8b03-781ba31759b7 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentEvolver: Towards efficient self-evolving agent system, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ae34fd-a085-46a5-9538-a4e50a4f620d · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5c6012-b77a-43f6-98a7-b8f0140b360f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lost in the Middle: How Language Models Use Long Contexts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ccb7fc-bc1a-44fa-9e25-e720b5f6ca86 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028e1dcd-ab43-4f06-921e-e162afc1fea9 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Git Context Controller: Manage the Context of LLM-based Agents like Git
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21703a96-ef58-4aba-924f-ca12aa1fea38 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Scaling long-horizon LLM agent via context-folding, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a8c101-fa00-4eed-94a8-fa0f4fb79d25 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ACON: Optimizing Context Compression for Long-horizon LLM Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c86055-aa32-468c-95fe-eb38d94e9afe · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Diagnosing and Mitigating Context Rot in Long-horizon Search
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc86b62-d07b-43a3-9b2b-15d3a3ad1905 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416faf43-b239-422d-952c-75b229eb1b1b · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9386a00-719c-4362-9d2b-47e1ba67ab20 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents O’Brien, Carrie J
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f1a38c-65e0-4521-ae94-a5d7bf04b98e · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd121b7-034a-4807-bff7-c0f3132f3d96 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Graph-based agent memory: Taxonomy, techniques, and applications
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a691d46d-79fa-4b54-aa69-31ff878f77e2 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad3c7ab-96a1-4af4-b83f-cfbd3fa8e789 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0df324-3392-4164-9585-507b372a872e · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a1bb9a-942a-4db0-8214-6ed790172bda · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents FadeMem: Biologically-inspired forgetting for efficient agent memory
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f72a3db0-9a42-4860-b3e8-e028ae15a2fa · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c123bf1-e8e8-47e0-a896-32a2a2db64e0 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f03b548b-bc5a-4e8c-9968-3d6b5dfc9262 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ed0683-b93e-444e-9d7e-ff74ac0c4d53 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Forensic Trajectory Signatures for Agent Memory Poisoning Detection
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc37d485-bad3-4b9f-be01-ba609ce4a7be · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a45079d-5852-4f74-bab1-0229a88e15c3 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f63ca0e-1696-4977-8787-6c7a4866e897 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3287544e-72b4-4d57-b3ce-b2c028532371 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec066b2d-7c9e-4b88-8e65-06046c159ad4 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45abc842-3186-4af0-bd7e-a95067b0d78f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901dfc73-dfa5-4c2d-843c-eef1836d35ea · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84dea266-5519-4e0f-8279-6bd3287d25fb · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691c522d-49e5-43a6-b8d7-ccb8cc017c5f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd39556-31b6-416e-919d-70fc6550e6e3 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb5527c-b2b1-417e-8002-2ce9847da674 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63528579-e520-4bbe-a426-ccf7c5517246 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents OmegaUse: Building a general-purpose GUI agent for autonomous task execution
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3c0ffee-606a-4584-97cb-e4cd47405017 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42aa1d8f-1201-41e9-9a3e-02e021fd7c3e · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation add4d526-bcf5-4272-b7cf-7238890e9cf5 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Governing AI Agents
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ea7510-3b8c-49a0-852c-a570ba15ded3 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The AI Agent Index
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9985a9-5b84-4abf-861f-88c391b0aa02 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Reflexion: Language agents with verbal reinforcement learning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5d97df-474b-4f8f-b5c5-96f76572501f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5753611a-4298-44c7-bd09-4acea5a5c805 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Can Self-Correct with Key Condition Verification
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307c0cdc-04a0-48be-a362-52d169bf18da · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a7b499-63cf-473e-8a30-198ba7f86666 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8685f8-54ae-4ee3-9784-32b8027f87f0 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cfafb1-596d-403e-bda3-fd2f5f9a17cb · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ee9040-cad6-44c8-914d-f865b460ce2f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0011074b-293d-4309-a6bf-df0031a71baa · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fd0de4-9bc6-48f0-84fa-fad8f6ca66fb · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b467a41-9e0c-4344-bb23-a87414d2db54 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341cc47c-06cd-4f9e-9f2f-7efe851364ed · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6470f50-8445-4a83-8d26-7cb5846d6b56 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18100c34-618d-4f6f-a317-ab6985c78114 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67b78875-b23f-4f73-99ed-7715ecf8436c · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f24c853-84c5-4563-a314-5993871c2b10 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a06d184-2b6d-4231-b3cf-177e3d10968f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Let's Verify Step by Step
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12b5d82-c357-4464-89c1-dcb0d289388e · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Entropy-regularized process reward model, 2024
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48e0268-e156-475a-ab03-ac7aadb5c4f9 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d8f352-658f-4f61-951f-23098439453f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GRPO is Secretly a Process Reward Model
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abcf2ec5-4a78-4c06-90b6-c33f2e7f18b6 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6761413-445d-484c-8e28-5896ccde7a29 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5be3098-e6dc-41b4-a776-5bec79c7e34c · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d856c48-6af7-4ea0-9cb8-12146571136b · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agentic rein- forcement learning for search misaligns instruction-tuning, 2025
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2169a623-151b-4748-b08f-1e94f0cacc5f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Self-evolving LLM agents with in-distribution Optimization
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b0478b4-4a35-464f-a68a-9a73bf9eea4b · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40f7a68-d53b-4b54-92b5-d969088ad129 · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabd52b4-b678-4be2-8d51-d1687327794f · outbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.