Pith. sign in

Paper Citation Record · LEDGER

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

As of 11 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2608.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06663 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.758160Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d91c039-ad49-4de0-8523-8e0ec040dd6f · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Measuring AI Ability to Complete Long Software Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.467237Z digest=sha256:b7ea6cddc09d058c807affc4175094b087c6c7637c98e171e65577b31b76366d

Observation 09b0131f-8e67-4138-8ced-bd0416dd7465 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Cannot Self-Correct Reasoning Yet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.471245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.471245Z digest=sha256:589feb18b2f75b83f7867b1e188d5220caf3cdcd9431f5e0fb71a6276f7f6f08

Observation 9a2f02d2-5ee1-4303-83f9-b10ec55959bb · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.474552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.474552Z digest=sha256:04880687ff5aa1e536571025fb8f48bc1f372999569361e45ad95a798c72ff36

Observation c637a348-385d-4bdf-9277-48b0c9453c64 · outbound

This paper cites The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.477854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.477854Z digest=sha256:79f8217788cca26a76e79783ed8df83268869d49d8bc3dccd76bb808f2b8a440

Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.480747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.480747Z digest=sha256:b441873ed5e77c3336748946c000053e4734c5c42fe354bf940e09ce3503db8c

Observation 79e13dfe-c8ed-4052-a90e-8d56c9c6b9c3 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Understanding the planning of LLM agents: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.483951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.483951Z digest=sha256:c1d46b1c7134eb5dc65ad7a841e50217a01f36189698361ce6d2070e81cfee86

Observation 9b045725-a423-42aa-9b11-3b0314362d1d · outbound

This paper cites LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.487410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.487410Z digest=sha256:5af5a0df1f1128291a9c890b5790218cc66e1a3f33af3609e6a19d30b54ec3ab

Observation 49569b95-0a7b-4b37-888f-b7de3ab610f8 · outbound

This paper cites A Survey on the Memory Mechanism of Large Language Model based Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on the Memory Mechanism of Large Language Model based Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.490120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.490120Z digest=sha256:b24f916a7df90d35cffe6da45556ba6b300c580613c2d76da6f04141ccc9a0bd

Observation 5d276526-6a02-481d-8c49-59bec7991584 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.493279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.493279Z digest=sha256:01800111c39107a7d2fb40bec0cbb91856dd9a43ca5ba6b515d62106cd4a59f0

Observation 05fff376-c46b-40a1-ac08-6742aabd3304 · outbound

This paper cites Large Language Model-Brained GUI Agents: A Survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Model-Brained GUI Agents: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.495876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.495876Z digest=sha256:f6164e6a9262e502604ced515637af5191d2aeccd1e5cac4c671dfdcaf9a5249

Observation e3c4aaa9-5303-41b1-9df1-04d0f9fb8a74 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.498394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.498394Z digest=sha256:42a1a954e3b96175f52703cf225bc62c210308c7c9baa9f8dbf11a3ffcd62066

Observation 5aa6c2f5-88bd-4b5d-8583-c74358440bb0 · outbound

This paper cites A Survey on (M)LLM-Based GUI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on (M)LLM-Based GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.501038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.501038Z digest=sha256:658c25067d7260ced6bb7978dd01fd3dc31a13858145d77a914f8625b6f1a324

Observation 79898d1e-c85f-4e7c-ac4c-01028f7fb264 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.503946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.503946Z digest=sha256:7e04f2e11f57987592d7bf400afbab3e71cb83d061f3b6462b776fe1e4e900a7

Observation 416228f4-611e-4317-816d-b6627376714b · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.506528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.506528Z digest=sha256:fe27a12973fb38161cd66979bcedc45bac098f34db6bcbf32eeefcd023d30702

Observation 495abdf0-ba25-461b-9bfa-39dccef6eade · outbound

This paper cites Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.509270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.509270Z digest=sha256:2a05569e83e4f3aa41e0983a68bb6ccc3e7d5c45c0f2b0854134c94ee27bb9fa

Observation 7611fd6d-329b-4548-a39b-882f21ab0d74 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Sutton, Doina Precup, and Satinder Singh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.512030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.512030Z digest=sha256:5eecd83a184ee723072c82ec37c3cd426ead73db0d1240396d69f2540088c975

Observation c59fb5cb-4685-4af3-ad59-040a0e204946 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.514691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.514691Z digest=sha256:49f6d78854fa7287459b4dc2807d9269ead9c616e2f5a74d400e3b88efb1cd85

Observation 8f23b146-909e-479d-ab6b-2f57f65a232f · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.517185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.517185Z digest=sha256:df5369282ecc6854e94a27264197084b3d93423c5840c8c9ae7023a4259d5bde

Observation 0bc5a8c1-5005-47b7-a175-aabf23646121 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.520046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.520046Z digest=sha256:bf6c54190c15c36556f485b982399488268fbc4a47ea2809d3128a172feddf52

Observation 805e7fd4-dd42-47f6-b2ea-372260a6d7cd · outbound

This paper cites PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.522839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.522839Z digest=sha256:44634424d7f970cc818cd0d2ed52f6d961905be901cd57efb51ac68fca491259

Observation 1debe8c0-46e5-41b1-90a9-54954f4e4cb5 · outbound

This paper cites Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.525639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.525639Z digest=sha256:b1e3df51cab648eb3c67b77ae46b93b48c6f58036a07aef2ec4282ec4b912e0e

Observation 09e9ca7e-7693-4537-99a7-81361682208e · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.528245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.528245Z digest=sha256:20db5de96ea2265dcc6e5524e6aaecc938585db5e88cabd52dc93b7f40b83545

Observation 0e3d99fc-a984-40f2-af62-86f7a7d69fcc · outbound

This paper cites PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.531161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.531161Z digest=sha256:582a8b08027985cf48446527ae1dba529d09fb99eb87251414830541dcb562ab

Observation 80a59d1c-9c40-4e89-9904-6c726f451504 · outbound

This paper cites Harnesses for Inference-Time Alignment over Execution Trajectories.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Harnesses for Inference-Time Alignment over Execution Trajectories

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.533892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.533892Z digest=sha256:10401da3397476fb448a9d2ef3e72c44f5b4f8b56c6081df97a072ceaa21de45

Observation e5b00e30-16ae-4bab-9c74-c69743b834ed · outbound

This paper cites PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.536440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.536440Z digest=sha256:9c8704cc97ea21751789399836d79d3f7af69428be3c304b31ae7eb2a0fc61ea

Observation d32cb368-2985-4bdb-ba00-13b110e59b19 · outbound

This paper cites ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.539470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.539470Z digest=sha256:04e153890a05c592c252080df167f37ea63660b026679a906d89c89d3409457e

Observation 58ba2e0e-129a-4325-9afb-b97ab8098bce · outbound

This paper cites Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.542342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.542342Z digest=sha256:5bd2493c8d44e819ac2c6d54af754e250dca38b6d1ad6c0846e8159e3f4350d5

Observation e2076d76-067d-49b0-b556-169498d713be · outbound

This paper cites The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.545080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.545080Z digest=sha256:5c4249fbec1fcea169cdc9f048c39e80d016ce58b62eea7859ff5814d0675761

Observation 4d9aaf94-2e8c-48c3-9c7b-37f5cb63231d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.547554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.547554Z digest=sha256:98754752043e5a50e821681982d0d3cd480426b65e335ce09137f2f76ae9d042

Observation 9b1e3e64-304b-4287-bd15-f678d3f508d2 · outbound

This paper cites CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.550199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.550199Z digest=sha256:2a5c9745a4082f3068f135973f6d8c9ba198739f56b2241c674c4e9c3e8b263d

Observation b223cb23-1271-4f9b-a6a2-139b36034026 · outbound

This paper cites SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.553083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.553083Z digest=sha256:98f292eda7dd720e71dbaeb95a4f02341d83de728719ab0e90d97ff1adc605cf

Observation f3ed3bc5-cb75-4177-b4a5-89dd8e52bbe8 · outbound

This paper cites W ALL-e: World alignment by rule learning improves world model-based LLM agents,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents W ALL-e: World alignment by rule learning improves world model-based LLM agents,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.555658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.555658Z digest=sha256:9bcf006b5c78dd61bba5eaa098a89ebeff33d9d46c5238f97d5da1f86a24d266

Observation 33743249-9f5b-4ae8-bcc4-d377b3b30ab2 · outbound

This paper cites Laird and Corey Clark.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Laird and Corey Clark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.561984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.561984Z digest=sha256:1c59ef8f438d21348d5b7f6f40a27b0591956bb2f7942e6b6b8f46b59d1d979b

Observation dd8b4f10-3641-48e2-acb3-f6ea52dcf7db · outbound

This paper cites MobileDreamer: Generative sketch world model for GUI agent, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MobileDreamer: Generative sketch world model for GUI agent, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.564416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.564416Z digest=sha256:641a76066ffb5b0efa81ddd3663c351dc7ba49de4ecbd1ee7ca36bf4592489b7

Observation 0fdff7de-756d-46cb-b1a5-3f3f35638a9b · outbound

This paper cites ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.567093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.567093Z digest=sha256:e13843c01b234439ca1db890ac15e71a2bb54721db97b76fb95e29f549228a12

Observation 304c4095-52cc-4303-8b03-781ba31759b7 · outbound

This paper cites AgentEvolver: Towards efficient self-evolving agent system, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentEvolver: Towards efficient self-evolving agent system, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.570037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.570037Z digest=sha256:08fdd0b7311433f1b514c64baf0d965529f2019fe00b29352f475cb5e38fdfa5

Observation a2ae34fd-a085-46a5-9538-a4e50a4f620d · outbound

This paper cites Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.572747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.572747Z digest=sha256:6afd79b9fe3f2d8cfe9047fb141a1f956803fd3b7c26dd4f32666b80aefde562

Observation af5c6012-b77a-43f6-98a7-b8f0140b360f · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lost in the Middle: How Language Models Use Long Contexts

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-10T23:01:35.575978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.575978Z digest=sha256:894416f7dfc78ccdafcddc756a274e8060b59275c7f4918a1f2a1e4993ff2a66

Observation 42ccb7fc-bc1a-44fa-9e25-e720b5f6ca86 · outbound

This paper cites Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.578943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.578943Z digest=sha256:42f27ba64b7f125bf9b019e644e2bfbff7558a31accc3855489b86c6cc2d800b

Observation 028e1dcd-ab43-4f06-921e-e162afc1fea9 · outbound

This paper cites Git Context Controller: Manage the Context of LLM-based Agents like Git.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Git Context Controller: Manage the Context of LLM-based Agents like Git

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.581794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.581794Z digest=sha256:ec5e86f01a64e067037fd97270d7c915ebd0461b3af7ddbe2a7a39465ebe9d5f

Observation 21703a96-ef58-4aba-924f-ca12aa1fea38 · outbound

This paper cites Scaling long-horizon LLM agent via context-folding, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Scaling long-horizon LLM agent via context-folding, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.584602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.584602Z digest=sha256:3d8ef5b5417657183707f9907f7ff28490fa52f7a79a93b03d0585711e4aaa93

Observation 27a8c101-fa00-4eed-94a8-fa0f4fb79d25 · outbound

This paper cites ACON: Optimizing Context Compression for Long-horizon LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ACON: Optimizing Context Compression for Long-horizon LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.587099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.587099Z digest=sha256:38164a8cce09ae4308d3c0f724f34181cb823a30b9c15d46fc3fe3bfb3196dca

Observation b5c86055-aa32-468c-95fe-eb38d94e9afe · outbound

This paper cites Diagnosing and Mitigating Context Rot in Long-horizon Search.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Diagnosing and Mitigating Context Rot in Long-horizon Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.589850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.589850Z digest=sha256:3e96a5c8fbe1896380165e601fb555223bdb057b6dae141b24f40e49ca4ecfd8

Observation 8bc86b62-d07b-43a3-9b2b-15d3a3ad1905 · outbound

This paper cites Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.592719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.592719Z digest=sha256:6576f752dd660bc54fd9db7162743c505c6486c95754f9a46ab27fdeae1a0370

Observation 416faf43-b239-422d-952c-75b229eb1b1b · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.595485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.595485Z digest=sha256:d677b081de6d49ffa4a26ddda6bcc4695f2b819ce2f1e3e868bd7070cca3a232

Observation c9386a00-719c-4362-9d2b-47e1ba67ab20 · outbound

This paper cites O’Brien, Carrie J.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents O’Brien, Carrie J

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.598199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.598199Z digest=sha256:ffc56775eea32793aa62fd251fb65394a4c7a356b1315765fae451a7886fa79b

Observation 70f1a38c-65e0-4521-ae94-a5d7bf04b98e · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.600632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.600632Z digest=sha256:f850fe24e79879f14cf529b1cfce9347791685676d7ddbd9e327a0d5f199897d

Observation edd121b7-034a-4807-bff7-c0f3132f3d96 · outbound

This paper cites Graph-based agent memory: Taxonomy, techniques, and applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Graph-based agent memory: Taxonomy, techniques, and applications

Reference 48

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.195059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.603257Z digest=sha256:b8d36d5f61256984c40f03e4616dab66edaf673fd3391fb8a5d184d588f7451d

Observation a691d46d-79fa-4b54-aa69-31ff878f77e2 · outbound

This paper cites WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.605771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.605771Z digest=sha256:2f9e407fd86beca131383c5bc41d102b4c5b8790c518e8394d03c070622b0fe5

Observation 3ad3c7ab-96a1-4af4-b83f-cfbd3fa8e789 · outbound

This paper cites When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.608319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.608319Z digest=sha256:160876f1bf75f2e8d02e2f04eac69acfa1c61352498c590faca4cdfe71a4e2cf

Observation 5a0df324-3392-4164-9585-507b372a872e · outbound

This paper cites MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.610938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.610938Z digest=sha256:2689c9db976521b622e4b889984750cbe51fe974eb508a7831ed46fc07e14364

Observation c9a1bb9a-942a-4db0-8214-6ed790172bda · outbound

This paper cites FadeMem: Biologically-inspired forgetting for efficient agent memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents FadeMem: Biologically-inspired forgetting for efficient agent memory

Reference 52

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.134245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.613747Z digest=sha256:8b43e76f47e4053936a756242b0797eee4d7746b496909217df2acf32ff13862

Observation f72a3db0-9a42-4860-b3e8-e028ae15a2fa · outbound

This paper cites MemPO: Self-Memory Policy Optimization for Long-Horizon Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemPO: Self-Memory Policy Optimization for Long-Horizon Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.616146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.616146Z digest=sha256:3473fe75eff468759b031cad6da5a8cf1552e23df2601a3421a73be225274e65

Observation 1c123bf1-e8e8-47e0-a896-32a2a2db64e0 · outbound

This paper cites Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.619058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.619058Z digest=sha256:7fe8f05b0365752409d07246cf2462d702198f28fb4ecb1caa15a603c3f69b15

Observation f03b548b-bc5a-4e8c-9968-3d6b5dfc9262 · outbound

This paper cites Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.621794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.621794Z digest=sha256:0d553eb5d3573a55546053add401d7f068957b6ed7f81888281024e1c21de61e

Observation 78ed0683-b93e-444e-9d7e-ff74ac0c4d53 · outbound

This paper cites Forensic Trajectory Signatures for Agent Memory Poisoning Detection.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Forensic Trajectory Signatures for Agent Memory Poisoning Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.624528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.624528Z digest=sha256:7fff2c8d4912432b653dd10aed30e7c1fd08dae388e5ac3a827f7bb93c2f181e

Observation fc37d485-bad3-4b9f-be01-ba609ce4a7be · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.627306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.627306Z digest=sha256:fd85a12886f8051657df24a11981fbae896d4267af3e6464d13014512139edf4

Observation 1a45079d-5852-4f74-bab1-0229a88e15c3 · outbound

This paper cites Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.630273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.630273Z digest=sha256:33b6223de1273ca3ef49c84d9d83e1e308281a3e97eb2f1897217a2672992805

Observation 1f63ca0e-1696-4977-8787-6c7a4866e897 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.633179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.633179Z digest=sha256:2757abd4dab25bdeac7f5d3da3e12d42e14ebfb988149fe88fa0b38fcb7907dc

Observation 3287544e-72b4-4d57-b3ce-b2c028532371 · outbound

This paper cites HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.636179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.636179Z digest=sha256:a8715e4123de06efb67d72e842ba95ad9397bb1eacb566a4b944031bf53e8003

Observation ec066b2d-7c9e-4b88-8e65-06046c159ad4 · outbound

This paper cites iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.638736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.638736Z digest=sha256:fd0322e43d7193f5768f892f35bcf2a537527de1335b4b7f1c29d0d60cbe8c08

Observation 45abc842-3186-4af0-bd7e-a95067b0d78f · outbound

This paper cites SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.641693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.641693Z digest=sha256:5ab42696a5bd4f53b0959b2ab7752ca31a7d96e8765a9e8b1cf888c2991aab62

Observation 901dfc73-dfa5-4c2d-843c-eef1836d35ea · outbound

This paper cites xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.647556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.647556Z digest=sha256:822dbb6071ebcf3ceb58fc2a63b518451091be989c8a5055b6a467ae7ba1b2c2

Observation 84dea266-5519-4e0f-8279-6bd3287d25fb · outbound

This paper cites Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.651564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.651564Z digest=sha256:0225f1c6847df165671d38896f561db0aa6099bc26465e4d8ca9ca671458dca4

Observation 691c522d-49e5-43a6-b8d7-ccb8cc017c5f · outbound

This paper cites Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.654095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.654095Z digest=sha256:2f7379d0610d878099e469ab4a8a58da7b2a6545f170375cc466543fe886587d

Observation 7dd39556-31b6-416e-919d-70fc6550e6e3 · outbound

This paper cites Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.656554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.656554Z digest=sha256:a70e5e695a1fed8fcd5a76fa6e59d87ba28ce58acbf4c35f92b5ed3a05eadbfb

Observation adb5527c-b2b1-417e-8002-2ce9847da674 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.659062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.659062Z digest=sha256:60347383706dcf5448fbddfc28ec462cb3822b445d360d137e2a7b67b5835990

Observation 63528579-e520-4bbe-a426-ccf7c5517246 · outbound

This paper cites OmegaUse: Building a general-purpose GUI agent for autonomous task execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents OmegaUse: Building a general-purpose GUI agent for autonomous task execution

Reference 68

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.047704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.662924Z digest=sha256:e8da67aca5d9ce939b5d7aadbf866da5d2140a410df4b147c3bdb55a8b0cdddc

Observation f3c0ffee-606a-4584-97cb-e4cd47405017 · outbound

This paper cites An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.665805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.665805Z digest=sha256:1c2caf61dec4dd56a2e8bfdc1bc8e950a835108fb06f0ddc1b28d51a90b400f2

Observation 42aa1d8f-1201-41e9-9a3e-02e021fd7c3e · outbound

This paper cites A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.875866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.668706Z digest=sha256:a92a5b0e604d4885cb608c77dbe0eec4c266dcc19f8fbbd992e08923e3d08fc1

Observation add4d526-bcf5-4272-b7cf-7238890e9cf5 · outbound

This paper cites Governing AI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Governing AI Agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.671459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.671459Z digest=sha256:2721983bbc0e31b38d5ec80e8d61b9dfee6e6153154ac0bb9ab8e9b63b731ca6

Observation c9ea7510-3b8c-49a0-852c-a570ba15ded3 · outbound

This paper cites The AI Agent Index.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The AI Agent Index

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.674505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.674505Z digest=sha256:f4f9e964b5afd00041df7bf0e33d309040e2eafe60d123557e1bcfd887b28350

Observation 5a9985a9-5b84-4abf-861f-88c391b0aa02 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Reflexion: Language agents with verbal reinforcement learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.677546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.677546Z digest=sha256:21325eaf3b4de9f4a078ec61a284203c8273cf10dd76bd26e201e3770c6821f1

Observation db5d97df-474b-4f8f-b5c5-96f76572501f · outbound

This paper cites Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.849405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.680054Z digest=sha256:c198a2955659360792a3c626ff8f3b93ee6212f8490b7e92992dfa7540750cc2

Observation 5753611a-4298-44c7-bd09-4acea5a5c805 · outbound

This paper cites Large Language Models Can Self-Correct with Key Condition Verification.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Can Self-Correct with Key Condition Verification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.682869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.682869Z digest=sha256:370d0d622de0358164d1e3aa3f2f32d7718603ec4fd1b32e229c1ca2914fa643

Observation 307c0cdc-04a0-48be-a362-52d169bf18da · outbound

This paper cites Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.685622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.685622Z digest=sha256:ff58a87b6640e0df41ffc051dbf9b6aa1bd9c758488ff77de2f86d1a9a4c01cf

Observation 24a7b499-63cf-473e-8a30-198ba7f86666 · outbound

This paper cites CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.688258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.688258Z digest=sha256:73ac7672b5822df564b993142e0bccea93f960b1e79808b04da693adc253c6da

Observation ac8685f8-54ae-4ee3-9784-32b8027f87f0 · outbound

This paper cites SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.691627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.691627Z digest=sha256:d2447e6772e1b162b1ca2d4420bb0493f4e96f11bf5a9892d106c5da19092e90

Observation 69cfafb1-596d-403e-bda3-fd2f5f9a17cb · outbound

This paper cites Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.694532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.694532Z digest=sha256:b9c9bd0b8fcf71a933a484b8ec513a1d90db7304eb36dcf839734c547a315867

Observation a1ee9040-cad6-44c8-914d-f865b460ce2f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.697245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.697245Z digest=sha256:74baccf7d8f211629fbc52a33fcbd2ec5da2dd1ade9196afe78733e3a2f8f1be

Observation 0011074b-293d-4309-a6bf-df0031a71baa · outbound

This paper cites ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.701005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.701005Z digest=sha256:0519a5e26bbcec4df70e9624aa5504c6109bde17ae200791ab324849592002a2

Observation e3fd0de4-9bc6-48f0-84fa-fad8f6ca66fb · outbound

This paper cites SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.703441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.703441Z digest=sha256:d86d51f0c785cc7f185b62b53c0ed21694d787c2d6fe549582fa099508508e55

Observation 5b467a41-9e0c-4344-bb23-a87414d2db54 · outbound

This paper cites Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.706348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.706348Z digest=sha256:627b879610471b7518156eeb577aef0c9ddb9c02000f1e1f6a94f4d1e47be0c5

Observation 341cc47c-06cd-4f9e-9f2f-7efe851364ed · outbound

This paper cites Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.708864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.708864Z digest=sha256:d8d11d5646366258859f9c5c22c1b10da5d176f2817eca334985985fa5ac6579

Observation c6470f50-8445-4a83-8d26-7cb5846d6b56 · outbound

This paper cites Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.558636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.712150Z digest=sha256:43954d667db0970812e1dfcf8a05fcffe8670873ec1a9007ea87215a89cfd9fa

Observation 18100c34-618d-4f6f-a317-ab6985c78114 · outbound

This paper cites MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:35.979608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.714955Z digest=sha256:bc184b63d0f6af38b675d736e05b40e42644a39dee262b57c7a0a1527a6d00e7

Observation 67b78875-b23f-4f73-99ed-7715ecf8436c · outbound

This paper cites Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.718280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.718280Z digest=sha256:654bccd5618d2c72a979f9f642899c5d782989219c9642d4fbc069353ff2d845

Observation 6f24c853-84c5-4563-a314-5993871c2b10 · outbound

This paper cites Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.721043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.721043Z digest=sha256:6618009ee92b60a607024fa67fd4e431079fdfbc8b9459bafaff0a466f75023e

Observation 1a06d184-2b6d-4231-b3cf-177e3d10968f · outbound

This paper cites Let's Verify Step by Step.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Let's Verify Step by Step

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.723496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.723496Z digest=sha256:e8d2f773f7c7a637e793f41a6344b236908abf4935e226c45d7d5520174e40e2

Observation d12b5d82-c357-4464-89c1-dcb0d289388e · outbound

This paper cites Entropy-regularized process reward model, 2024.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Entropy-regularized process reward model, 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.726497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.726497Z digest=sha256:7549c403126d11ede679c3e4cac11464e44c13551a37263351fb73cc269300fd

Observation f48e0268-e156-475a-ab03-ac7aadb5c4f9 · outbound

This paper cites GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.728986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.728986Z digest=sha256:f2bd2e6a912841b9b1cd8d9d6fe05d22f71bcdb3545ddfc70204fa453b19bfac

Observation f4d8f352-658f-4f61-951f-23098439453f · outbound

This paper cites GRPO is Secretly a Process Reward Model.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GRPO is Secretly a Process Reward Model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.731504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.731504Z digest=sha256:412300e90559809725ae840e762c50e7c6b2a70c6b305d71ee385b7bb6f6dd9d

Observation abcf2ec5-4a78-4c06-90b6-c33f2e7f18b6 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.734381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.734381Z digest=sha256:9dea16df5a8e226e40211e9f9e832ccdf6c4d545fd6d1247f5fe3d6002a57e56

Observation f6761413-445d-484c-8e28-5896ccde7a29 · outbound

This paper cites AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.737168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.737168Z digest=sha256:adccdcd385e453e4a24b90005e3b81066cd61d41aed3d9451332727a348c326e

Observation f5be3098-e6dc-41b4-a776-5bec79c7e34c · outbound

This paper cites SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.739622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.739622Z digest=sha256:2318918460627bf1eabb80fcaad377df1c1ec693277ec72a3fa9e5640b858b9a

Observation 5d856c48-6af7-4ea0-9cb8-12146571136b · outbound

This paper cites Agentic rein- forcement learning for search misaligns instruction-tuning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agentic rein- forcement learning for search misaligns instruction-tuning, 2025

Reference 96

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:01:37.234083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.742845Z digest=sha256:1bce07887ab10c40dc2fb4ff72b4ce7886163ed941ed6b9663a5c6e9d780f820

Observation 2169a623-151b-4748-b08f-1e94f0cacc5f · outbound

This paper cites Self-evolving LLM agents with in-distribution Optimization.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Self-evolving LLM agents with in-distribution Optimization

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.167467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.745579Z digest=sha256:ba4b26933ef8f4e8c8db4ec86530401c55ad48001573fa21dd270821b1aa8530

Observation 1b0478b4-4a35-464f-a68a-9a73bf9eea4b · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.749355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.749355Z digest=sha256:19c9e1ae17295b59834be3b7e13d7ca09ce6d93292fae7105e36e402ab6c0f33

Observation c40f7a68-d53b-4b54-92b5-d969088ad129 · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.755009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.755009Z digest=sha256:d926e7ec0b70b66ba2e2688ea4b15453637c25a782398d9619e6f9f54705ba72

Observation eabd52b4-b678-4be2-8d51-d1687327794f · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.758160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.758160Z digest=sha256:66b0167289aa78d63bd1596348f451581097ccdf5ca7611f8e8982beb4107d0b

Pith citing papers

No inbound Pith citation observations are available.