Pith. sign in

Paper Citation Record · LEDGER

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

As of 11 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2608.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06663 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.758160Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d91c039-ad49-4de0-8523-8e0ec040dd6f · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Measuring AI Ability to Complete Long Software Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.467237Z digest=sha256:2ce359d7911daa23df1ef0d51937f9fc4cdf7b80a3fc67e257405a344661bbe2

Observation 09b0131f-8e67-4138-8ced-bd0416dd7465 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Cannot Self-Correct Reasoning Yet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.471245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.471245Z digest=sha256:337e47623d0b0a2239a4a6c1dbc41d098c4fdfade724f999305f263578a57a3e

Observation 9a2f02d2-5ee1-4303-83f9-b10ec55959bb · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.474552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.474552Z digest=sha256:11d057ec0187b2c486e7466c03c825fa4554788527d7e8e3d4b1a0b34eceab1a

Observation c637a348-385d-4bdf-9277-48b0c9453c64 · outbound

This paper cites The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The SWE-bench illusion: When state-of-the-art LLMs remember instead of reason, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.477854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.477854Z digest=sha256:3b1343337cf840e59dd80c721adc61a17ec0dde2283ac7b38e2bf87b0d0e9bb5

Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.480747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.480747Z digest=sha256:a12264cef57db49d5be7a1b89c0b70e90717eb8151796dc78c1c245d4e153c69

Observation 79e13dfe-c8ed-4052-a90e-8d56c9c6b9c3 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Understanding the planning of LLM agents: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.483951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.483951Z digest=sha256:381c3916cfe14d622a0bcb658726e72758d9b4d8e8c11fcbb0ea342fe10ecd98

Observation 9b045725-a423-42aa-9b11-3b0314362d1d · outbound

This paper cites LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLMs as planning formalizers: A survey for leveraging large language models to construct automated planning models, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.487410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.487410Z digest=sha256:4af255adec7db10c64ee63be234e50e6d78f2f73c41b9f791aec4d1cb4eb4cfb

Observation 49569b95-0a7b-4b37-888f-b7de3ab610f8 · outbound

This paper cites A Survey on the Memory Mechanism of Large Language Model based Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on the Memory Mechanism of Large Language Model based Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.490120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.490120Z digest=sha256:07d598cd6f6214dac4c473e40f4c9947de56df5e8862170187a89f599aff0d72

Observation 5d276526-6a02-481d-8c49-59bec7991584 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.493279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.493279Z digest=sha256:c49da545b9fa960de7e63b5c8a6302d67acdbfa97c4ddcf793bfdd525fc9c8aa

Observation 05fff376-c46b-40a1-ac08-6742aabd3304 · outbound

This paper cites Large Language Model-Brained GUI Agents: A Survey.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Model-Brained GUI Agents: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.495876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.495876Z digest=sha256:30684155e83783f303f6c8b7d31770e9ac2e81be607ca7fab492e281fcce26d5

Observation e3c4aaa9-5303-41b1-9df1-04d0f9fb8a74 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.498394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.498394Z digest=sha256:9f68a46e0a22610dcda5ee2a056badcfcba47ea75c6f842d6fe4c4634e17ad8c

Observation 5aa6c2f5-88bd-4b5d-8583-c74358440bb0 · outbound

This paper cites A Survey on (M)LLM-Based GUI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Survey on (M)LLM-Based GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.501038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.501038Z digest=sha256:94f918dc0ff12eed6870d7bb6bc29cdca15ca6fbe3100ff87d706b8da6ec2694

Observation 79898d1e-c85f-4e7c-ac4c-01028f7fb264 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.503946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.503946Z digest=sha256:7e3c73d90ec031e6508f38dc6308e0dca3387483b4eca9386e443531a11af412

Observation 416228f4-611e-4317-816d-b6627376714b · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.506528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.506528Z digest=sha256:683c2d1d5182fe02cfbd94ec8d7f04da7953049ea52359b3a400c2ac50d87e18

Observation 495abdf0-ba25-461b-9bfa-39dccef6eade · outbound

This paper cites Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.509270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.509270Z digest=sha256:3a3d7f37dd17e4f755be6dd01fe77dade63ee321bee7157608889e69ed3562f1

Observation 7611fd6d-329b-4548-a39b-882f21ab0d74 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Sutton, Doina Precup, and Satinder Singh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.512030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.512030Z digest=sha256:33a441e59000b1ca90ee6230911c3456476489408c72397cc0f9244b62dd1c06

Observation c59fb5cb-4685-4af3-ad59-040a0e204946 · outbound

This paper cites an unresolved cited work.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.514691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.514691Z digest=sha256:a7ac7109e70124cedbf97736d0a03c6de46b8730d07a403e54f98b6d8cb395e8

Observation 8f23b146-909e-479d-ab6b-2f57f65a232f · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.517185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.517185Z digest=sha256:831f931fba3d4c9701da1a0306e8dc7103fb0873b31dee279dc0e004c6607168

Observation 0bc5a8c1-5005-47b7-a175-aabf23646121 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.520046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.520046Z digest=sha256:108f01283067cbdaddc8084aca9a912807110c8eb4b8ad255328b0ba31782669

Observation 805e7fd4-dd42-47f6-b2ea-372260a6d7cd · outbound

This paper cites PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.522839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.522839Z digest=sha256:0b80a352738462bec31c111ff2e1ebd05b95f9293cd380d9abf2145779e6b73d

Observation 1debe8c0-46e5-41b1-90a9-54954f4e4cb5 · outbound

This paper cites Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.525639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.525639Z digest=sha256:fe0b5563255e0029364e816ef10e2cd25c403569f2f31934e0ca614f873afbf2

Observation 09e9ca7e-7693-4537-99a7-81361682208e · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.528245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.528245Z digest=sha256:b16f36746825ae50bb388ed8a4f0f1287ad00e3082274852f78fb106d6b6f33a

Observation 0e3d99fc-a984-40f2-af62-86f7a7d69fcc · outbound

This paper cites PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.531161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.531161Z digest=sha256:9a7353d32f98c511db8a91cb094633bdd8bb9dc60fd82dd7bd001491c26333b6

Observation 80a59d1c-9c40-4e89-9904-6c726f451504 · outbound

This paper cites Harnesses for Inference-Time Alignment over Execution Trajectories.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Harnesses for Inference-Time Alignment over Execution Trajectories

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.533892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.533892Z digest=sha256:2e3708cc19fd824d54ffe9758a9a89ae046ea0c87bd4fa83f7bddcbffc6377f1

Observation e5b00e30-16ae-4bab-9c74-c69743b834ed · outbound

This paper cites PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.536440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.536440Z digest=sha256:b68e595660d47164e92618b5546720cf66bd8bdca74a383bff12bab00a82fb9a

Observation d32cb368-2985-4bdb-ba00-13b110e59b19 · outbound

This paper cites ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.539470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.539470Z digest=sha256:b34f9dc9b4f612b59681366c1f425536060e60207aeb53cb7351ad7f2742ae9e

Observation 58ba2e0e-129a-4325-9afb-b97ab8098bce · outbound

This paper cites Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.542342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.542342Z digest=sha256:87e703f4faec1e4f0c09b955a0a005b19f7fa5250706abe0077fc119f3d62236

Observation e2076d76-067d-49b0-b556-169498d713be · outbound

This paper cites The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The cognitive bandwidth bottleneck: Shifting long-horizon agent from plan- ning with actions to planning with schemas, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.545080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.545080Z digest=sha256:09b090dd0db45c7822c22429291165ad8f7f2557052a17c21a72b5c3e9f8cf38

Observation 4d9aaf94-2e8c-48c3-9c7b-37f5cb63231d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.547554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.547554Z digest=sha256:ffc57aef7506b35e3097cccbcd55bc55fd786d7cb54c4ac4dce3b225d2027b2b

Observation 9b1e3e64-304b-4287-bd15-f678d3f508d2 · outbound

This paper cites CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.550199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.550199Z digest=sha256:060312d15413ffeee47b7079229570df87daaf9d5213f642aa2679ea016145fb

Observation b223cb23-1271-4f9b-a6a2-139b36034026 · outbound

This paper cites SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.553083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.553083Z digest=sha256:6b2b475d5248c9038c189b7719ddd67fecfa54363a1b88e096cf744d70490772

Observation f3ed3bc5-cb75-4177-b4a5-89dd8e52bbe8 · outbound

This paper cites W ALL-e: World alignment by rule learning improves world model-based LLM agents,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents W ALL-e: World alignment by rule learning improves world model-based LLM agents,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.555658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.555658Z digest=sha256:53db960ccceaca47ade0a7706ba9f574a6e8ac26b70433712ee67cc1602b73c9

Observation 33743249-9f5b-4ae8-bcc4-d377b3b30ab2 · outbound

This paper cites Laird and Corey Clark.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Laird and Corey Clark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.561984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.561984Z digest=sha256:539805c05ed7e9f3b1a9b60e5422985a13058d92d17ff8bdda72d1d55e8691a0

Observation dd8b4f10-3641-48e2-acb3-f6ea52dcf7db · outbound

This paper cites MobileDreamer: Generative sketch world model for GUI agent, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MobileDreamer: Generative sketch world model for GUI agent, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.564416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.564416Z digest=sha256:cd802dafc7d2faabbb7fef951634392cdaf35331e5a6ab70075ded9fd2940037

Observation 0fdff7de-756d-46cb-b1a5-3f3f35638a9b · outbound

This paper cites ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.567093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.567093Z digest=sha256:473b6502c6a862e31fb14db825d18684bb77a01f290f0eb54d554c63b1a48f11

Observation 304c4095-52cc-4303-8b03-781ba31759b7 · outbound

This paper cites AgentEvolver: Towards efficient self-evolving agent system, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentEvolver: Towards efficient self-evolving agent system, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.570037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.570037Z digest=sha256:3aafe596b4e32da7369a4d0a0bd1e080e4f0e0c528553867e9ddfd0a1e26c482

Observation a2ae34fd-a085-46a5-9538-a4e50a4f620d · outbound

This paper cites Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.572747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.572747Z digest=sha256:e219f5d52ee9811c811bf0ac9ac81370bbb089bce374117cf21f9be83d07e04d

Observation af5c6012-b77a-43f6-98a7-b8f0140b360f · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lost in the Middle: How Language Models Use Long Contexts

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-10T23:01:35.575978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.575978Z digest=sha256:cbe36a2ee431c28179ebaad8952aee513659faea3394dfc372e36e0aef42e666

Observation 42ccb7fc-bc1a-44fa-9e25-e720b5f6ca86 · outbound

This paper cites Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.578943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.578943Z digest=sha256:714df2019fca2c21eb52b3a61f205362fa9feb1cb78d46361c6b776bcec24ac3

Observation 028e1dcd-ab43-4f06-921e-e162afc1fea9 · outbound

This paper cites Git Context Controller: Manage the Context of LLM-based Agents like Git.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Git Context Controller: Manage the Context of LLM-based Agents like Git

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.581794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.581794Z digest=sha256:1340f3b6216cc3f85232bfc239d09843108d5b58ad131b06284bcf843cb42d55

Observation 21703a96-ef58-4aba-924f-ca12aa1fea38 · outbound

This paper cites Scaling long-horizon LLM agent via context-folding, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Scaling long-horizon LLM agent via context-folding, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.584602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.584602Z digest=sha256:3f22c1116b04e3d58aa2f23ad1c638efb2475791b964a8980afe854ed6c39139

Observation 27a8c101-fa00-4eed-94a8-fa0f4fb79d25 · outbound

This paper cites ACON: Optimizing Context Compression for Long-horizon LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ACON: Optimizing Context Compression for Long-horizon LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.587099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.587099Z digest=sha256:47bc138886806b0cd9cff7c2a8b84624e39a5f8f4effbf89520597f6c8a032c0

Observation b5c86055-aa32-468c-95fe-eb38d94e9afe · outbound

This paper cites Diagnosing and Mitigating Context Rot in Long-horizon Search.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Diagnosing and Mitigating Context Rot in Long-horizon Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.589850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.589850Z digest=sha256:a08848dda5e50e7afea0e228a087e1e95da2f37cecc21b46e90255f892492220

Observation 8bc86b62-d07b-43a3-9b2b-15d3a3ad1905 · outbound

This paper cites Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.592719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.592719Z digest=sha256:1f7113d79a9d33effd6c56c22815da1be375874f9f04c208fdcfe668f4866d59

Observation 416faf43-b239-422d-952c-75b229eb1b1b · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.595485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.595485Z digest=sha256:6831432185c1ee2e03e7d648fe9d55c72196db1237fc5a0f62ba8dd35b11d318

Observation c9386a00-719c-4362-9d2b-47e1ba67ab20 · outbound

This paper cites O’Brien, Carrie J.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents O’Brien, Carrie J

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.598199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.598199Z digest=sha256:6313c472280838637a3a27fb54d77b54ef717cf9f8ade95c2d7163156a513c66

Observation 70f1a38c-65e0-4521-ae94-a5d7bf04b98e · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.600632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.600632Z digest=sha256:37001ef57edc56302afcc96ce2aa40c1a595867ce3f58b00f6de9dff7fab06e8

Observation edd121b7-034a-4807-bff7-c0f3132f3d96 · outbound

This paper cites Graph-based agent memory: Taxonomy, techniques, and applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Graph-based agent memory: Taxonomy, techniques, and applications

Reference 48

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.195059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.603257Z digest=sha256:775702cbccf1dd75cd17731b22f41e81cfc615f4080687b77a72141a722da30f

Observation a691d46d-79fa-4b54-aa69-31ff878f77e2 · outbound

This paper cites WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.605771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.605771Z digest=sha256:c8a67b57acfee50878b9a8ad1c982f1f498c0210c6aca05e33667b9b423c7e86

Observation 3ad3c7ab-96a1-4af4-b83f-cfbd3fa8e789 · outbound

This paper cites When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.608319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.608319Z digest=sha256:01da54e049952ebc89613283b752cb3f537bbfa7647bcd577d3a40e6a9016e90

Observation 5a0df324-3392-4164-9585-507b372a872e · outbound

This paper cites MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.610938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.610938Z digest=sha256:8a175eb00c116e530766790c30100530faf4dc01346344bdc2fe5d7e2aef327c

Observation c9a1bb9a-942a-4db0-8214-6ed790172bda · outbound

This paper cites FadeMem: Biologically-inspired forgetting for efficient agent memory.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents FadeMem: Biologically-inspired forgetting for efficient agent memory

Reference 52

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.134245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.613747Z digest=sha256:182e32712e78f70ec33c2630474f917272d3fe21871795268cb3c2f49f9a28f7

Observation f72a3db0-9a42-4860-b3e8-e028ae15a2fa · outbound

This paper cites MemPO: Self-Memory Policy Optimization for Long-Horizon Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MemPO: Self-Memory Policy Optimization for Long-Horizon Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.616146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.616146Z digest=sha256:3ef02d574622b74c55466d5b9e4c6a16ae83a7fb9e07bee81f0bca520ed16586

Observation 1c123bf1-e8e8-47e0-a896-32a2a2db64e0 · outbound

This paper cites Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.619058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.619058Z digest=sha256:811e10fad2315261996c085b9bb6a16e7d72baf25c2f951a27d3d68ba4582ce0

Observation f03b548b-bc5a-4e8c-9968-3d6b5dfc9262 · outbound

This paper cites Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.621794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.621794Z digest=sha256:faedebf446efe8cfb735c4c80c31e8f7978135f7ceed414212bea80e0d070bc2

Observation 78ed0683-b93e-444e-9d7e-ff74ac0c4d53 · outbound

This paper cites Forensic Trajectory Signatures for Agent Memory Poisoning Detection.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Forensic Trajectory Signatures for Agent Memory Poisoning Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.624528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.624528Z digest=sha256:a3b2b69ddd8d3fa0a5fee73b7cc267961c3549bb536f5ce4103dd3646d043b40

Observation fc37d485-bad3-4b9f-be01-ba609ce4a7be · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.627306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.627306Z digest=sha256:9a3a5c4255601b821fc2468e1078a5a0905078c79498bf617a0eb62c099b4e12

Observation 1a45079d-5852-4f74-bab1-0229a88e15c3 · outbound

This paper cites Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.630273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.630273Z digest=sha256:912a80ee35f6b22e5c6c65b6dffd8c2015e2d062afe46e38fc8689bdc4acc477

Observation 1f63ca0e-1696-4977-8787-6c7a4866e897 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.633179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.633179Z digest=sha256:e3c15755e252bd44f6b90138e2bc695fd83e60a9672d2bd2c08d2a774814e247

Observation 3287544e-72b4-4d57-b3ce-b2c028532371 · outbound

This paper cites HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.636179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.636179Z digest=sha256:ab1aad15c71e5accb32e92cc126bbb8c5c27ff3687ffbeaa51086dcba78fa7f9

Observation ec066b2d-7c9e-4b88-8e65-06046c159ad4 · outbound

This paper cites iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.638736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.638736Z digest=sha256:e8a17edb488ff4f06af822b990ffdf9ac75d489d24bf8474e9aa029c65797102

Observation 45abc842-3186-4af0-bd7e-a95067b0d78f · outbound

This paper cites SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SentiMM: A multimodal multi-agent framework for sentiment analysis in social media,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.641693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.641693Z digest=sha256:6d1fc7c05acf1b32c63ffa7020a3fa5115041829ad963e22c5169181e71d7115

Observation 901dfc73-dfa5-4c2d-843c-eef1836d35ea · outbound

This paper cites xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.647556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.647556Z digest=sha256:41ddd7825b7afe4b61355c4cc9412c0479518da31f4b572bcc4f563738898091

Observation 84dea266-5519-4e0f-8279-6bd3287d25fb · outbound

This paper cites Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Ho, Anna H

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.651564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.651564Z digest=sha256:b298d737e773a00b9de09223910b3109bbbb8595aadb870827a92da5229003c8

Observation 691c522d-49e5-43a6-b8d7-ccb8cc017c5f · outbound

This paper cites Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Generative AI-driven hierarchical multi-agent framework for zero-touch optical networks, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.654095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.654095Z digest=sha256:6cae501e49bbb13d8a38301b701b01a318e8270563ea337ee012a064b060998c

Observation 7dd39556-31b6-416e-919d-70fc6550e6e3 · outbound

This paper cites Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-agent LLM orchestration achieves deterministic, high-quality decision support for incident response, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.656554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.656554Z digest=sha256:bad934566dfdebe6f7b548b4ba88dd6e1774a355bf35bb1e76b862666c34433b

Observation adb5527c-b2b1-417e-8002-2ce9847da674 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.659062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.659062Z digest=sha256:f6eb71ea96da0e8f5e5deb5f61de53d395c27bc7602ae902270d1b4fc5f8ca17

Observation 63528579-e520-4bbe-a426-ccf7c5517246 · outbound

This paper cites OmegaUse: Building a general-purpose GUI agent for autonomous task execution.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents OmegaUse: Building a general-purpose GUI agent for autonomous task execution

Reference 68

Resolution
verified exact
doi, observed 2026-08-10T23:01:36.047704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.662924Z digest=sha256:4ad1b8e61b26cf31ec3455be93b86614bfc75113a293864fdd491d02ae80f888

Observation f3c0ffee-606a-4584-97cb-e4cd47405017 · outbound

This paper cites An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.665805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.665805Z digest=sha256:3c002b7dfab7b8f37fd5ac3bb56f456a15a45586fd34e47f2382ffd6a7bb5274

Observation 42aa1d8f-1201-41e9-9a3e-02e021fd7c3e · outbound

This paper cites A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.875866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.668706Z digest=sha256:e8b18b6bc0477171192493b3c743cf9a38dee24300d5be4ee3ae638d8a414929

Observation add4d526-bcf5-4272-b7cf-7238890e9cf5 · outbound

This paper cites Governing AI Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Governing AI Agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.671459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.671459Z digest=sha256:c403653d6d17fe7ef2bdbfe35df4c2ff0d478bde1ffe4739c00faa785bfa1f43

Observation c9ea7510-3b8c-49a0-852c-a570ba15ded3 · outbound

This paper cites The AI Agent Index.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents The AI Agent Index

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.674505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.674505Z digest=sha256:29c701a941c7d2c9cf4b853ef5c1687422e077b181071f0c4c4caabc9d45ea9b

Observation 5a9985a9-5b84-4abf-861f-88c391b0aa02 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Reflexion: Language agents with verbal reinforcement learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.677546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.677546Z digest=sha256:9b2f4e76a5990e8c6752ea7ac05586be28abd1b3df56b018210eae99933b93cb

Observation db5d97df-474b-4f8f-b5c5-96f76572501f · outbound

This paper cites Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.849405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.680054Z digest=sha256:c10d58c861917093ecd8ec39186d7eda014441f782108e630de9a0069e73d8d7

Observation 5753611a-4298-44c7-bd09-4acea5a5c805 · outbound

This paper cites Large Language Models Can Self-Correct with Key Condition Verification.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Large Language Models Can Self-Correct with Key Condition Verification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.682869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.682869Z digest=sha256:ef5c2455060fb662c40ebfe8d65da52d688973dadd891e8ed15cf82423ba899c

Observation 307c0cdc-04a0-48be-a362-52d169bf18da · outbound

This paper cites Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.685622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.685622Z digest=sha256:e2b91689c529883ee85b1411acc35902bec441f8bfd88659c3694b981d5613b8

Observation 24a7b499-63cf-473e-8a30-198ba7f86666 · outbound

This paper cites CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.688258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.688258Z digest=sha256:5c809ac83373738b3d4095239b266021aa3500c3da80c4a7a38517ef4dc68594

Observation ac8685f8-54ae-4ee3-9784-32b8027f87f0 · outbound

This paper cites SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.691627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.691627Z digest=sha256:8d2634f02aab5b6bf5a5d2fd7097eaaa2f0f1844d61b16376f67932998345ea2

Observation 69cfafb1-596d-403e-bda3-fd2f5f9a17cb · outbound

This paper cites Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond entangled planning: Task-decoupled planning for long-horizon agents, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.694532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.694532Z digest=sha256:54ebc1899f3f18a0dd96756cbead5619e3c04478653744dd263cec5a2d12e0b2

Observation a1ee9040-cad6-44c8-914d-f865b460ce2f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.697245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.697245Z digest=sha256:11a9d02d4580cd7ae385725578470303a5524b38235e9624fd85a8301b723e3e

Observation 0011074b-293d-4309-a6bf-df0031a71baa · outbound

This paper cites ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents ViReSkill: Vision-grounded replanning with skill memory for LLM-based planning in lifelong robot learning, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.701005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.701005Z digest=sha256:646a808bcd4acbf5d43cec253bf1ac1c7ade8a3b0fad86e8666cf47c3821e334

Observation e3fd0de4-9bc6-48f0-84fa-fad8f6ca66fb · outbound

This paper cites SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.703441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.703441Z digest=sha256:7a2a92553c69ecbeeedb86aa269bd9f741d226aa836ca019fb18001b640a0643

Observation 5b467a41-9e0c-4344-bb23-a87414d2db54 · outbound

This paper cites Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Segment policy optimization: Effective segment-level credit assignment in RL for large language models, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.706348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.706348Z digest=sha256:d238bc33f65ea6dc71e1ce0e9a8d6d14b2b560a86dc99d7a80e615574b01cecd

Observation 341cc47c-06cd-4f9e-9f2f-7efe851364ed · outbound

This paper cites Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.708864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.708864Z digest=sha256:75ec8d2774199842f21f3fb8d717e8222cb8980a45076ab5789c8a0fccc7db21

Observation c6470f50-8445-4a83-8d26-7cb5846d6b56 · outbound

This paper cites Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.558636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.712150Z digest=sha256:ea2dc83accd8369f0158faf366e9cc47fb6febd68b6134234a2a4a8476023345

Observation 18100c34-618d-4f6f-a317-ab6985c78114 · outbound

This paper cites MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:35.979608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.714955Z digest=sha256:f47e9232f86f9e5bb8868697ba72e26cf76db7166dc0df7d1132c1d45d4c957a

Observation 67b78875-b23f-4f73-99ed-7715ecf8436c · outbound

This paper cites Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.718280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.718280Z digest=sha256:99485b785836486389f97ebb34e66e13ec9fe3d24439a9336f466b5e9b46d3e6

Observation 6f24c853-84c5-4563-a314-5993871c2b10 · outbound

This paper cites Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe, 2026

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.721043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.721043Z digest=sha256:c5737a0e7c16e6a392e1b0e081ed47875237c3c32236d8c3c0ff510a222aa5fe

Observation 1a06d184-2b6d-4231-b3cf-177e3d10968f · outbound

This paper cites Let's Verify Step by Step.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Let's Verify Step by Step

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.723496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.723496Z digest=sha256:1562e47c430ddb53a8fcf7afd6f3a377882e34c23315ce7ba6a7b1e20f4c55e4

Observation d12b5d82-c357-4464-89c1-dcb0d289388e · outbound

This paper cites Entropy-regularized process reward model, 2024.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Entropy-regularized process reward model, 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.726497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.726497Z digest=sha256:7cc7b6cca1784d1775b7cd0659d3a4510b27eaa28b8fb62ef4420db3a63ff1fa

Observation f48e0268-e156-475a-ab03-ac7aadb5c4f9 · outbound

This paper cites GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GroundedPRM: Tree-guided and fidelity-aware process reward mod- eling for step-level reasoning, 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.728986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.728986Z digest=sha256:1129e57de1121c26285291f32cb200051735ed7af43236bc246390ede9e432b8

Observation f4d8f352-658f-4f61-951f-23098439453f · outbound

This paper cites GRPO is Secretly a Process Reward Model.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents GRPO is Secretly a Process Reward Model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.731504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.731504Z digest=sha256:3f4f0b216c2d8a0fe4871fc62f0f88cd2d20d2c7a861e91ea62115464e8d3e93

Observation abcf2ec5-4a78-4c06-90b6-c33f2e7f18b6 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.734381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.734381Z digest=sha256:9e6755736103281e7afd9efd9d85b2934bbbe20c11305ce748fc9c16e72913ee

Observation f6761413-445d-484c-8e28-5896ccde7a29 · outbound

This paper cites AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents AgentPRM: Process reward models for LLM agents via step-wise promise and progress, 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.737168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.737168Z digest=sha256:ab4115c438eed378f21b0462bd284ad0cc7a6f6e6a17c1a1a13b0a1b9e00b95f

Observation f5be3098-e6dc-41b4-a776-5bec79c7e34c · outbound

This paper cites SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.739622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.739622Z digest=sha256:a8c2e3406846452ffe12af48a431fba85666e20bf76badbd839bab0ebcd4f5eb

Observation 5d856c48-6af7-4ea0-9cb8-12146571136b · outbound

This paper cites Agentic rein- forcement learning for search misaligns instruction-tuning, 2025.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Agentic rein- forcement learning for search misaligns instruction-tuning, 2025

Reference 96

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:01:37.234083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.742845Z digest=sha256:be0cb4a447fa7ddac0e14b7719ee6a8c8c26a4ee448cc7fe85182e51b96b168a

Observation 2169a623-151b-4748-b08f-1e94f0cacc5f · outbound

This paper cites Self-evolving LLM agents with in-distribution Optimization.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Self-evolving LLM agents with in-distribution Optimization

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:37.167467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T23:01:35.745579Z digest=sha256:ab7fd21d1de9e0aba96199b2a7615d50db745544b1df32b9fee4c6e183dca725

Observation 1b0478b4-4a35-464f-a68a-9a73bf9eea4b · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.749355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.749355Z digest=sha256:ef3ca7faa4e79836beefef37bbd212dbf22edb221260167eacdb0fbd3f7aa321

Observation c40f7a68-d53b-4b54-92b5-d969088ad129 · outbound

This paper cites SWE-bench-java: A GitHub Issue Resolving Benchmark for Java.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.755009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.755009Z digest=sha256:e3d22cdb58e17e131616d74d4c5ca975ba2d0d7fe9b251568adfb0326be6dfed

Observation eabd52b4-b678-4be2-8d51-d1687327794f · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.758160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.758160Z digest=sha256:ac021c0d13cc62d0ff4694c4bd1ddb9d79f4f8ebdd32042e6efeb7cd6c8a651f

Pith citing papers

No inbound Pith citation observations are available.