Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2607.05378.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T04:48:26.776852Z
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3e0b7f1f-21c6-4f78-b210-85ce19c6a25e · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7927d0d-817a-4c0d-8243-53ec0a377332 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a288b75-3e56-455b-8f45-c710aa0c68be · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Longformer: The Long-Document Transformer
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 023bb951-e804-417d-9340-f44cace5c373 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 045d7020-f785-4f0e-91f0-16c2aaa76866 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6123da0-4012-48b2-8d19-fb1ac55d7222 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2aa56711-e4d4-46a5-8abe-be2e62d0cd56 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55055d7b-fee2-4555-b94a-908aaff47fc9 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Scaling llm multi-turn rl with end-to-end summarization-based context management
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7ff7ccf-8645-4c87-8cde-57fa11fcedee · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Scaling llm multi-turn rl with end-to-end summarization-based context management
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df19cf28-7086-41d6-9002-eb6addeb02d7 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b84c590-9e9f-48f7-825e-ef3cd51ff423 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents MemGPT: Towards LLMs as Operating Systems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5c6f8f2-8623-4812-901f-ae9ee5edc6ad · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c193882f-a41e-4693-97d3-6eb68216b44c · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b5a9f53-0e91-4fac-b2f2-5eeb8d8377a4 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0cb92ccc-8750-4222-8b7b-396eba6f9d8e · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Scaling long-horizon llm agent via context-folding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8af86be-f3c9-4f83-8bdb-e65689f503f3 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Solving math word problems with process- and outcome-based feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac13f391-c494-4674-8ae9-2517d8748159 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Resum: Unlocking long-horizon search intelligence via context summarization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 874c9524-eb42-4c89-a0bf-5b3db07e460b · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Resum: Unlocking long-horizon search intelligence via context summarization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 799b7099-f76b-4e41-8096-63bf8c3c7a98 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9204ddb4-64b8-4377-918b-0875d07f1c33 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bbf5e80-5a18-4823-a0c3-2e3b10152ac7 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0504468-7755-4dd9-a45f-4f2ae2ccb475 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc0fbd2b-181a-47c9-86f7-6ca5775ab26b · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Group Sequence Policy Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a88d4ce-846d-43ce-9970-23ab6f4a0469 · outbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 881ff980-e98d-4b97-a9c8-8046ec27c3da · inbound
Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e24622-8465-421c-ac21-ea8a57905800 · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.