Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2509.08755.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:17:44.662944Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation ab033493-353e-4c3c-86e5-a01590349bc5 · inbound
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1bac4949-4614-4099-84cb-bf272862a396 · inbound
Graph-Enhanced Policy Optimization in LLM Agent Training AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ecc095-b3c4-4c04-8fd5-45bd71d0c2f1 · inbound
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 521c8bd2-31d2-4ddf-bc5f-0bfa840077e2 · inbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62076d23-5f87-47e3-b40e-a118ffdc1d62 · inbound
HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80e94209-9660-4f03-9d68-3a17f887eee7 · inbound
Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a0a00fb7-0848-4b85-af89-ff144e01af2a · inbound
Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f2fe3ee-5043-4fd1-86a5-ce7f9e1e0052 · inbound
Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d56ae31f-1127-4d63-8d06-b6f79feefa8a · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 62f24da8-3e94-423f-9baa-ad1e2f5b43ee · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e85163d-8550-442e-9751-e64bd6830b91 · inbound
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8088afa4-dd0d-4fa3-8c33-a1292a4c46b5 · inbound
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6a7d221-1648-47b6-aefe-2964ec339729 · inbound
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f9d5f345-c44a-4e88-8985-0fe8803df340 · inbound
GRAFT: Graph-Tokenized LLMs for Tool Planning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d857d5c2-67c7-46cb-a55c-bc15d944de6b · inbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 199a2e9b-eb9a-4a3a-940b-4e6a27353c78 · inbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72827b3b-f976-4a54-9e0e-970c19d50e4d · inbound
PriorZero: Bridging Language Priors and World Models for Decision Making AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 468889c8-af3e-4213-baa3-d9b4c468334a · inbound
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1ac94f0-db25-4541-9ebc-593bb43dfd4e · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b59c1e80-b311-4087-a119-9dbbe090a57a · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d583fbf-e3c7-403d-b117-3b59177e3305 · inbound
SkillGrad: Optimizing Agent Skills Like Gradient Descent AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b1e739c-bb86-4664-9402-992f048bad50 · inbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763e6f28-3062-4056-96fa-b9052d9b42fb · inbound
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a45e017-0649-429c-b45c-3054a1e94154 · inbound
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1934517f-cac9-4a64-9ac3-74851e6dce68 · inbound
When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2c48f9-e951-48b3-8d4a-71238d638059 · inbound
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.