Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:31.002411Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08255.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:31.002411Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d90394ae-8bf4-421a-81a6-97acf69a6280 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc7db82-d1bf-4bd7-a164-b663e62831fd · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb549041-a873-461a-8fb4-e4a36e6be90a · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f552ab7-7564-4b60-9ea0-0f450dc67aa4 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Shuo He, Lang Feng, Qi Wei, Xin Cheng, Lei Feng, and Bo An
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1bdb373-645c-4c2d-b478-0a7e8fc05de0 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20d99a4f-11f2-4944-b987-753ae45f1b6c · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Sparse Rewards Can Self-Train Dialogue Agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 951ad88b-d8f3-4151-8833-7524d42e3994 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Let's Verify Step by Step
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a1e9ed0-3451-4c16-857d-22bd399370dc · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9385ed41-af56-4b1d-8836-f1b0e782bdd0 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0a00af-47bd-490e-adbb-aca16bc14dad · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284107c9-dd1b-4155-b113-ab125fd0a0fc · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15cddf84-14e0-4aa5-990a-35b31922c348 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Chenglong Wang, Hang Zhou, Yimin Hu, Yi Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e85fd38-0071-47a6-93d1-ab29797e7564 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Qwen2.5 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9b2aac-c0ec-4774-ac53-d88079310eaf · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af1ee5bb-ac15-47a5-a867-87075d555589 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e83f44-986d-4074-9e94-65de6da1013a · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8165469f-0a38-4d57-b00c-5c60b6b73de1 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ac6398-d796-43b4-a6a3-86c19e1940d7 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Hindsight credit assignment for long-horizon llm agents.ArXiv, abs/2603.08754,
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5caeca4-f423-4630-a461-424602baec20 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db603c01-77a3-440a-b6ab-8265efc89ac8 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6092662-97d3-483f-9352-65d6c435a8a9 · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ec0b2d-640a-441c-8c8d-848892cd796c · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9020c1-d71a-4fc9-9863-446c487e5f0a · outbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.