Pith. sign in

Paper Citation Record · LEDGER

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08255 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:31.002411Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d90394ae-8bf4-421a-81a6-97acf69a6280 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.881986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.881986Z digest=sha256:14b1c9d06b106a23271d93f4f9390dc7bbeca919547390dd73ebb02985c13c1e

Observation 5dc7db82-d1bf-4bd7-a164-b663e62831fd · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.898822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.898822Z digest=sha256:196cc86acf4e162a1db0ba085e9d0631fedb798e917f092908adbf0e95e3c84a

Observation fb549041-a873-461a-8fb4-e4a36e6be90a · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.903845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.903845Z digest=sha256:0b5a1d513160267369a309e85dfd40d977677e7a8a8712c750b9e7346cec1b85

Observation 5f552ab7-7564-4b60-9ea0-0f450dc67aa4 · outbound

This paper cites Shuo He, Lang Feng, Qi Wei, Xin Cheng, Lei Feng, and Bo An.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Shuo He, Lang Feng, Qi Wei, Xin Cheng, Lei Feng, and Bo An

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.915025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.915025Z digest=sha256:2da294fe62e1e6dcf717450156bebac0ad02773486c674d13a035660bacf00f7

Observation f1bdb373-645c-4c2d-b478-0a7e8fc05de0 · outbound

This paper cites SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:20:31.418525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:20:30.920209Z digest=sha256:6e461391239d9e1788abba89f8e11f0e24c1c30eabfefc35aaa3d970991f5218

Observation 20d99a4f-11f2-4944-b987-753ae45f1b6c · outbound

This paper cites Sparse Rewards Can Self-Train Dialogue Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Sparse Rewards Can Self-Train Dialogue Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.924964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.924964Z digest=sha256:271b562cef1c41cda77cd08c6ad957a583811384083ff3d99902a426849694f2

Observation 951ad88b-d8f3-4151-8833-7524d42e3994 · outbound

This paper cites Let's Verify Step by Step.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Let's Verify Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.935549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.935549Z digest=sha256:1394b976a77fddc156c8d4840cad5644d2d0b487e321e30d59fe777674717204

Observation 0a1e9ed0-3451-4c16-857d-22bd399370dc · outbound

This paper cites A Survey of Temporal Credit Assignment in Deep Reinforcement Learning.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Survey of Temporal Credit Assignment in Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.941220Z digest=sha256:379df55df463ceb4e309392394c4171835ace568f348c1f560d8c259fb06d55b

Observation 9385ed41-af56-4b1d-8836-f1b0e782bdd0 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.951443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.951443Z digest=sha256:b449ace27b64a817dddfafb6b52001c888b14e06f807272b369ed1850fa5f071

Observation ef0a00af-47bd-490e-adbb-aca16bc14dad · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.956539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.956539Z digest=sha256:dac310b3c2ebee8aa359bc25a79df3a33628aec2445e6c79a9bee2d6c36989a3

Observation 284107c9-dd1b-4155-b113-ab125fd0a0fc · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.966366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.966366Z digest=sha256:8fbd93994375dbbfb501497790ee5980eebd89400193fbc2b461ff5b2cea0f89

Observation 15cddf84-14e0-4aa5-990a-35b31922c348 · outbound

This paper cites Chenglong Wang, Hang Zhou, Yimin Hu, Yi Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Chenglong Wang, Hang Zhou, Yimin Hu, Yi Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.977076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.977076Z digest=sha256:9f3e0fc2de967579c3edb9fa09a84137e554f5912b03620078109d08523bcc33

Observation 1e85fd38-0071-47a6-93d1-ab29797e7564 · outbound

This paper cites Qwen2.5 Technical Report.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Qwen2.5 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.982441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.982441Z digest=sha256:3d497c251f4e6fc9ef5721774e73b5fc687551c9b1c9ae519206648d48926a68

Observation ea9b2aac-c0ec-4774-ac53-d88079310eaf · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.987486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.987486Z digest=sha256:d78b9a1b34ca8e3b6b097eefab800144361aa65f5d54b13cc06ea0d3464dd92f

Observation af1ee5bb-ac15-47a5-a867-87075d555589 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.992560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.992560Z digest=sha256:c9fafe04e183b89b75e828e64e5a8a57b6557de124083232d188b2ea76b0bb96

Observation 47e83f44-986d-4074-9e94-65de6da1013a · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.997565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.997565Z digest=sha256:1dd9b0d827089ac17b52713de749714dad3a34fbeb96276823c2d36a662b09da

Observation 8165469f-0a38-4d57-b00c-5c60b6b73de1 · outbound

This paper cites SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:31.002411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:31.002411Z digest=sha256:2606e9677591b45ab2faab53e78fecd1b6faf3407359e5bf2cd8f29a7555ec39

Observation 40ac6398-d796-43b4-a6a3-86c19e1940d7 · outbound

This paper cites Hindsight credit assignment for long-horizon llm agents.ArXiv, abs/2603.08754,.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Hindsight credit assignment for long-horizon llm agents.ArXiv, abs/2603.08754,

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.972193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.972193Z digest=sha256:d118690678c5bab5b561f6770493f11e63cc4d8d6a7d7a30343cdbdf4a88ddf6

Observation d5caeca4-f423-4630-a461-424602baec20 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.961438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.961438Z digest=sha256:c38749f5f9eeec9cc64e57412f9b9941ae882e55a99a4fcc7380a6a801070c0e

Observation db603c01-77a3-440a-b6ab-8265efc89ac8 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.909215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.909215Z digest=sha256:79c3e71e6cde2e4b1077040f6fba061ab9a39bc5a3ae222426f645e84988c6d1

Observation d6092662-97d3-483f-9352-65d6c435a8a9 · outbound

This paper cites Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.930351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.930351Z digest=sha256:5e1643526e0afd27fea8efda3f5b9aac2046029ad8636f143df22827ff527da8

Observation e5ec0b2d-640a-441c-8c8d-848892cd796c · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.893043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.893043Z digest=sha256:2ca5c0e6d0eec6512108feebb59c77f942be34990abb81de6280ea762cd1e028

Observation bf9020c1-d71a-4fc9-9863-446c487e5f0a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.887706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.887706Z digest=sha256:37ea04082897c92a0380e76dce9c09d01ec21fa904890b3d33907daa943e2fde

Pith citing papers

No inbound Pith citation observations are available.