Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2505.20732.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:41.709278Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 3457b9dc-10c2-4919-a7e1-b14f2383e45d · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0b54aee-9a64-4042-8e4e-87c7a403a11f · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 239
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27fec0da-b0ed-4402-8749-072a4a15b231 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 304bfd78-5a1d-49be-85a6-c7d1143f7396 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 172
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e81255b8-ca8b-4fa4-a25f-17e737f2a235 · inbound
MASPRM: Multi-Agent System Process Reward Model SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de32141f-abe3-4509-93cc-45d894a952b1 · inbound
Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f2de34-a371-4ce2-b11a-eef553f488d6 · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9666a72f-dfef-43e1-bd3e-ee0228a6f0c9 · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50229213-8e0e-4377-b128-5992bc15bec6 · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 382e89b3-92ad-4b34-bb7b-3d76051a0f45 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d52c3d8e-7bd5-4694-94fe-a29cfe5534c4 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c3b6a66-ac9c-40cf-821e-dc58d0261018 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2934e963-22e7-4960-a723-0874c1b24a90 · inbound
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1433239-4e12-45fa-82cb-2a0b6dc58a1b · inbound
Learning CLI Agents with Structured Action Credit under Selective Observation SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ecac87d8-5a88-43b1-9ccb-d7f7f7d9cb13 · inbound
TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3855bf6f-9883-4cc3-a03b-cf51c2395fa6 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5acadc1f-90eb-4d04-9270-06b3eefa2206 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23f6867-d8b8-48f4-8732-18d302cf835e · inbound
Trust Region On-Policy Distillation SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c73e13c1-8c37-46c2-acd5-04a4f3a70e37 · inbound
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9f02781-c1f8-4734-b987-c1863f1b9afa · inbound
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8a41327-e0fb-49c0-96b9-f28ddc71507e · inbound
HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c18f4942-7020-4032-9661-2d3c217938cd · inbound
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8f816ba-2c2f-4e0f-921f-f10b18a962b6 · inbound
StepGuard: Guarding Web Navigation via Single-Step Calibration SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 570946be-579d-46af-94b6-074622e7402d · inbound
Learning with a Single Rollout via Monte Carlo Pass@k Critic SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 113cc57d-5f88-4b68-af43-ac634976f5e7 · inbound
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c434fbc-c40c-4bbe-92b9-e9566d9fb4e0 · inbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af49ee17-bcf4-4c1f-b102-7978cdce4208 · inbound
What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8148a7e5-e65c-4119-b98f-7bc318c68640 · inbound
RLVP: Penalize the Path, Reward the Outcome SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74c730fc-3178-4f5e-97f1-0176e0f8fc29 · inbound
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff0fa1f-36fc-4df7-9af0-d5877a854285 · inbound
Process Reward Informed Tree Rollout for Effective Multi-Turn RL SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69741546-e41a-4506-a1df-59b33b37d407 · inbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dc83e9-f27f-480a-9822-b56bb3da7a01 · inbound
CAST: Game Solvers as Turn-Level Teachers for LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.