Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2505.15107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:20:40.072107Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T13:04:56.727934Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5f9244b7-87c3-4e89-b3db-e8c95144cc37 · inbound
Group-in-Group Policy Optimization for LLM Agent Training StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82c3f128-cd6b-4ca1-b174-dfc382795218 · inbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a1e444-233c-42af-bd0b-a92e31b51857 · inbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1666c2-d673-4616-a540-ed43729ee018 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 280
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e9fd6de-99b2-4e1f-b9e1-580c3c687b17 · inbound
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bdd4d1db-c793-41da-ab9e-65725839cfbe · inbound
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e100f0-2904-4f0b-aea0-c0055fab618a · inbound
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dda22789-8f60-40e1-b75b-0e9db858154c · inbound
Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89acafe-afeb-498e-9828-b297d9155d1f · inbound
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90f0cf76-e5e0-4776-9f45-f82b1adae920 · inbound
$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a715c2fd-ff47-4ab7-94d8-a02b8f4ac245 · inbound
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62eb5c5d-e137-4c90-b007-1781216907f4 · inbound
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89303f49-90db-40e4-9daf-e8468738297e · inbound
Harnessing LLM Agents with Skill Programs StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5554b0c3-3278-4092-a2dc-63a5476a0f77 · inbound
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04b16e29-d8c8-413a-88da-6f308ad9e551 · inbound
Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5a9e757-4c28-4535-bde2-fbadbae3ab0f · inbound
Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67705cae-0a56-41ee-9152-ac4e29548a0d · inbound
Trust Region On-Policy Distillation StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 80832a9b-e435-4a00-afdf-f06b668aa714 · inbound
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae6b84d7-f4b5-4510-9421-cecfa4c3c4cb · inbound
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffaea00e-6581-4880-a5f1-68dd30ab359c · inbound
Co-Evolving Skill Generation and Policy Optimization StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 538e4a03-48c6-4322-8bc7-350f39348636 · inbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d52464a9-a930-4b0c-8ca9-24415466a351 · inbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91115b55-54b5-4efc-b863-5981c9d9aaba · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1b21114-9d2a-4b4f-aed3-ec6ed1c30fda · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2409857c-a3f5-4f08-aab0-b5a6b817a4cd · inbound
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 733fe686-4d61-4648-8494-88de7018c776 · inbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10c737c3-645d-49ab-9e8a-d5b20ac9ffb9 · inbound
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab23ebee-b632-4e80-ae75-1775ba2bac1d · inbound
EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c9d078-39ed-4e86-877e-921490f8de89 · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.