Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T08:51:10.098370Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2607.07508.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T08:51:10.098370Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.548470Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T00:16:33.996094Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 533f0531-72c7-4a80-be28-209459ab1997 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d71f88d-b416-485d-b658-4db627090fae · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1535a6b-7bb4-4c04-9275-05960d480bab · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62ee85c2-6cbe-4527-a66b-b2854a66f22e · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a1b20da-e3a2-4b93-8472-8b2f65e442fb · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Acme: A Research Framework for Distributed Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb616be2-e8c4-4786-b318-de0a9c8e7843 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b458b75f-a787-4470-b38d-a48f0b94318b · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Let's Verify Step by Step
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de2b4abc-161c-4bc5-8aa5-27fba8e177cd · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd9748e5-9bdf-404b-b861-64a172b03968 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f38ed09-ba6e-4358-b6de-8abe62e05649 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning V olodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 868684ee-bdfe-4db4-b959-e46c40df7f67 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb2a4a22-1319-451c-813d-68375c26c8ff · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6880e6f-ff2c-4e2b-9cc3-b29ec19eb916 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0b09ed3-9e6a-430e-bab5-aed6b585c98f · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c3d12bc-f903-4de7-8813-003e0a82207d · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Every step evolves: Scaling reinforcement learning for trillion-scale thinking model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d802991-aa15-4f98-a56c-7d128d623fa6 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1627bf83-e4d9-4f99-87bf-caf7d7137fad · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb48c571-9088-411c-987c-19551079569c · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2edcb82-d888-4486-af58-067824bc5ac7 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e0b0ac7-cdff-448e-a7e1-ef20f930a758 · outbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Group Sequence Policy Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 884b4557-7dad-476e-b828-55a4aa438cee · inbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc26dc60-50b0-4b78-87d0-2c2eac0d56a0 · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e750cc-d962-4755-b375-d5cd91e43c79 · inbound
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54bc4f66-0d05-4dd2-92e8-85cace2ffe15 · inbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.