Pith. sign in

Paper Citation Record · LEDGER

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2607.07508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07508 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T08:51:10.098370Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.548470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:16:33.996094Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact18
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 533f0531-72c7-4a80-be28-209459ab1997 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.437175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:228d34ed09cc84bece9fe0910a85807cad73d18f229ec857db3488180ffce1dd

Observation 2d71f88d-b416-485d-b658-4db627090fae · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.422874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:cfcee7ac7381107f9b14c345190f49242c65dad9b370a98d6c9d685b91379417

Observation b1535a6b-7bb4-4c04-9275-05960d480bab · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.425617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:a6df57d8464344817748759a9cca10f0ed8728264fa750c26845002ed4e4d520

Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.428422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:b7e3655b6b96d955704c248e612c35d93de00971ab4bef08da4b96b3453bd883

Observation 62ee85c2-6cbe-4527-a66b-b2854a66f22e · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.399622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:f8a768150c7e766403a654089f7c9579026c71546c2e5e857b99c9d8c98ef1f1

Observation 3a1b20da-e3a2-4b93-8472-8b2f65e442fb · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Acme: A Research Framework for Distributed Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.431374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:36892562f51c41e870ac47b79fe8ee8804ff00fb3fd434d75fea6f06d4d3d83c

Observation cb616be2-e8c4-4786-b318-de0a9c8e7843 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.405180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:13ad86e1256c6750337c46acef85128d19b55c5d05d2a086ca9f8b605e2ecd85

Observation b458b75f-a787-4470-b38d-a48f0b94318b · outbound

This paper cites Let's Verify Step by Step.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Let's Verify Step by Step

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.446011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:4a80afc984e252afe009619163f040fdb7096c57f914d25ed8708a22abca5f0b

Observation de2b4abc-161c-4bc5-8aa5-27fba8e177cd · outbound

This paper cites Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-09T08:56:06.419907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:87d57a8528c5d88fff9035c40ae1afb186fdda99588f18a2170c9c48d966642e

Observation fd9748e5-9bdf-404b-b861-64a172b03968 · outbound

This paper cites an unresolved cited work.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-07-09T08:56:06.846916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:8a01c20ca5ab9ddb31a97061e2d425e28a8e5ad6c7b05df2411a37b5bc8cc3b2

Observation 3f38ed09-ba6e-4358-b6de-8abe62e05649 · outbound

This paper cites V olodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning V olodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:56:06.848602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:091ecad452c5d24292308e47a1d7f6f03f0a1eb6a7266190434e1438c22156cb

Observation 868684ee-bdfe-4db4-b959-e46c40df7f67 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.448589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:4cadea706c37c4c3f9f36e914fa9cb4e8b90811e06afa3a8e7fa449048884162

Observation cb2a4a22-1319-451c-813d-68375c26c8ff · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.402413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:1005eb0807b9ccc199913e2aefc1bd4a0259b059e49810f686bbce03eab5e9bb

Observation b6880e6f-ff2c-4e2b-9cc3-b29ec19eb916 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.416635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:460f50541b1993fa6c429640abea131f8aef4b0db4ab0475e08e6ee68de697b9

Observation f0b09ed3-9e6a-430e-bab5-aed6b585c98f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.439895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:7dd1fd122a9536dc3c19e82fc9f1076b26fcd38c18985abcc0cf38fe81cb2a0c

Observation 4c3d12bc-f903-4de7-8813-003e0a82207d · outbound

This paper cites Every step evolves: Scaling reinforcement learning for trillion-scale thinking model.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Every step evolves: Scaling reinforcement learning for trillion-scale thinking model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T08:56:06.443232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:01fbc035b5c73cf50fa5f54fea987ed4f17d0b2d9729a2e21e3600ad8dd71479

Observation 1d802991-aa15-4f98-a56c-7d128d623fa6 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.408245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:cd5d359d2bd6ef79ec3ecdd841b27f656fcc089583707156c7a0554bf3c6100c

Observation 1627bf83-e4d9-4f99-87bf-caf7d7137fad · outbound

This paper cites Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T08:56:06.451773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:82a16ec0eb1bf647f8c453d4574910146e7a93a677275697ec9e17a002604409

Observation eb48c571-9088-411c-987c-19551079569c · outbound

This paper cites Qwen3 Technical Report.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.411146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:be4746aa781ce49a33d7d938a52f6295ab98f21c15c3ed79e223824b4bd8681e

Observation f2edcb82-d888-4486-af58-067824bc5ac7 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.413615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:c5616c014f52624764d48791144a80bc230a288d600deceb916662e0e28ef710

Observation 4e0b0ac7-cdff-448e-a7e1-ef20f930a758 · outbound

This paper cites Group Sequence Policy Optimization.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Group Sequence Policy Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.434205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:a66e725ec045d492154a7ef297be29c512abe893f974d57d969638e4142603f2

Pith citing papers

Observation 884b4557-7dad-476e-b828-55a4aa438cee · inbound

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning cites this paper.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.203143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.203143Z digest=sha256:45e27a44421dda35239fd3aa8ff127badc84be789b94067fc38db0340fb0a408

Observation cc26dc60-50b0-4b78-87d0-2c2eac0d56a0 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:25.886104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:25.886104Z digest=sha256:88abdd5056b1dfe29cc5e62eb617e71e7ed0a7e46479cecbb3cc113e26d6a69b

Observation c3e750cc-d962-4755-b375-d5cd91e43c79 · inbound

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents cites this paper.

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:46.052248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:46.052248Z digest=sha256:417a9b8a075a099f01d6976d2d93563d6a88adc2aa66d7e61ec84020d8ab9d5b

Observation 54bc4f66-0d05-4dd2-92e8-85cace2ffe15 · inbound

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning cites this paper.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:16:34.000797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:16:33.548470Z digest=sha256:6175c12d155e64d8f7035783f09a61c4a7ba9834cce30d0bf8ff1cf765658cb0