Pith. sign in

Paper Citation Record · LEDGER

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.18722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18722 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:36:34.322864Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6413e05-596f-4d37-be35-57411ae27a57 · outbound

This paper cites Staleness in fully asynchronous RL.https://appliedcompute.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Staleness in fully asynchronous RL.https://appliedcompute

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:31.831428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:31.831428Z digest=sha256:f909c417836551892ff173a753af9731758ccdb0030f12a2be80a01571fb5a42

Observation b6cb9c57-9e63-4c53-bf68-528e5929e1b8 · outbound

This paper cites Olmo 3.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Olmo 3

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:31.937887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:31.937887Z digest=sha256:32be9917f5f72fe76e67563d642b34fda21ab6a199e62bcee16aadecf11ee17c

Observation b8b5cc90-592b-426c-b6ce-460b27712920 · outbound

This paper cites Kpop: Taming training–inference mismatch in reinforcement learning with adaptive masking regions, May 2026.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Kpop: Taming training–inference mismatch in reinforcement learning with adaptive masking regions, May 2026

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.058621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.058621Z digest=sha256:aaa957d761c4cf589b93c7b4ea6cdda708369b28d04d5b4eea3d62838c6e9264

Observation 884b4557-7dad-476e-b828-55a4aa438cee · outbound

This paper cites Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.203143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.203143Z digest=sha256:440dea7c1a6b76627e19d22234632d5f23700a4d6513ef6b71a3a7942821c728

Observation 64398b4c-f0f4-4f83-89fe-d0e1a2fdda98 · outbound

This paper cites Stabilizing rlvr via token-level gradient diagnosis and layerwise clipping, 2026.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing rlvr via token-level gradient diagnosis and layerwise clipping, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.304147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.304147Z digest=sha256:40f2adb821262b706e439ccf85cd860c574c2adb829a8f08ef00948283190a35

Observation a46f3cb1-3564-4238-98fa-791a20249e55 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Approximately optimal approximate reinforcement learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.449198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.449198Z digest=sha256:56d3a88362fdaf39f6c32aeedbd11fd40d47ef7868423498555fdcff8b701fa2

Observation 2cae68bc-4ccf-4684-b76f-fa09566df595 · outbound

This paper cites Stabilizing MoE reinforcement learning by aligning training and inference routers.arXiv preprint arXiv:2510.11370, 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing MoE reinforcement learning by aligning training and inference routers.arXiv preprint arXiv:2510.11370, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.634733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.634733Z digest=sha256:4830725282721b43535d6692ceea660890b6b1d680cfe9a0b7d1384bb5d5cfc1

Observation 3c811da1-d109-4b48-9278-e72bb60c6bfb · outbound

This paper cites Rethinking the trust region in LLM reinforcement learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the trust region in LLM reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.845639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.845639Z digest=sha256:7b6716b9fc38759f61926449958203c7d5275c42855aadbe1109af72cfef45fe

Observation 738ca189-9408-4045-8f74-1be796585e9f · outbound

This paper cites Trust Region Policy Optimization.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Trust Region Policy Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.958394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.958394Z digest=sha256:506c1be2acc3ea6d54bcbe5c5bb4dfa3a08c9da0375c97cfab93bec136a9fed3

Observation beec5b53-d5e2-4e73-ba9e-d12b235f85f8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.084469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.084469Z digest=sha256:7bec03ef7406913c1cbd38d9d794cd315c9fe3c88d12aadaba98f431bfb1fd3b

Observation 0d9e1a7f-c816-4895-ab79-cb4bbabb0842 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.192530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.192530Z digest=sha256:5d4be40c337f80cfb6b5b2a5cf8475a80c39572f3f3e917ab141a2f26b418136

Observation be9955a2-bf4f-4a54-a983-4bc8373cd7d4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.391991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.391991Z digest=sha256:95a5370467930da24ac80f54f825ffa419264031e17640e587afe86241976b7c

Observation 5d1bd3b8-31c6-4612-af43-5d8203e0db37 · outbound

This paper cites Qwen3 Technical Report.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Qwen3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.544666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.544666Z digest=sha256:aaeb5c7a746439cbcced0d95bfc88ebbb655cae1834ac8de8705883c5e2b2a0d

Observation b288e0a7-fbd5-4981-9ec2-f356ff9bcba6 · outbound

This paper cites Rethinking the Divergence Regularization in LLM RL.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the Divergence Regularization in LLM RL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.642069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.642069Z digest=sha256:31872f5fb8e61d2d74864ab302d23bcf1f70faa89fac3a21cc829b19359500d6

Observation 2da0c69c-9f37-48ac-9d82-5b7cfe2ab2e5 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DAPO: An open-source LLM reinforcement learning system at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.791105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.791105Z digest=sha256:2dbf5a9654f184820562b1f84450ec8d7358f71cf8ce13a8c108aa63735d8545

Observation 1e6148d5-599a-405e-8982-b49caa7c7880 · outbound

This paper cites Small leak can sink a great ship–boost rl training on moe with icepop!, Sep 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Small leak can sink a great ship–boost rl training on moe with icepop!, Sep 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.909133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.909133Z digest=sha256:2366a16048386183dc4f3e52da88cb76363a8cead86cc52040243bbfdfe96a75

Observation 12399fa0-bf49-48f4-9c37-add1d4c02137 · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices, 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing reinforcement learning with llms: Formulation and practices, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.995325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.995325Z digest=sha256:39b9d2756e98ab0c189a57748492d491e968099dc58035d4fa0589dc51fa2342

Observation 21e3219c-4068-4cba-a11e-089608d7fb36 · outbound

This paper cites Group Sequence Policy Optimization.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.082604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.082604Z digest=sha256:6ff1fd8ae91bc9bd8c9769957d0fe3586a1fecee48c76e36d93848e66529bc1f

Observation b8d1cb8a-44e8-445d-8086-f3b1bc2763ca · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.155714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.155714Z digest=sha256:39a869492b67090d221bff36df644ea3cb7f3ddc37238983db5a5935c11a4810

Observation b3d1f6f7-a2c5-44b6-a836-00045af49020 · outbound

This paper cites X yt µ(yt |s t) π(yt |s t) µ(yt |s t) −1 # = 2ξ TX t=1 Est∼µ h 2DTV(µ(· |st)∥π(· |st)) i = 4ξE y∼µ.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning X yt µ(yt |s t) π(yt |s t) µ(yt |s t) −1 # = 2ξ TX t=1 Est∼µ h 2DTV(µ(· |st)∥π(· |st)) i = 4ξE y∼µ

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.238953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.238953Z digest=sha256:7ca5c2ef57a4e7e19a821c3a406df6ac4fc70f5c78933d9b34ac13484e3bdde0

Observation fd5ebb47-0db7-44c1-b13d-b11f3608725d · outbound

This paper cites This is consistent with a sampledDTV gate improving stability on the reported stack, but it does not explain the remaining gap toSAT-GSPO w/ R3 or even toSAT-GRPO w/ R3.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning This is consistent with a sampledDTV gate improving stability on the reported stack, but it does not explain the remaining gap toSAT-GSPO w/ R3 or even toSAT-GRPO w/ R3

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.322864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.322864Z digest=sha256:1fe93347bcaa851317623c99e330cd13d27c2b28ca3bf5da11ca9fa001b25247

Pith citing papers

No inbound Pith citation observations are available.