Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:36:34.322864Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.18722.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:36:34.322864Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e6413e05-596f-4d37-be35-57411ae27a57 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Staleness in fully asynchronous RL.https://appliedcompute
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6cb9c57-9e63-4c53-bf68-528e5929e1b8 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Olmo 3
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b5cc90-592b-426c-b6ce-460b27712920 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Kpop: Taming training–inference mismatch in reinforcement learning with adaptive masking regions, May 2026
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884b4557-7dad-476e-b828-55a4aa438cee · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64398b4c-f0f4-4f83-89fe-d0e1a2fdda98 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing rlvr via token-level gradient diagnosis and layerwise clipping, 2026
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46f3cb1-3564-4238-98fa-791a20249e55 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Approximately optimal approximate reinforcement learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cae68bc-4ccf-4684-b76f-fa09566df595 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing MoE reinforcement learning by aligning training and inference routers.arXiv preprint arXiv:2510.11370, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c811da1-d109-4b48-9278-e72bb60c6bfb · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the trust region in LLM reinforcement learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738ca189-9408-4045-8f74-1be796585e9f · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Trust Region Policy Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beec5b53-d5e2-4e73-ba9e-d12b235f85f8 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d9e1a7f-c816-4895-ab79-cb4bbabb0842 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9955a2-bf4f-4a54-a983-4bc8373cd7d4 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1bd3b8-31c6-4612-af43-5d8203e0db37 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Qwen3 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b288e0a7-fbd5-4981-9ec2-f356ff9bcba6 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the Divergence Regularization in LLM RL
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da0c69c-9f37-48ac-9d82-5b7cfe2ab2e5 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DAPO: An open-source LLM reinforcement learning system at scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6148d5-599a-405e-8982-b49caa7c7880 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Small leak can sink a great ship–boost rl training on moe with icepop!, Sep 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12399fa0-bf49-48f4-9c37-add1d4c02137 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing reinforcement learning with llms: Formulation and practices, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e3219c-4068-4cba-a11e-089608d7fb36 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Group Sequence Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d1cb8a-44e8-445d-8086-f3b1bc2763ca · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d1f6f7-a2c5-44b6-a836-00045af49020 · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning X yt µ(yt |s t) π(yt |s t) µ(yt |s t) −1 # = 2ξ TX t=1 Est∼µ h 2DTV(µ(· |st)∥π(· |st)) i = 4ξE y∼µ
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5ebb47-0db7-44c1-b13d-b11f3608725d · outbound
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning This is consistent with a sampledDTV gate improving stability on the reported stack, but it does not explain the remaining gap toSAT-GSPO w/ R3 or even toSAT-GRPO w/ R3
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.