Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:57:49.286939Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2605.12070.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:57:49.286939Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:56:36.611575Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-02T12:16:14.744948Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a5e0c94-92bd-403e-8da9-d4993c90e2a6 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 86d40475-e663-4cba-b2b3-22c2e2a98b27 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 58d05ae4-e8f2-4793-8c7d-7813d40ae441 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 243d28b7-4cab-4dcb-aa7b-30b7bef80327 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d4f29b51-a9d4-4994-8318-68e2048ae54d · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Agentic Reinforced Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d6efb9a3-b5a1-4520-8e78-302bf9fe19b1 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Areal: A large-scale asynchronous reinforcement learning system for language reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 80e444e8-3257-47cd-9d51-1af09203b22a · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 23f2fcd8-c8fc-4b4e-bc46-400f0d170df4 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a77987df-108c-4545-9349-22f3e287bc09 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Batch size-invariance for policy optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6083ae74-36f4-4064-aea4-5312a80d8551 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stable asynchrony: Variance-controlled off-policy rl for llms
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b3444ef2-1355-4175-a5ed-9daa6d8c2ab2 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Efficient memory management for large language model serving with pagedattention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9934be26-66fd-4b6b-9e49-ce933009aa73 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction A-3po: Accelerating asynchronous llm training with staleness-aware proximal policy approximation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b829d819-d35e-4b0c-9cd7-5b8de5cde411 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction When speed kills stability: Demystifying RL collapse from the training-inference mismatch
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 736074fc-6e75-4b52-a98a-b7676343e11d · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b2d56bc3-d8b7-4202-9c32-dc064c6545e0 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Rethinking the Trust Region in LLM Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 80cabeb2-26d8-4292-8949-9e574954a27f · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation da8b7d41-10d5-40c8-9ec7-8343a3ae2312 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Proximal Policy Optimization Algorithms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c1cfbc40-68d0-45cd-9e44-0d264e6ef316 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a3b1edf2-5521-487b-a66f-f39d305372c9 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 22f492b9-3b1c-4b79-992f-c61670db853c · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Laminar: A scalable asyn- chronous rl post-training framework
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd249724-412f-4d96-8d1a-171b34535e21 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Hybridflow: A flexible and efficient rlhf framework
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 63cb8be7-490c-480a-9b90-15f5bd778d5b · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b20ee37d-8e70-47fa-b22a-ab0bdb940594 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Kimi K2.5: Visual Agentic Intelligence
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4dab0159-f0c5-4a6b-b257-86db61318b35 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Every step evolves: Scaling reinforcement learning for trillion-scale thinking model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3665664e-ce1f-4476-8860-bfcb504be7e3 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Ernie 5.0 technical report
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f3660f3f-d005-46a3-a023-10a31f34b7e2 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1b87c563-85bb-43ec-9f24-4f665cfaa99a · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b029efc2-fed6-4bf5-823d-cfee5f79ba8a · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiMo-V2-Flash Technical Report
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 07b90231-bb85-44e8-94df-75d358c4b7cb · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, August 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 32303e9a-21cb-447b-b9b0-b2788502cca5 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 26ad966d-6e42-4bc6-b296-acd744d36442 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 141c3492-816f-4804-b0a7-14d2c7e4c070 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction GLM-5: from Vibe Coding to Agentic Engineering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 99f9e4b1-c3e1-4ede-a81c-539796e9fa94 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction The landscape of agentic reinforcement learning for llms: A survey.Transactions on Machine Learning Research
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8b232ab9-5f0a-47da-a982-e7ab8a5792f7 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Small leak can sink a great ship–boost rl training on moe with icepop!
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c7831df4-df25-4f4c-a67b-36852ae4c24b · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9238bbf6-a52f-46fe-991e-3fdacb4541c9 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Group Sequence Policy Optimization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 806f8306-3f5d-4164-bd7e-59b441558a1c · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Prosperity before collapse: How far can off-policy rl reach with stale data on llms?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b9b6f318-4fad-404f-ac68-15a895208590 · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5a4380e2-768b-4371-8664-b6fae4aaf663 · outbound
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 66624dd4-4111-4234-8af2-517fe92cdfea · outbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Therefore IRB approval is not applicable
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f891b82f-e12e-4294-8fd7-2d1871075466 · inbound
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.