Pith. sign in

Paper Citation Record · LEDGER

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

As of 4 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2605.12070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12070 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:57:49.286939Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:56:36.611575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-02T12:16:14.744948Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a5e0c94-92bd-403e-8da9-d4993c90e2a6 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.922725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:fa6eccade909275a3fa97cc9fa5d0e8c7b9ecef44f3fb544b35db3b9bd2e57eb

Observation 86d40475-e663-4cba-b2b3-22c2e2a98b27 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.920908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:3a5f6527617bba330ae9d724c886f4d615969108ff3c224696414fc51c2706b3

Observation 58d05ae4-e8f2-4793-8c7d-7813d40ae441 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.557756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:df28cc7f2878276b259ff5f379c729de3e6e59530991db066ba3e28799496a1f

Observation 243d28b7-4cab-4dcb-aa7b-30b7bef80327 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.573116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:7d886069d2674b8fa3aef9f84707488d963912cd244dbb4228093dd2e5d75bdf

Observation d4f29b51-a9d4-4994-8318-68e2048ae54d · outbound

This paper cites Agentic Reinforced Policy Optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Agentic Reinforced Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:57:12.173630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:aa3904d41b3737eb6a71597cdffe666057e719e80d72e07536798bc426c3178e

Observation d6efb9a3-b5a1-4520-8e78-302bf9fe19b1 · outbound

This paper cites Areal: A large-scale asynchronous reinforcement learning system for language reasoning.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Areal: A large-scale asynchronous reinforcement learning system for language reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.918869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:5a4341347e7c0d6ebccaed29825ff393866bfe8a29827a2d801741f0e4f774d6

Observation 80e444e8-3257-47cd-9d51-1af09203b22a · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.570402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:bfb5b4ca53477987f66836eef942fd6eb0f4e54743b8a17f6492951111636d12

Observation 23f2fcd8-c8fc-4b4e-bc46-400f0d170df4 · outbound

This paper cites Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.564409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:22cd2b860e5a82e65b54e191786ad05b3e5c766bcb33e8081b7a4fddb5afb40c

Observation a77987df-108c-4545-9349-22f3e287bc09 · outbound

This paper cites Batch size-invariance for policy optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Batch size-invariance for policy optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:76b6a375c4d2c4a44762d6d8152e4545a57b85eed345ab043737d91abe863853

Observation 6083ae74-36f4-4064-aea4-5312a80d8551 · outbound

This paper cites Stable asynchrony: Variance-controlled off-policy rl for llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stable asynchrony: Variance-controlled off-policy rl for llms

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.561246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:92e05fb2a2296fcabd9c8d889a1465d1e17b4ace63195671ea2ca672ad203810

Observation b3444ef2-1355-4175-a5ed-9daa6d8c2ab2 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.902836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:a0fc760353cd5bee41df4bdb474f99e669939712c69364b9a34369cec9116118

Observation 9934be26-66fd-4b6b-9e49-ce933009aa73 · outbound

This paper cites A-3po: Accelerating asynchronous llm training with staleness-aware proximal policy approximation.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction A-3po: Accelerating asynchronous llm training with staleness-aware proximal policy approximation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.528218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:26597e47b99f6501dbd91df15a566c11415e28c8ce88cf3d80b8317b267d816d

Observation b829d819-d35e-4b0c-9cd7-5b8de5cde411 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction When speed kills stability: Demystifying RL collapse from the training-inference mismatch

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.900937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:8a1b2264b0867614b8d2739918df003873bb87d6eb0f157b3756007b9e41ec14

Observation 736074fc-6e75-4b52-a98a-b7676343e11d · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.522294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:40711bf13f35ab85c3eaa9395529aefb646c4628a54f3c1d599716fe254c55d8

Observation b2d56bc3-d8b7-4202-9c32-dc064c6545e0 · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Rethinking the Trust Region in LLM Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:05:11.603431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:910299cc3d82639f6c9a8b735a6a5420cbd89bf422437011d0bef84ce24f1458

Observation 80cabeb2-26d8-4292-8949-9e574954a27f · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.512700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:454b3a6cd4c4930159c92a6bf5c89db1808437618e140f37831e792763601fb7

Observation da8b7d41-10d5-40c8-9ec7-8343a3ae2312 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Proximal Policy Optimization Algorithms

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.502320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:091b7b414e0f4ba8bef021b7d7c9d8060a96215053f89a78641b530833e39ee8

Observation c1cfbc40-68d0-45cd-9e44-0d264e6ef316 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.505475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:2a1d82c5b9d1a904e80e98e8a4f731b1fe28100bfbe1d04a40cede340f862b06

Observation a3b1edf2-5521-487b-a66f-f39d305372c9 · outbound

This paper cites VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.516131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:9037c95d84e652307cdce62d831afd42e12d7d1b5abd621e0d12477f9239646f

Observation 22f492b9-3b1c-4b79-992f-c61670db853c · outbound

This paper cites Laminar: A scalable asyn- chronous rl post-training framework.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Laminar: A scalable asyn- chronous rl post-training framework

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.508737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:57e273ba001c97cbe44f00392c8c8763da00c66db4d98957ad5dd60ac228aef7

Observation fd249724-412f-4d96-8d1a-171b34535e21 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Hybridflow: A flexible and efficient rlhf framework

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.915476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:8916dbcc66157faca9df1b6680abc18cddf8707d1eacdf90868201bdbc706b17

Observation 63cb8be7-490c-480a-9b90-15f5bd778d5b · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.549528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:3acc1ea215653f30e5e308c77a0884c3a6cb6a9593a4d2225cbbc9ee5c8461b8

Observation b20ee37d-8e70-47fa-b22a-ab0bdb940594 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.540340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:7dcb8532a589c2972f9418a13532cea322fb828c6b8598bc7e9ac690e16f353d

Observation 4dab0159-f0c5-4a6b-b257-86db61318b35 · outbound

This paper cites Every step evolves: Scaling reinforcement learning for trillion-scale thinking model.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Every step evolves: Scaling reinforcement learning for trillion-scale thinking model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.546490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:e02d04bcc15c536255f55bac8a1964a13005b51c02ca11e1f192f77fc6c7e264

Observation 3665664e-ce1f-4476-8860-bfcb504be7e3 · outbound

This paper cites Ernie 5.0 technical report.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Ernie 5.0 technical report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.543107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:383742713380814066e0f7a0b2ea0aac409b230d46c6ebb69c3440f880d524b9

Observation f3660f3f-d005-46a3-a023-10a31f34b7e2 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.534275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:58586e372ca0aea6cd850b19e52e491e63ddf829d7583b9b2e7d140425991360

Observation 1b87c563-85bb-43ec-9f24-4f665cfaa99a · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.537635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:6d2b0d510113d19821401c89d6d63f0fd2a53812904858e0536b7dc785e2a7a8

Observation b029efc2-fed6-4bf5-823d-cfee5f79ba8a · outbound

This paper cites MiMo-V2-Flash Technical Report.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiMo-V2-Flash Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.524895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:1ddc9dca94d053521693577f099e9441e08fd8521f8a1e814c28076ca32f29ab

Observation 07b90231-bb85-44e8-94df-75d358c4b7cb · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, August 2025.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, August 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.910212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:962da440ee5b08efadb6da1f1e4260d1bb623f1e7abc52110316e1e55ae54c60

Observation 32303e9a-21cb-447b-b9b0-b2788502cca5 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.917217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:f859911416abadf59545b6d26ea4f825b253a8ecb96a99ccf1f30a3c3b582e9a

Observation 26ad966d-6e42-4bc6-b296-acd744d36442 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.530861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:f5a827f011a28e0456bf2dd44b2bccab123f58b903e547f75c91fdd9db6f20ae

Observation 141c3492-816f-4804-b0a7-14d2c7e4c070 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction GLM-5: from Vibe Coding to Agentic Engineering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.554770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:b924dbf6264d17845d67dfc6c45c125966dfe0c788c9829248cf03e186a94131

Observation 99f9e4b1-c3e1-4ede-a81c-539796e9fa94 · outbound

This paper cites The landscape of agentic reinforcement learning for llms: A survey.Transactions on Machine Learning Research.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction The landscape of agentic reinforcement learning for llms: A survey.Transactions on Machine Learning Research

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.912014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:c38c824c51543943ad8dd317b88c1b3e14c93f1df7cd4c77279f1c2b6da1bee7

Observation 8b232ab9-5f0a-47da-a982-e7ab8a5792f7 · outbound

This paper cites Small leak can sink a great ship–boost rl training on moe with icepop!.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Small leak can sink a great ship–boost rl training on moe with icepop!

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.908495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:e36906028f0739708f4b4caa1f7ea5ba2e557822e823681f4208ed40be2595e2

Observation c7831df4-df25-4f4c-a67b-36852ae4c24b · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.496418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:b73f83d0703b23f4fb37cb66ce49d6981ae1d37fc0bd45ab1a468e4d8f173805

Observation 9238bbf6-a52f-46fe-991e-3fdacb4541c9 · outbound

This paper cites Group Sequence Policy Optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Group Sequence Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.552263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:d5d6b51e5c11ec36fd2930bc03d20a66b28d2a5c69ca7e71bf2796f20acbca27

Observation 806f8306-3f5d-4164-bd7e-59b441558a1c · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.499553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:a26aef91924f00e2d59f031cb6b149a8630941bd8da3adb712fe920c0694b161

Observation b9b6f318-4fad-404f-ac68-15a895208590 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.906563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:c1aea90638dadf198990b2af7b166ef3a7fe215db22e2c0c19e471a8370f8e4b

Observation 5a4380e2-768b-4371-8664-b6fae4aaf663 · outbound

This paper cites Limitations.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Limitations

Reference 39

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T09:42:33.904611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:436a2580c2dbc8c23a6844423dfc02d7a84f3ad8b75888ab2bf5d06825ed04cd

Observation 66624dd4-4111-4234-8af2-517fe92cdfea · outbound

This paper cites Therefore IRB approval is not applicable.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Therefore IRB approval is not applicable

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.913667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:de33c07c10c5ab93f873e2ece9235cb0f55267d5900fca195c5056e6bb64cb14

Pith citing papers

Observation f891b82f-e12e-4294-8fd7-2d1871075466 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-02T08:58:21.176076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-08-02T08:56:36.611575Z digest=sha256:2f54ba1cd3295f087e5463e188312a0804acd07278992f0f0f9d2479de0d5c2d