Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:39:04.306845Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2605.12380.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:39:04.306845Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4c243dda-e7cb-451c-878d-e652851278d1 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Amo-bench: Large language models still struggle in high school math competitions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71356322-f774-4891-b7e6-8a78977a6580 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training What matters for on-policy deep actor-critic methods? A large-scale study
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14d801f2-3185-4f2a-a32c-fa250507843a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Bridging the training-inference gap in LLMs by leveraging self-generated tokens.Transactions on Machine Learning Research
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e713ee0d-d843-4a8d-88bc-adb122cdaa83 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.Nature, 645:633–638
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ba607bc-0d2d-4c77-909d-a0e3b21c8406 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Implementation matters in deep policy gradients: A case study on PPO and TRPO
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b40fdd02-35e4-42ec-8e59-9fd7e789199a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 354d4686-ca58-4f1e-8a32-0c34540e9f08 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f999cdf-f5d1-4288-90bf-0f3f34668aa1 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DDPG++: Striving for Simplicity in Continuous-control Off-Policy Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e3b96a8-358f-4460-807b-976e67f738d6 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61e706a3-94d9-459b-8ae9-df42c4a0f63c · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Lipton, Pratik Chaudhari, and Alexander J
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49cee8d0-dc64-45e9-afc4-145bdf1dc056 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Continuous doubly constrained batch reinforcement learning.Advances in Neural Information Processing Systems, 34:11260–11273
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f82cbb90-947e-4ca4-8a2f-87a56eeb337a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deep reinforcement learning that matters
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f49c6e4b-aa79-4c2f-b54f-b06aa19579ff · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Batch size-invariance for policy optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acce5128-763f-4f8d-a388-c8df6444d004 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training The 37 implementation details of proximal policy optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ef1f5c4-d947-4f09-90d8-4824bbca8958 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training A note on importance sampling using standardized weights.Technical Report 348
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4182725d-9c30-4766-b8dc-ffa9773b8dd5 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 892412d3-78f1-4e34-9199-afed08f9ee8f · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Budgeting counterfactual for offline rl
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd61730b-01e7-4b40-9f0c-17346b1a7ac7 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a947907-4c24-4181-99e3-13b8c5a89aa2 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cda9dfa-5239-4d64-86d4-d5c33b221210 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training 2023 american mathematics competitions (amc 10 and amc 12)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd2b72ee-884a-4499-a9a6-18ca007fddd7 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Training language models to follow instructions with human feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6eea29ba-3c0e-4598-80c5-2a268297c87a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Defeating the training-inference mismatch via fp16
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 099a9314-ec89-4b4f-a306-e420fdf99257 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Springer Science & Business Media
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8f8f53a-2266-4d34-a9f7-bc115d662516 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Trust region policy optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d472ba1-ecd6-41c2-8d3e-5a5fd4ed2db4 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2326cc52-513c-49e3-8af6-e7a2870fd943 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be00e02d-b485-48a2-844d-3c8bd9006959 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b495a167-96d8-4f6b-bae2-077b10a8bf82 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Sutton and Andrew G
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83357d86-8e76-47b8-8529-4baf4de39174 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Qwen2.5: A party of foundation models, September 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37a9a01a-18e1-46b6-a354-933ded269136 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Qwen3 technical report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec249c40-fe0a-42c7-b01a-5129ba583d0a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Fp8 quantization — vllm documentation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63f70013-7a03-4d88-85cf-34bce9076dca · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Williams
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 892475b6-3b25-4c68-a834-ba959471bee0 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Jet-rl: Enabling on-policy fp8 reinforcement learning with unified training and rollout precision flow
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c50ed725-5222-48e1-a881-d227efddd534 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9b6840d-02a1-4000-9d8f-5a2c899b8ae3 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f952669a-9237-4939-9ad7-4ea5e9e98524 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9e4ca4a-4bd6-45c6-a842-346f026050b8 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2026
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6945279-67c8-4ce3-abce-d53b6aa41c7a · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Group Sequence Policy Optimization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb8982e8-5366-41c5-9355-71f1628fd6e5 · outbound
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Fine-Tuning Language Models from Human Preferences
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.