Pith. sign in

Paper Citation Record · LEDGER

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2605.12380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12380 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:39:04.306845Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact9
  • verified fuzzy27
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c243dda-e7cb-451c-878d-e652851278d1 · outbound

This paper cites Amo-bench: Large language models still struggle in high school math competitions.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Amo-bench: Large language models still struggle in high school math competitions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.146229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:74b99b9ca03ee2ec85c78fb4d01d22d5ed4f6e1981cf2dfddda432668b24e86a

Observation 71356322-f774-4891-b7e6-8a78977a6580 · outbound

This paper cites What matters for on-policy deep actor-critic methods? A large-scale study.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training What matters for on-policy deep actor-critic methods? A large-scale study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.079693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:7478653f9e0d3d5d25b6e0961d4f837d758e724bacd7f0ef2a1bbeae2da91317

Observation 14d801f2-3185-4f2a-a32c-fa250507843a · outbound

This paper cites Bridging the training-inference gap in LLMs by leveraging self-generated tokens.Transactions on Machine Learning Research.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Bridging the training-inference gap in LLMs by leveraging self-generated tokens.Transactions on Machine Learning Research

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.003365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:8f5fb08c77217cd273eb913aae987960d5f7da7b62e886d8e0d68efffc5e6b39

Observation e713ee0d-d843-4a8d-88bc-adb122cdaa83 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.Nature, 645:633–638.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.Nature, 645:633–638

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.083999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:f31ff72b038786419bff88be31a146225c36deda0d20f665b0add5831ff8271c

Observation 1ba607bc-0d2d-4c77-909d-a0e3b21c8406 · outbound

This paper cites Implementation matters in deep policy gradients: A case study on PPO and TRPO.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Implementation matters in deep policy gradients: A case study on PPO and TRPO

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.075355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:22ff746b355a7ba6020c2b762e47b6af5a78eee2dab1a3658432d0146ba678ef

Observation b40fdd02-35e4-42ec-8e59-9fd7e789199a · outbound

This paper cites IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.066735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:06006bc5b9e08d0b57d1179c799b4d5e4b5e0a1995f192fb86e690906276fe62

Observation 354d4686-ca58-4f1e-8a32-0c34540e9f08 · outbound

This paper cites an unresolved cited work.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-13T10:17:39.046404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:c834bce3fc5dad14c1af5f181d25211a8103ec7a0f9fc1d625cebf04965d8d87

Observation 8f999cdf-f5d1-4288-90bf-0f3f34668aa1 · outbound

This paper cites DDPG++: Striving for Simplicity in Continuous-control Off-Policy Reinforcement Learning.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DDPG++: Striving for Simplicity in Continuous-control Off-Policy Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:42:21.015162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:e3d2faab2b2229d15fa37f4198118f05bbe6bb5af245094b69fd360e7e43f731

Observation 2e3b96a8-358f-4460-807b-976e67f738d6 · outbound

This paper cites an unresolved cited work.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-13T10:17:39.054682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:cfbb4cac47d837626ce342400141f92c532eeb02fd595174479000007b7c212d

Observation 61e706a3-94d9-459b-8ae9-df42c4a0f63c · outbound

This paper cites Lipton, Pratik Chaudhari, and Alexander J.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Lipton, Pratik Chaudhari, and Alexander J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.058601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:6602c204fa06703800b02594a0c7daa522e74ce7fcb51645acb1849fcee6fbe3

Observation 49cee8d0-dc64-45e9-afc4-145bdf1dc056 · outbound

This paper cites Continuous doubly constrained batch reinforcement learning.Advances in Neural Information Processing Systems, 34:11260–11273.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Continuous doubly constrained batch reinforcement learning.Advances in Neural Information Processing Systems, 34:11260–11273

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.062522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:14d152bb31a78738665c7bd0164511198f038808d87d94ab4e5c5ed49e30a3d9

Observation f82cbb90-947e-4ca4-8a2f-87a56eeb337a · outbound

This paper cites Deep reinforcement learning that matters.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deep reinforcement learning that matters

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.071370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:f193303e5b8d8483334155759d18941565621310a50444f7e0dd214831696ccd

Observation f49c6e4b-aa79-4c2f-b54f-b06aa19579ff · outbound

This paper cites Batch size-invariance for policy optimization.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Batch size-invariance for policy optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.091633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:89d9dc1b190b00246be3bca2f7cb16c8132f145429b51a484b931cdd127e0dea

Observation acce5128-763f-4f8d-a388-c8df6444d004 · outbound

This paper cites The 37 implementation details of proximal policy optimization.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training The 37 implementation details of proximal policy optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.141074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:4ddd6c4bf635bfcdedf2fcc8d0f2686809159ae638eba132c91ca2ce2c4f28e3

Observation 5ef1f5c4-d947-4f09-90d8-4824bbca8958 · outbound

This paper cites A note on importance sampling using standardized weights.Technical Report 348.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training A note on importance sampling using standardized weights.Technical Report 348

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.017058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:21607913f39d652fcadc91177e324ca599bfd893abb207600409453f8b994627

Observation 4182725d-9c30-4766-b8dc-ffa9773b8dd5 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.047235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:4347e8c0ae0cfa8da39ee87fcd924fd5d4ed7bfc47f0bb082fecec0f30d0b27e

Observation 892412d3-78f1-4e34-9199-afed08f9ee8f · outbound

This paper cites Budgeting counterfactual for offline rl.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Budgeting counterfactual for offline rl

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.021166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:e38f327a558264f419b33265fc0e0b5a651c35baa3040dfe68f80aca09bb1b0f

Observation bd61730b-01e7-4b40-9f0c-17346b1a7ac7 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.033251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:48a6428895b37195070c0d06cb9ada164007938a0373aee21d92e2a848051da2

Observation 2a947907-4c24-4181-99e3-13b8c5a89aa2 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:42:21.043232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:c22cf7b7514f0b4d1be1d0718ca0d170745adae59bc3f2f509bb4b1d975677fd

Observation 6cda9dfa-5239-4d64-86d4-d5c33b221210 · outbound

This paper cites 2023 american mathematics competitions (amc 10 and amc 12).

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training 2023 american mathematics competitions (amc 10 and amc 12)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.037920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:0abbd2025cc4055c4e4e13f05f0a1c1a0186bd69387327512ec48c9e333ee4fa

Observation dd2b72ee-884a-4499-a9a6-18ca007fddd7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Training language models to follow instructions with human feedback

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.051093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:b21c41331627ddb1f3ab6a122ce9a97b2a896c2df18e3aa4286170164a5e89f6

Observation 6eea29ba-3c0e-4598-80c5-2a268297c87a · outbound

This paper cites Defeating the training-inference mismatch via fp16.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Defeating the training-inference mismatch via fp16

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:42:21.009935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:969a4f15fda44c034dbaaa7ceb46fac54a6bd7a7edb94c788a50b02ee1423b59

Observation 099a9314-ec89-4b4f-a306-e420fdf99257 · outbound

This paper cites Springer Science & Business Media.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Springer Science & Business Media

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.008314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:4e5ae4785fde26f181acfbbefd03b725a4732b9ee861da54441ae5fdab8045ae

Observation b8f8f53a-2266-4d34-a9f7-bc115d662516 · outbound

This paper cites Trust region policy optimization.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Trust region policy optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.095829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:2e60c18eff73d821de046cccf7b7761d4a4ada86bfe11ef1b3582bd033728962

Observation 2d472ba1-ecd6-41c2-8d3e-5a5fd4ed2db4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Proximal Policy Optimization Algorithms

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.038720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:9152f21fd17cfceb89d979a0069268fcd646983d1542191f0d87cf9104412be6

Observation 2326cc52-513c-49e3-8af6-e7a2870fd943 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.030358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:0671e116dbdc2a9551c45eedb5e52bfb32c5466f4546a6a32717d6e696ca582a

Observation be00e02d-b485-48a2-844d-3c8bd9006959 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.012612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:a4ab22f49e557754c7a74780eed7f9d187ce1977e4bd8b52a96a1c150393ca79

Observation b495a167-96d8-4f6b-bae2-077b10a8bf82 · outbound

This paper cites Sutton and Andrew G.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Sutton and Andrew G

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.087800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:e618e0b25863d3f41f80cc33d2dd97bc2277b01aa2445aa939916a98b4f2e249

Observation 83357d86-8e76-47b8-8529-4baf4de39174 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Qwen2.5: A party of foundation models, September 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.029351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:868327a662438bada0ece7a8c7b049f17939f39cec87e547dfeb839e7789908d

Observation 37a9a01a-18e1-46b6-a354-933ded269136 · outbound

This paper cites Qwen3 technical report.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Qwen3 technical report

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.118278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:1cda513037e6f8976eb22617e6e78f533ba05e25bb343a0e7f577d29120deb79

Observation ec249c40-fe0a-42c7-b01a-5129ba583d0a · outbound

This paper cites Fp8 quantization — vllm documentation.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Fp8 quantization — vllm documentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.121978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:119156fdd18360d6334f4652a8da0f3f05d5e4e211c3c8f71cec651779fdbf5e

Observation 63f70013-7a03-4d88-85cf-34bce9076dca · outbound

This paper cites Williams.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Williams

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.133369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:072d0195fb91b6d87cbb686775a8870398a12cbd97fb6005445f80638d120383

Observation 892475b6-3b25-4c68-a834-ba959471bee0 · outbound

This paper cites Jet-rl: Enabling on-policy fp8 reinforcement learning with unified training and rollout precision flow.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Jet-rl: Enabling on-policy fp8 reinforcement learning with unified training and rollout precision flow

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.114332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:d79e1d9234a3ff7a5edf64a93444217e6dccf7462d25b52fc6c8ccc994dc8187

Observation c50ed725-5222-48e1-a881-d227efddd534 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.035003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:898cd3950bb60a8c1f5efe743d5c9fc55f69b169258d5222abda08eb8a3b9b81

Observation b9b6840d-02a1-4000-9d8f-5a2c899b8ae3 · outbound

This paper cites American invitational mathematics examination (aime) 2024.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.137116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:c7d4bb66ca3337df1b0161610c00de3c84335b5efa70bb611c397f8cefff8cde

Observation f952669a-9237-4939-9ad7-4ea5e9e98524 · outbound

This paper cites American invitational mathematics examination (aime) 2025.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.103092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:b5e1636d3c5ecbe9a5154b8a436c608efd2b2f7960caab6780c5e9fac5ad6771

Observation d9e4ca4a-4bd6-45c6-a842-346f026050b8 · outbound

This paper cites American invitational mathematics examination (aime) 2026.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training American invitational mathematics examination (aime) 2026

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:17:39.110087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:295679f5aa82813a4d245a56e865592244924242bff973b4c43b424912453ed9

Observation f6945279-67c8-4ce3-abce-d53b6aa41c7a · outbound

This paper cites Group Sequence Policy Optimization.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Group Sequence Policy Optimization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:42:21.025655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:dd4c18f7f436fb285b307a63063848ce47bbd71a2ec8630bbdf97e56a351adef

Observation cb8982e8-5366-41c5-9355-71f1628fd6e5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Fine-Tuning Language Models from Human Preferences

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:42:21.020354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:39:04.306845Z digest=sha256:59f7366751caef356624b7355b3945b63c43e948ee9ecf89551cea3f738dec97

Pith citing papers

No inbound Pith citation observations are available.