Pith. sign in

Paper Citation Record · LEDGER

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2606.04272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04272 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:37:53.152370Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:26:33.858637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c50e5123-6655-4a91-909c-42f8823b2885 · outbound

This paper cites Reinforcement learning on pre-training data.arXiv preprint arXiv:2509.19249.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement learning on pre-training data.arXiv preprint arXiv:2509.19249

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.028151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:dbcc40743ce5c9a84b3827e5252289f725b95b28afd45b82aaaf93ef2a01a8be

Observation 4383ca90-3c58-411b-a610-ad941016c0a5 · outbound

This paper cites 2026 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:c63b128108c643dcb44b8407b55e6ee5256c1e6335758e4a071f102be1ac2297

Observation ba6d57a6-4722-4e50-a417-4552047a6fb5 · outbound

This paper cites Reinforcement Pre-Training.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement Pre-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.032828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:03e9050dd38592255d6a11b8e0cef4d3fa727849d796c3d372a9cd87ab1224e7

Observation ca0cfcdb-d01f-40f2-93f1-4ab0139559b1 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , volume =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:ff0b8e74186ced422b50c4f6bf802143166dcd5f7d1033d59e90b65742250f76

Observation a28dced6-e67a-4efa-a83c-a474f8a32073 · outbound

This paper cites an unresolved cited work.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:05426e6ae11789c0ca307dda7d646cc67b0710a287f27ddb2986dc91f606deef

Observation 7982491a-0431-440e-a9d2-9b0e517a6f18 · outbound

This paper cites International Conference on Machine Learning (ICML) , pages =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Machine Learning (ICML) , pages =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:7bf5953e27e868bca6d3bb18d9eda0a069da7b301cb9afe3bcbfe4bd656c566c

Observation a5283c9e-cc67-45b0-b767-d4500200055a · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:6c6d2f64cba07c9b3c4dcc6ab284edbf13418407ffeeed9c76e53cc71a118989

Observation cb40cfcb-cb99-4967-875d-99f831075796 · outbound

This paper cites Secrets of.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Secrets of

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:14c90ce553589d9c7f9830f383c6bd7794afc5c647c67bc14f2d8f873d171da5

Observation a6e9194c-8070-4ce0-8925-91144a9be7aa · outbound

This paper cites 2024 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:13fd4c1ad0e1777f2dcc7a2934df075498a8d40397e9102f5dd5de9955145b74

Observation 20e26056-393f-43ec-b02f-ee0fecee4344 · outbound

This paper cites On the Interplay of Pre-Training, Mid-Training, and.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training On the Interplay of Pre-Training, Mid-Training, and

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:32d5a146815a2233587a0d6a6a56ef27db687cea92e3793053cc69388efd891f

Observation aed3fbaa-6ff7-406f-a60d-e470c49527ca · outbound

This paper cites arXiv preprint arXiv:2510.15020 , year=.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2510.15020 , year=

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.010483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:2f58c2ae9143f393b89db988c286aca76ed25a546b23509a0f0f55c21f67b877

Observation 121a3219-5509-4dcb-8413-376fa84b50a1 · outbound

This paper cites Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.015086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:8f78773643a3547b9defb8890037b8c5e177a93a7f6f39795337f05865ba3e64

Observation bc25554b-1c10-40a0-8f70-03b0c2821027 · outbound

This paper cites 2026 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:b81b9be8748d0a0b2692ccd33384c665cc9062459cc7711cd15680aa7248dca5

Observation 8aaedb04-4fc7-4dfb-83eb-95d393d6bef6 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:8307f1ead6ff5495431c7f860e4ac58160c2c864afc512749c684007fbe789e6

Observation dddd8ad2-1cf7-4d2e-858e-a42cbad3a3e9 · outbound

This paper cites 2023 , organization =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2023 , organization =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:194146a80912caef7d9be9c8fcff410abc4c2431929cfef410e126422616a2f9

Observation 9c625d6f-91c0-4344-932a-ec98327e595a · outbound

This paper cites Conference on Language Modeling (COLM) , year =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Conference on Language Modeling (COLM) , year =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:c9c345a2e89a80128a94b5ca7713540cbd692bbe5d8f9fbf7d4b4c82e2bfd9d6

Observation 8dcf0d89-f4f1-4bd9-afbf-aa2501561406 · outbound

This paper cites 2024 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:b1f4ece25d2a46697e6057726223a1a82b8ce4484eae6c306346ce5d4f5efb82

Observation ef484a2d-e6c3-430b-8d7b-36e01b57746a · outbound

This paper cites 2024 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:c1c29fffe4e8566d26bc1e0993966a642cd1fea04f13bd9d18cd5c15d0fa5cf5

Observation 8c3a323d-dcc1-422a-9e8f-690dbd1ff8aa · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Learning Representations (ICLR) , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:a1c5063aa89192bddf37c1a4720904c677737c6064c23f194c45e023116ae136

Observation 458c0839-00bf-4f33-bb69-c71c4583681e · outbound

This paper cites 2021 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2021 , url =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:dcd191b291cc459c1cc074d0b71693e682c80b713ecb0daed27079a6ac3aa927

Observation 44f43cd7-81e3-4636-ad1b-e2a264d52ca3 · outbound

This paper cites 2020 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2020 , url =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:2329abc3be02ba54f25d6877f753547b03ce4d00797431306c53d04381925e98

Observation 74d714f0-b551-408f-b3c5-f57ad6dff0cc · outbound

This paper cites Advances in Neural Information Processing Systems , pages =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , pages =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:f542084da6ff6df77a1e5eb353dd1a589c9099a20c0d8b616210281e2a80e626

Observation 46eafb37-2a1a-4896-bd5c-f3a838eb87cd · outbound

This paper cites 2019 , eprint =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2019 , eprint =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:22f7a975cad3b0d139804e313e4f2f3f774350d6e6ebfa878d711cf45e9b1042

Observation 2da44d80-0f2b-40c8-9865-e1fb13575ba5 · outbound

This paper cites 2024 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:68aa02d9c9218aeeb82971a51039060add5708883d4c222a1290eac78e55e2b8

Observation 08fb571e-44b2-4a4b-ac05-09818c4d67c9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Evaluating Large Language Models Trained on Code

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:29.005163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:cc5bc9bedcc146b8ee5daba1f03e161a3d2b86557a9e65fee7f29259567968fb

Observation 070a267f-bb99-4b75-8997-9cbec284a564 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Training Verifiers to Solve Math Word Problems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:29.000903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:bb18fe8e10fda74a00f3daa30ed8160031ed0619540e88cdc26db1e7053dfe2f

Observation 76901fb7-a0f7-4e7d-b8e5-331645346cce · outbound

This paper cites Measuring Mathematical Problem Solving with the.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Measuring Mathematical Problem Solving with the

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:ef7776ebf684e6c41b392df745ab3753fbd43122386cbd68ebfcc127ba0354ae

Observation 5d879b86-e998-44ce-9933-a219600d4a5b · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:ecadc88a7eb8e77075de8dede6aa884de1acb822c997d6c31175bd433a9bbd90

Observation 48876333-6ea8-4b76-9383-9b5c824b7b12 · outbound

This paper cites Training Compute-Optimal Large Language Models.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Training Compute-Optimal Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:29.019630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:9b8d51f1546f9585610eb553ab516080ed62a1eb865c85eee5f3aa84a7a8cd5b

Observation df50f29f-2842-4f13-b1c0-130c4a385d85 · outbound

This paper cites The Invisible Leash: Why.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Invisible Leash: Why

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:2f3156b935763ad6aba63b7550551fcde82eb8800703e46883fffac9b3dba9f1

Observation 8d37cba8-1c47-487c-ab06-85da1c859b13 · outbound

This paper cites 2023 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2023 , url =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:23d11b14f9cf7cc6a2de722ea40ab4157bdb1f83202d05513d55fd2e18e53218

Observation 9a3f247f-3421-45dd-8857-e1b8b5637a01 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:f7fc4db2a96b2bb9dfbecd06a0d560be1a1b218eb674da1d6ab400e87caa329d

Observation 105465d2-7129-4f7d-bb0e-caeb01e29c86 · outbound

This paper cites Olmo 3.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Olmo 3

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:29.023535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:a68600c5a75fa478b99b390d6e40a9d9fc4a66231cb2cce39d1a233d6d4c0e6d

Observation 5728a67e-fc49-43da-bdbf-369bcfcf5586 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Learning to Reason under Off-Policy Guidance

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:28.991589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:8ac4b3676ef52d265ef2fe94eb7049a80b1b09ed0b702e487a6f8dabc253d5dc

Observation 336a2e91-d373-4b2f-9122-4a4757173a7e · outbound

This paper cites 2026 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:9ec7737ddcf6471e69bc3373a7156d4e34827ccef1ebc1b591022b1296b89c3b

Observation 02792363-90c7-4c3d-8015-dfe1491fb7e3 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:6f8bbe846b717e043cf0959b9b58b68cb8445a1f7cb7c67fb250a4f11be305cc

Observation a645d370-9dcd-4c3a-ab77-486656631ffb · outbound

This paper cites International Conference on Machine Learning (ICML) , year =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Machine Learning (ICML) , year =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:716aadb83b04d721a552f9938dc80431527ea7c95f3b3a6544fffb9219c12ea5

Observation 540bccad-7215-41fb-853e-e5f01f4aadf3 · outbound

This paper cites Towards a unified view of large language model post-training.arXiv preprint arXiv:2509.04419.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Towards a unified view of large language model post-training.arXiv preprint arXiv:2509.04419

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.998009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:64831f2b60941048ae8ed878e1e1cb20b3565faaeb9b7b0302b29d0fb5860f62

Observation 7804ce0a-a8ac-4ce4-8ba0-18495b214c2a · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:14d386d082c782bb1dbe6a6c1ba1b69a82ca7069bcd54dc20a92f10043982fd0

Observation 0a052105-5b53-4020-a542-742571cee4f4 · outbound

This paper cites arXiv preprint arXiv:2508.11408 , year=.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2508.11408 , year=

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.984570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:a9bc272d6d3e19b0fe98459ae4bc33d31531c43915e63227690a0937fa346e98

Observation 0b6ae605-f2f8-44b8-91c1-46d0617c85a8 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:28.988369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:b5f8dd15802be651acd72f530ebd16e15ac3c7932af0dbec40f917e5d469c57d

Observation a0a1b4ef-e6b4-48a5-a829-75a391b07927 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:7ef80c2e37893e13fb90d06a160ed2a8deb3eeba222b7ad9f09bee2c2d28ad7a

Observation e199826d-5b70-4d5a-ba19-84ea89c70b74 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:c3ee5bbdf8fd2317d0a4e2e6f58554bb79e3f553983dab705ea79cfc1bb6e07e

Observation b23aab82-b223-4a18-9989-2398ff665ae0 · outbound

This paper cites Proceedings of the National Academy of Sciences , volume =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Proceedings of the National Academy of Sciences , volume =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:82e0b6482111d650f098408b322ed6d3168c492c8bebb2ac706eb5cf9919e0e3

Observation 673226cb-ef8f-4a45-8026-8d9c8c439844 · outbound

This paper cites 2026 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:572540abdbaa1b1293fa13e328ce061fc3e606d6a45a7f7d8790792179309b26

Observation 1c4ea56d-bbe6-4ac9-ac4f-f2239273c100 · outbound

This paper cites 2024 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:d0d3505624927fd6917c9d1839d216a0ae82cd630d238b22cde10e1b9383790f

Observation 9eba6ef1-07e6-40bc-bdbe-0a3752181737 · outbound

This paper cites 2025 , url =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:2fbaeb214de861beb7694da16095c9275790db096ef33367072842e22fc151e4

Observation d40eddb5-23b1-42d0-9853-690827aeed2a · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:28.994858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:34b05e20e236aaf2a9c6d18bd8a38f4add2940cbcaeefc94134a46deb95bcc2e

Observation e6b48ae6-c1de-452f-9bb0-097b29c260aa · outbound

This paper cites 2025 , eprint=.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:bea5dd001831654db0bddac739f0f6ecbeb8adb3c444c3701a6cda1fcef6f7e5

Observation 64043b4a-9574-4f4e-b31e-51a8bbbe24ae · outbound

This paper cites 2026 , eprint=.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T10:37:53.152370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:2b9301f8b1bb9ccb1ba731adad1ae387fa78f782cad0aec54030519595fbd274

Observation da46f109-766e-4bfe-9bb9-aef13c358ebd · outbound

This paper cites arXiv preprint arXiv:2412.04619 , year =.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2412.04619 , year =

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:28.980781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:e98ea94b8d6823c6a78368d4890d432661d0d48d29f4e39c3ae4f5c2c7a1c5d6

Pith citing papers

Observation 9a5d67da-066c-4e9a-9def-35c2a8f57ba4 · inbound

Understanding Reasoning from Pretraining to Post-Training cites this paper.

Understanding Reasoning from Pretraining to Post-Training RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:33.858637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:26:33.858637Z digest=sha256:54a9f4d50ffbf348f94c00b1c009de40278414f34688ba8372d1f008464915b4