Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:37:53.152370Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2606.04272.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:37:53.152370Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:26:33.858637Z
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c50e5123-6655-4a91-909c-42f8823b2885 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement learning on pre-training data.arXiv preprint arXiv:2509.19249
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4383ca90-3c58-411b-a610-ad941016c0a5 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6d57a6-4722-4e50-a417-4552047a6fb5 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement Pre-Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ca0cfcdb-d01f-40f2-93f1-4ab0139559b1 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , volume =
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28dced6-e67a-4efa-a83c-a474f8a32073 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7982491a-0431-440e-a9d2-9b0e517a6f18 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Machine Learning (ICML) , pages =
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5283c9e-cc67-45b0-b767-d4500200055a · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , volume =
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb40cfcb-cb99-4967-875d-99f831075796 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Secrets of
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e9194c-8070-4ce0-8925-91144a9be7aa · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e26056-393f-43ec-b02f-ee0fecee4344 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training On the Interplay of Pre-Training, Mid-Training, and
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed3fbaa-6ff7-406f-a60d-e470c49527ca · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2510.15020 , year=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 121a3219-5509-4dcb-8413-376fa84b50a1 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc25554b-1c10-40a0-8f70-03b0c2821027 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aaedb04-4fc7-4dfb-83eb-95d393d6bef6 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddd8ad2-1cf7-4d2e-858e-a42cbad3a3e9 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2023 , organization =
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c625d6f-91c0-4344-932a-ec98327e595a · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Conference on Language Modeling (COLM) , year =
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dcf0d89-f4f1-4bd9-afbf-aa2501561406 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef484a2d-e6c3-430b-8d7b-36e01b57746a · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3a323d-dcc1-422a-9e8f-690dbd1ff8aa · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Learning Representations (ICLR) , year =
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458c0839-00bf-4f33-bb69-c71c4583681e · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2021 , url =
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f43cd7-81e3-4636-ad1b-e2a264d52ca3 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2020 , url =
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d714f0-b551-408f-b3c5-f57ad6dff0cc · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Advances in Neural Information Processing Systems , pages =
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46eafb37-2a1a-4896-bd5c-f3a838eb87cd · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2019 , eprint =
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da44d80-0f2b-40c8-9865-e1fb13575ba5 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fb571e-44b2-4a4b-ac05-09818c4d67c9 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Evaluating Large Language Models Trained on Code
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 070a267f-bb99-4b75-8997-9cbec284a564 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Training Verifiers to Solve Math Word Problems
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 76901fb7-a0f7-4e7d-b8e5-331645346cce · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Measuring Mathematical Problem Solving with the
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d879b86-e998-44ce-9933-a219600d4a5b · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Does Reinforcement Learning Really Incentivize Reasoning Capacity in
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48876333-6ea8-4b76-9383-9b5c824b7b12 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Training Compute-Optimal Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation df50f29f-2842-4f13-b1c0-130c4a385d85 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Invisible Leash: Why
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d37cba8-1c47-487c-ab06-85da1c859b13 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2023 , url =
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3f247f-3421-45dd-8857-e1b8b5637a01 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105465d2-7129-4f7d-bb0e-caeb01e29c86 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Olmo 3
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5728a67e-fc49-43da-bdbf-369bcfcf5586 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Learning to Reason under Off-Policy Guidance
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 336a2e91-d373-4b2f-9122-4a4757173a7e · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02792363-90c7-4c3d-8015-dfe1491fb7e3 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a645d370-9dcd-4c3a-ab77-486656631ffb · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training International Conference on Machine Learning (ICML) , year =
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540bccad-7215-41fb-853e-e5f01f4aadf3 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Towards a unified view of large language model post-training.arXiv preprint arXiv:2509.04419
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7804ce0a-a8ac-4ce4-8ba0-18495b214c2a · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a052105-5b53-4020-a542-742571cee4f4 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2508.11408 , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b6ae605-f2f8-44b8-91c1-46d0617c85a8 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reasoning with Sampling: Your Base Model is Smarter Than You Think
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0a1b4ef-e6b4-48a5-a829-75a391b07927 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e199826d-5b70-4d5a-ba19-84ea89c70b74 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23aab82-b223-4a18-9989-2398ff665ae0 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Proceedings of the National Academy of Sciences , volume =
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673226cb-ef8f-4a45-8026-8d9c8c439844 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , url =
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4ea56d-bbe6-4ac9-ac4f-f2239273c100 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2024 , url =
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eba6ef1-07e6-40bc-bdbe-0a3752181737 · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , url =
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40eddb5-23b1-42d0-9853-690827aeed2a · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6b48ae6-c1de-452f-9bb0-097b29c260aa · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2025 , eprint=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64043b4a-9574-4f4e-b31e-51a8bbbe24ae · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training 2026 , eprint=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da46f109-766e-4bfe-9bb9-aef13c358ebd · outbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training arXiv preprint arXiv:2412.04619 , year =
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a5d67da-066c-4e9a-9def-35c2a8f57ba4 · inbound
Understanding Reasoning from Pretraining to Post-Training RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.