Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:26:08.536798Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2507.08267.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:26:08.536798Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:39.658589Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T22:43:37.868176Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b57bf2b3-6c37-4caf-9e90-3603c8f86b2f · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f6567c-e674-4803-a3b2-9a077cd0b661 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ed019e-744e-403f-88e5-6f7c2f51ed3d · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe5eb46-0af6-40a4-9cdd-03814bc37907 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Alphamath almost zero: Process supervision without process
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e909581-2699-444c-a2d5-827300282a19 · outbound
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37aca5f0-c980-4c84-ae41-f2358fc22c83 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580f2210-02d4-4d7c-aa92-f4e60816498c · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6258d02-8f85-4857-8d2d-d5583d8f9fdc · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C., Buzzard, K., Gowers, T., Liu, P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6dfa5ca-d1dd-4f0a-88ba-63076368c673 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Measuring mathematical problem solving with the MATH dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71208300-83d6-435a-8ee8-eb585246e9b9 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Training Compute-Optimal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf86600-963a-49d0-b315-4e4c9a5dd9a4 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34d686bd-584d-427a-a481-721473558fc3 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Scaling Laws for Neural Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a5c783-791f-435b-8a72-10a9db1142a4 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Solving quantitative reasoning problems with language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86cb55e2-32e5-41a4-917c-00cdee1a533a · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competition-level code generation with alphacode
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67c5727-2112-4cbd-8b15-c01819965be2 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6a561d6-ea7c-40c9-9b64-6c5474e58b1c · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34e44a7-7b90-4f77-b857-fadfa7d5e00a · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Wider or deeper? scaling llm inference-time compute with adaptive branching tree search
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540f1a2e-80f3-4915-ab92-4d4b1f879a47 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning s1: Simple test-time scaling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a50dcea-77aa-48c6-b002-3b3d8b4a68ee · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b1e904-03b7-4c56-81ca-614ffd5da247 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning OpenAI o1 System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd696b0-af35-491a-bd05-22225f5cf6a0 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competitive Programming with Large Reasoning Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02302a27-eca4-4170-b3c1-f439d6e8558b · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8615f2d8-2c23-4121-9b1b-6851900bd419 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e81936-ede2-4c91-8adf-4e591b2ceaf4 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa33f8c4-dc5e-4f4f-b329-bcc927e8dd22 · outbound
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df7098c-9852-4f97-8a62-9c8852132e61 · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21d4e395-d743-4d01-aa42-52e48ab6cfea · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning T., Wang, W., and Li, W
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0d3622-a270-4a90-b1b8-f81966e32d0a · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning LIMO: Less is More for Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb22ee4b-3048-4dc8-8a0d-33c50d00818a · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Demystifying long chain-of-thought reasoning in LLM s
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 432f3e6a-5dd0-4835-afcf-14f230158f4b · outbound
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5cb283-56ab-47d4-a8ee-94f90b6acb05 · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73d7339c-e9ef-4e58-b417-a13640e62122 · inbound
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0d6a44c-6a40-42d8-bd33-517004870c9c · inbound
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4757bf6b-42e1-436c-812b-69a03b3cd5cf · inbound
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7e02888-5225-4b44-9d33-6f18a8d8d872 · inbound
CRAFT: Learn the Schema, Execute the Plan A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b467e5-494d-45ae-ab3b-100edb5b663d · inbound
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.