Pith. sign in

Paper Citation Record · LEDGER

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 6 inbound Pith citation observations for arXiv:2507.08267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08267 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:26:08.536798Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:18:39.658589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T22:43:37.868176Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b57bf2b3-6c37-4caf-9e90-3603c8f86b2f · outbound

This paper cites write newline.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.461183Z digest=sha256:8fc11e4bac15a2f5d79bffc3eaa72925ee5b1ba045eff138e766d818fc5444d5

Observation 64f6567c-e674-4803-a3b2-9a077cd0b661 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.464611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.464611Z digest=sha256:4022a4c431eaf02891aae9a4e387f67f52e47881d0892bfa720bbf9c2c7e12d3

Observation 92ed019e-744e-403f-88e5-6f7c2f51ed3d · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.467197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.467197Z digest=sha256:c0dece962112cd43ddb8434a19b9e8301fc7aae489df8435a7c922d0bceda3e5

Observation 3fe5eb46-0af6-40a4-9cdd-03814bc37907 · outbound

This paper cites Alphamath almost zero: Process supervision without process.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Alphamath almost zero: Process supervision without process

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.949427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.470895Z digest=sha256:e92e300bfaa1aab1ab12a71bb7edee0dc69fd1cc27484dd34151f11b2be8c9be

Observation 7e909581-2699-444c-a2d5-827300282a19 · outbound

This paper cites and Ngo, C.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning and Ngo, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.473351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.473351Z digest=sha256:887e30e46799f0d40a556a3ceecb9214c2035a62225887ed6d7c234828319b54

Observation 37aca5f0-c980-4c84-ae41-f2358fc22c83 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.475643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.475643Z digest=sha256:f474f008d6857744d6f63b40ff4951646ddc33265db98c36688e70cb10d13484

Observation 580f2210-02d4-4d7c-aa92-f4e60816498c · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.477985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.477985Z digest=sha256:511584e0d7f3b82aa377f5e39073134b38a23eaa5c9897038f17800d63ad694f

Observation f6258d02-8f85-4857-8d2d-d5583d8f9fdc · outbound

This paper cites C., Buzzard, K., Gowers, T., Liu, P.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C., Buzzard, K., Gowers, T., Liu, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.938412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.480618Z digest=sha256:4d3670914447ba87fd74573b9963e58b25cb55cf80147605db5214d38e1c64e0

Observation f6dfa5ca-d1dd-4f0a-88ba-63076368c673 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.483094Z digest=sha256:0ec401aed7a95ea5eba78cb603489c7c4499cce26e945fe38b2a770b8fb2570e

Observation 71208300-83d6-435a-8ee8-eb585246e9b9 · outbound

This paper cites Training Compute-Optimal Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.485151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.485151Z digest=sha256:1b74dc38bf059e504d99b451f3b3de92a4007b906f693370e44bdf0f476b1b91

Observation 3bf86600-963a-49d0-b315-4e4c9a5dd9a4 · outbound

This paper cites C3ot: Generating shorter chain-of-thought without compromising effectiveness.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.924347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.488176Z digest=sha256:f0ded786f9dec9fae526b4ac73147286c1ff7941142d7b86a61957da1de5087a

Observation 34d686bd-584d-427a-a481-721473558fc3 · outbound

This paper cites Scaling Laws for Neural Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.490527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.490527Z digest=sha256:13cd1d0782b120836497662aadbf6fa56faf7d77fc45f23001850af5722de130

Observation 97a5c783-791f-435b-8a72-10a9db1142a4 · outbound

This paper cites Solving quantitative reasoning problems with language models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Solving quantitative reasoning problems with language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.917123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.492780Z digest=sha256:a459a848ed011d0cd5ba9ed9cde54e7130d7b42e66be8ad780409bd7894c3b63

Observation 86cb55e2-32e5-41a4-917c-00cdee1a533a · outbound

This paper cites Competition-level code generation with alphacode.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competition-level code generation with alphacode

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.494732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.494732Z digest=sha256:3a91a09505f220dd424652880d14600238ce7ca88d65ef0e7f1a44d3bceddc6e

Observation d67c5727-2112-4cbd-8b15-c01819965be2 · outbound

This paper cites Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.906089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.496800Z digest=sha256:aae96e2006a6fd6dea0af5c3f5d537089479ada0a675d9580dfd9e378013e121

Observation b6a561d6-ea7c-40c9-9b64-6c5474e58b1c · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.499132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.499132Z digest=sha256:8d9b9c1a5252cec74a9d6deab3e8fbe781b5b39b8a8408899d32c76e78792e46

Observation b34e44a7-7b90-4f77-b857-fadfa7d5e00a · outbound

This paper cites Wider or deeper? scaling llm inference-time compute with adaptive branching tree search.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Wider or deeper? scaling llm inference-time compute with adaptive branching tree search

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.501675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.501675Z digest=sha256:2c2a1f44cff0777fa7284ac82384fbe3294000e92b8078831f15b27a0650c319

Observation 540f1a2e-80f3-4915-ab92-4d4b1f879a47 · outbound

This paper cites s1: Simple test-time scaling.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.503965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.503965Z digest=sha256:0987eaa609901a9939bf2da816b8f89bbad34307649f2f18d1c9cf7f1267f994

Observation 2a50dcea-77aa-48c6-b002-3b3d8b4a68ee · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.506610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.506610Z digest=sha256:469312d2b4701b27d9813c404a33edb85a97bb9cb6d6d7ef6c9ebb0f6b1b49fd

Observation 58b1e904-03b7-4c56-81ca-614ffd5da247 · outbound

This paper cites OpenAI o1 System Card.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.509681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.509681Z digest=sha256:a4de86cfaf84ab3391ff28c4a4e7b463caabcb762b99df4ca9d1c7c792cf293f

Observation 8bd696b0-af35-491a-bd05-22225f5cf6a0 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.512634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.512634Z digest=sha256:051cb6655268e6d8e2bf664a0af45190d57c1d053c7a47df2e9bcce3e65f0490

Observation 02302a27-eca4-4170-b3c1-f439d6e8558b · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.898157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.515249Z digest=sha256:974b3c8ed1b77f12edd83baa03f892550936d22122ec0d0297ffd2350becef1d

Observation 8615f2d8-2c23-4121-9b1b-6851900bd419 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.517455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.517455Z digest=sha256:5c7715a59e281138d2a507e9edc0b8f1a1e22ef0aa20c5cb7797957da01d45b9

Observation c9e81936-ede2-4c91-8adf-4e591b2ceaf4 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.520090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.520090Z digest=sha256:917274adf2c53b151ffb34b789247c76349f8d36d08eb21aee439c81d94ae16e

Observation aa33f8c4-dc5e-4f4f-b329-bcc927e8dd22 · outbound

This paper cites H., Le, Q.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning H., Le, Q

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.522615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.522615Z digest=sha256:ff723844d2afe86e9ee78f5371cf846f4def2e0a0dd128bd2364b0e9f4d71271

Observation 6c52c4c3-7204-4084-8ba8-a61935fa4f62 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.525387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.525387Z digest=sha256:e9aecd0b86b2ba2d05e8cfcd5aad9f5f16d42fdaabc2df609e24857f56ecc338

Observation 2df7098c-9852-4f97-8a62-9c8852132e61 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.885946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.527957Z digest=sha256:37b42a40210ba386057382dadf79545b02a41199ddae8ba11e7db974a06c1b31

Observation 21d4e395-d743-4d01-aa42-52e48ab6cfea · outbound

This paper cites T., Wang, W., and Li, W.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning T., Wang, W., and Li, W

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.530115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.530115Z digest=sha256:ab794791bf0e6d42da0cfd8a01d9a3b20c43bc829d3afcd1ee7b29e7dc6328e7

Observation 8d0d3622-a270-4a90-b1b8-f81966e32d0a · outbound

This paper cites LIMO: Less is More for Reasoning.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning LIMO: Less is More for Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.532302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.532302Z digest=sha256:0f9eb34f43f2d6f752c5a08cc9aee75772cac0d29e9e6b7c51b7cbf904019986

Observation cb22ee4b-3048-4dc8-8a0d-33c50d00818a · outbound

This paper cites Demystifying long chain-of-thought reasoning in LLM s.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning Demystifying long chain-of-thought reasoning in LLM s

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:26:08.878484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:26:08.534689Z digest=sha256:ee7747519e14983406d1bece0e3379ed3e571514829ac46dcacb2a8069c1ef94

Observation 432f3e6a-5dd0-4835-afcf-14f230158f4b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:26:08.536798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:26:08.536798Z digest=sha256:a30f5334db5c865b657b027a7ecbcb9302f03af9dc1705e7e536634131298774

Pith citing papers

Observation cc5cb283-56ab-47d4-a8ee-94f90b6acb05 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.870506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:530d255ebb3dc2b98940baf19d4795f7de53f4732dd2a4253a513faa5bb7455c

Observation 73d7339c-e9ef-4e58-b417-a13640e62122 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:45:37.299666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:1456d4a7b4f490c8cf40c213afcf59fb04cda6b174e51472aae30d935d932b5d

Observation d0d6a44c-6a40-42d8-bd33-517004870c9c · inbound

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning cites this paper.

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.376328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:45:40.915428Z digest=sha256:fe5934e0bb9e4fe2f441a482655218d8bf4e15ebf10c99cddcef8be68c036a22

Observation 4757bf6b-42e1-436c-812b-69a03b3cd5cf · inbound

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning cites this paper.

Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:06:00.003903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:49:26.829527Z digest=sha256:a05792892140df23100515d93a41e079994aa4b8b940bf3d2e175424eff2867e

Observation b7e02888-5225-4b44-9d33-6f18a8d8d872 · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.658589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.658589Z digest=sha256:2cecee2e55ee633f0e05dc3db2ea378e33be92508041f4786d0ae52b301716d8

Observation 94b467e5-494d-45ae-ab3b-100edb5b663d · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:06.724007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:06.724007Z digest=sha256:42dc6c939cd8bda1015bd972c4393cd39ad530102ec2bd3186f7fd6e836a4691