Pith. sign in

Paper Citation Record · LEDGER

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2507.07498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07498 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:44:00.020826Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:40:40.910461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.143168Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e97d3ca6-0613-403f-a167-1f9d2740b4e0 · outbound

This paper cites Program Synthesis with Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.513180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.513180Z digest=sha256:874d71e4beca0848fe67115871180f06205b9deae813e2a7c7f8dbe0065309d8

Observation 3498fbf5-206a-4956-b0ee-4d4b3abd7f41 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.595227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.595227Z digest=sha256:193b90e1b42e12c3e26c074973162d7d409ee4eaad05b1f8b5a9ac1d85cb235c

Observation f7db6b0e-5d22-48a2-b67e-999a3b922f9f · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.693037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.693037Z digest=sha256:502e2974d473b81a3c6b725aa893fd80517c26d6e1796506edaad0ce81c75ba0

Observation f3e74ca6-911d-4584-919d-901d130cc690 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.702849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:56.789538Z digest=sha256:edaf39ea3a5320c2be50aca612ddb1c8ad8c44a73731d455729d25cb5a932ef8

Observation a182342f-0939-4216-a5ad-c4014354a35f · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.867692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.867692Z digest=sha256:bc87ae808a5c826e0bb9edc896229c3c3f13a915e965a524c23960d429081ac9

Observation be213f86-508b-4eec-bd9d-7ce01221e1b0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:56.929622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:56.929622Z digest=sha256:feca777b3c763e5c59ae20e58b1d38aab258dae24f0718003830edbc7f388e3c

Observation d7be902d-f847-4705-b203-fc78ebf9373e · outbound

This paper cites Min, Gail E.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Min, Gail E

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.001314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.001314Z digest=sha256:ca8d7abf26a16385be78425ced871a2f379572180f7b789822eb56d60954ea1d

Observation 00ba49be-2db1-48ac-bad9-5f6480f21142 · outbound

This paper cites SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.092562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.092562Z digest=sha256:7f6c52dc6ab1cc52a1b44047f236dc091decff1097f1acb5829b0d6bfad3462a

Observation c2b5160f-9463-4fb9-82bb-23e5fda2d77e · outbound

This paper cites Min, Gail E.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Min, Gail E

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:01.615383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.183728Z digest=sha256:e98253520fe4bbaa837315054762c486c2122a5eb3f829916817d961a208e872

Observation 63d3f85b-8bcb-49b3-acb4-4b9c2a70d255 · outbound

This paper cites Kaiser, Wei Le, and Baishakhi Ray.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Kaiser, Wei Le, and Baishakhi Ray

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.248831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.248831Z digest=sha256:cb0b2c517259f75d3ecd910f8b959d0fc2a4533fde4bfa57d784e5cca92df238

Observation f8999b04-aa9a-4d01-b602-c9f5584ecc5c · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.306817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.306817Z digest=sha256:3b7dbbea6690e5f15f961d0f3d09ce9e53f2016ea970d1a1a5577cde163ed542

Observation 9933a054-1d12-4da6-8fde-87d0afd5cedc · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.538429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.360557Z digest=sha256:d3aa79df5bdf39affae95861d3f13b011e740f8c2545bcdfd516ea58bcbec36e

Observation 9ee7c40b-946e-46c1-9991-3dda6ac8a2ae · outbound

This paper cites Are We Done with MMLU?.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Are We Done with MMLU?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.410818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.410818Z digest=sha256:2aef887cf2fc65e1cc266fa8d13897876b6f2508864cb62fe5a4ab67e28c2121

Observation 2524e37f-3834-445b-ad3f-2766ed9f426a · outbound

This paper cites Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.462689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.462689Z digest=sha256:4167e246d8ed9ac38447a8cc195f20e29d7901444b5bac92303738bb9ec25cc4

Observation 7ea1aca0-17cf-44ce-bab3-d6912a0b2bc1 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.449710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.510655Z digest=sha256:7188046786ffd013d0d4dffa410ce4f28739728839ec0efd819ba2aaf811464f

Observation 89b15af0-e0ae-43b1-876b-86fffe9e5918 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.355905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:57.558937Z digest=sha256:649710f4fdec6129b1bf41c958767c9f002a1909e7635be62ea726bd63c0f42c

Observation b770dbb4-18f9-4cd7-8879-0dc42f6fe197 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.608469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.608469Z digest=sha256:daa954cbf00f8ec77708f3657ea6eed75a2f39f8fb916465dc05d8a628ba561c

Observation e5066ef0-7bd0-4cd0-84fa-93f0e1316849 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.651700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.651700Z digest=sha256:79867b2875a0fb32456329c2d92570cfe00b72b9943b491a74f04aa4c6931ba2

Observation 81f6f63a-f19c-4e28-9807-46b81b6a8dbd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Measuring Mathematical Problem Solving With the MATH Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.708756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.708756Z digest=sha256:e9d69daa39de35538e4f7c2828f6a39fb9b52f5fe3916dfb0c2d631df972b8af

Observation 4a38939c-a5c4-4976-952f-08e297d0c85a · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.785882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.785882Z digest=sha256:abe0a62d1c32eaf4d869e26414262d0c8b5c2e739201a677934adbdff3511326

Observation 0f54101d-8a92-4ba9-ae5e-9ccf6a4cad05 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5-Coder Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.839394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.839394Z digest=sha256:fedf509c73a5bae2b9a3c9df4fd0b251665c24b7db50f36e46a658676920c24e

Observation c01f51ce-a523-47e5-ba9f-cdb2e14be61a · outbound

This paper cites OpenAI o1 System Card.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.895612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.895612Z digest=sha256:a2102d4d11c0331fa223f866f95c999904420e1899b617a58dfef6f669c520ac

Observation fc6c8bd9-bc18-447b-8b49-db539a07e775 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:57.969802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:57.969802Z digest=sha256:70956a46e921643dc405b316d759547ea14b244ad74daf853cbd0eaddfc62965

Observation f775d378-1c6f-41d0-93b1-f1171f34f13f · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.283436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:58.020560Z digest=sha256:49b8059cfecb67ecce4861fda68595031ed7222426bb2e7b0ac2152d6a0e2794

Observation 493bf569-5162-4e7c-a312-81eb33bb145a · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.069083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.069083Z digest=sha256:0ca6ede35340ed063a0854661ed501d1ee3b2052768691e44e8541cc1aaccbbb

Observation 5ca038d3-4070-4412-93fe-258d7d9f6090 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.194426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.194426Z digest=sha256:3ab4f1ce51f8332a6934a2103ed397281e9cfdc6c843062c4920cbe2ab6c07d3

Observation 705bca43-4958-4332-8364-62c8169b340e · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.301904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.301904Z digest=sha256:c5410c3ad3c0fdab682dae02477f419035b5eaebb99290e94382e6a0e7f9d205

Observation c8d37bcf-d2fd-4899-ae03-7805c3e87a0c · outbound

This paper cites MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.430688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.430688Z digest=sha256:286a8d0335a565cb800393014f1e4a0edcfdaca11e424127801415206c964f9e

Observation abe9e047-01dc-4272-a012-e5869a9c0d20 · outbound

This paper cites HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.486504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.486504Z digest=sha256:819e31cd64cf1f302ec765ff47bd4d8ad81d382d0de02d03d3a26e4675987895

Observation cf41a461-6ba7-4a9d-97a6-6c755177b0fc · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.158488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:58.628898Z digest=sha256:8f07ed17614f80ee028e6e100548fec2f6c11d5d13b3a145e7636cdbfd2c2006

Observation 351480f6-b042-44dc-a133-69877045730e · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.771660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.771660Z digest=sha256:a7a8aee68605dc038f6f07a6fe1adfc82595d82b2e69df6556a76f4a5fe9389b

Observation 711345b0-5fcf-4d58-aae5-ae6f48db94f9 · outbound

This paper cites On the Impact of Fine-Tuning on Chain-of-Thought Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.869826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.869826Z digest=sha256:4f3c2da54e97052d4f910a47abc083174307580a3323e35b10b90a67bc1219f7

Observation 43a09bcd-2a34-4758-b151-876c4c53b183 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.929656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.929656Z digest=sha256:c6969630eccdf97cd55009916353ac93a7abf0fa6b95cf2696ef5ae46a3f26b2

Observation b02a2f70-fb09-44ee-a3ad-ec2234bab7ff · outbound

This paper cites KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:58.987141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:58.987141Z digest=sha256:de854f10814751f3387e5f4d28ae5687ca0723ede515f18828a916b317dd99d2

Observation 0270f618-f1c0-4595-84dd-2b347d0ec3fc · outbound

This paper cites Qwen2.5 Technical Report.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.046822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.046822Z digest=sha256:3544a8e4d3cb3835ee651f9756515b48c2186878720e041e7597732e8691371d

Observation 17feeeea-e0d2-4784-99ce-83f14916907e · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.109983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.109983Z digest=sha256:faf2267bc6cf3888743c960ae32fa36b5804a208d5c39e2180b64d25b5770c10

Observation 67fb2b71-9929-4c31-8644-0cc18dac0ee4 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.155801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.155801Z digest=sha256:37ebed64f88c666478718a5b91bd1f029cce00d48ca49136761162f6b0668f1b

Observation 2b48e226-7534-4af7-a9c1-1b5fcc0ea1a2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.205671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.205671Z digest=sha256:46fc81cf35169b1bdcec3cfe9e9ea8a08ef3c0b991b319503c25d59686a6e90a

Observation 62d1e1e1-8df0-4206-9df7-212aa93f435d · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:01.049327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.267078Z digest=sha256:3a8a2fda713f5b1b1e7956b223ac4aa926a48559ccdb3bb067b2b6ab43aaf055

Observation 4c774638-4500-4e1d-beac-2ba61ae4e4b4 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.311348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.311348Z digest=sha256:6448a540817b6bf29272216e00f02d0fa262d85b4adce570a9724049203b2b49

Observation d7dde5ff-939b-4562-b0e6-95af6b34ff72 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.388088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.388088Z digest=sha256:640afecbca58b6a0721466de6328ed4f71efff5fbf19d9d7d61e323ad1621fc6

Observation ff4a753f-d84a-4aaf-b0b8-e6b537bc62dc · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.471326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.471326Z digest=sha256:5ecfa46fe5d2fd3e6d2334c08f6d0e5f37f38434abeca68e6ad8c175bb83c81b

Observation 88366e24-9723-48b8-933d-b688da49350d · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.969329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.551743Z digest=sha256:90435466741c48358049575101a8f4bab4ac89d653b88c1c351e730864a50761

Observation 38559dbf-f1ae-4979-aae0-abcdcdc53f6f · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.614488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.614488Z digest=sha256:f989e045e3121d1746726b2ba4ba2bdbf2343766dc3b25689ed6365a3eea5823

Observation db14ff2f-6586-4fba-a676-ab1b1c522760 · outbound

This paper cites Le, Ed H.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Le, Ed H

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:00.892980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.669231Z digest=sha256:fe8b7b076e78bd8563d029b4f55327b49311df08a06c576cbb76d169ef177ed1

Observation 2e0b8cd6-af3b-40a5-8efa-256c0de4201b · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.785136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.719102Z digest=sha256:01736b071d78d16301d9f93bb5b990260e69b1bd3261898e14cac8d98d7347ae

Observation bf2044ff-eb66-4004-bf51-d582acf36600 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:44:00.703241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.764041Z digest=sha256:8e5dbc9ede13948cf91b23c6a5a168e18fdc1157ccefbd27d2c0c29e0399e59b

Observation 92874828-cc70-4d5e-b29b-aec626613d50 · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:44:00.584735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:43:59.808781Z digest=sha256:f82517b86ae88f0a2879dbd09ae715c7f777a9d384679786e979041beada684e

Observation b8b0e9da-da06-47ad-8c4d-09724995ea1a · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.848050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.848050Z digest=sha256:c3202011346e89ee439330ffcdedd86ac8434ccb4fff0b819bc3edf72cb21b44

Observation 99a139d5-32ad-498a-880e-a8739508a026 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.893251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.893251Z digest=sha256:08d11c9e30a2a1df09566fa6028a4f46c8af9586a08a99ef0837d8bbdb9c5d39

Observation 498cf9e9-8b5a-42f6-9848-f9b4f31190ec · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.932862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.932862Z digest=sha256:b25ac954151fdae7c732513b0aa1bdd85c56d66f8dfd4141addc84f54475fc1d

Observation 536563bd-f0ae-4a38-ac31-fe7cc09dc943 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.959396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.959396Z digest=sha256:70c4dd32eee4b1b8a8097413fcecd52d7ef8f7aebc486579425328a98132e02a

Observation 85f4a8c5-380c-40c7-ac76-6aedd63ed9e2 · outbound

This paper cites Learning to Execute.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Learning to Execute

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.965594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.965594Z digest=sha256:7cd0109b983631e7522dd04572e66cccfb3ea9f34764786084b075bc35f7d54a

Observation f7351096-4071-4a91-8b50-e1129681d1ba · outbound

This paper cites an unresolved cited work.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.970269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.970269Z digest=sha256:1742dc605f63b041badeba537f5931a935845b5d9ea848cbc3bc6acea9fc949f

Observation a5e5406b-ede2-49b5-b534-889a281da0ac · outbound

This paper cites online" 'onlinestring :=.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:59.987191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:59.987191Z digest=sha256:48cf4440401c73b43677bd098f7a206d7b8ba6ea76f3ac840aa36d28fc62785f

Observation 9cff075c-7ed1-42d4-929a-396445ae561e · outbound

This paper cites write newline.

Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:00.020826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:00.020826Z digest=sha256:380f5624222268ace8ce082b25bef5f1bdc13dbee7595a1f00f13b8566b72d52

Pith citing papers

Observation bd083f39-b368-43e4-831b-cbe8ff80f22e · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.146090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:b3dce5c75aa3ef1690f3e082f0f914347f373fc446db796133eb578480c636fb