Pith. sign in

Paper Citation Record · LEDGER

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model

As of 6 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2606.22317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22317 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T11:10:56.755558Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:08:22.453766Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:08:30.124848Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact20
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e539e23a-43b7-4966-b543-bbe156f47c81 · outbound

This paper cites Rl for reasoning by adaptively revealing rationales.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Rl for reasoning by adaptively revealing rationales

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:6d0401656d4aca196655694b3e6456bb9d3e0ccd35b3caadbbaf8bfbf1fef384

Observation 787e5272-d8f9-4bc4-b96c-1ffedd6855f8 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Online difficulty filtering for reasoning oriented reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b8f004888976288f7dd3f270bf6e344379fd7d4a96c40bd22214356804a1c478

Observation 4a135eb6-9869-4771-b9e7-e43eefdb1b66 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.281464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:e743be6e61ebbe6198ea650c8785c093ea5880d8aed328f09dfebacc030c020a

Observation 53c92407-fa65-4bcd-a92d-6d8c29db6b0e · outbound

This paper cites Unveiling the key factors for distilling chain-of- thought reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Unveiling the key factors for distilling chain-of- thought reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:d3bfe550cb1e3ccc2b0cbddb8dbfaa49d1ac1d6e4930b5959aa9818af1e9f901

Observation b8e2a08b-f532-431c-846a-ddc673baa1ad · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.312906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:324e59ed4524974a09682bb30c32a76327950644cae53d840566eacf9563fc51

Observation 64a05973-4fc6-49f6-af85-097bb1ea1560 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.306067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b4b8e2a2bfe4629253e3c1485db569a4be2a8e2ed23bcc50abe5478e413e7460

Observation 3151905a-5e48-41e1-b11a-8801af193a51 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.Nature, 645(8081):633–638, 2025.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.Nature, 645(8081):633–638, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:5a2ed31ae76c9f4b3bf97af81d1b961a93258e1c917a4d21f0c0854962e3b8db

Observation ad351469-98a8-4e59-8bd9-405ba2ef7126 · outbound

This paper cites Omni-math: A universal olympiad level mathematic benchmark for large language models.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Omni-math: A universal olympiad level mathematic benchmark for large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:1f9a13f09dba524442fde5d7ec3e8ae86fac1e3a24380bdebabb6ae5424d6b21

Observation f8d10b29-0fe4-4ea3-b166-67a56fbff509 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.300162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:fa3a5f214d630f05238991f1c9d88764cf75213b496399e3b0fbbdbc4b221d9e

Observation efe3cd24-14c7-4002-9329-7febe15112df · outbound

This paper cites Rewarding the unlikely: Lifting grpo beyond distribution sharpening.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Rewarding the unlikely: Lifting grpo beyond distribution sharpening

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:de6ac1e49ed7ad575b72b1f54f9bd7a6ff08a7c3a548ec63f77a8d6f88fae9ed

Observation d75c0abc-3b57-43a6-87f4-ba96f7b181e6 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Measuring mathematical problem solving with the math dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:c97bb750a0293f4bd0daf2d9dbe1ddec0747973037b750bfe1bcf97e145a4354

Observation fe5532e9-d0c2-4fd2-ae39-4ada51a27e06 · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:8f67cdc11d3f63f5e65a0c44b24a12668162b7dd751445bb2228d33f113ad3eb

Observation fb9e5769-20cf-4bb0-9e43-29e2eb5fca8d · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.309663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:60f773926894189cef772cfe858ad603d23023dcd18f2bdf6ec4c26683ad9d5d

Observation d3229f20-a436-4298-9568-fa00c3fededc · outbound

This paper cites Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.303465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:8e1b9f4e3e40e0100163ce9b04796d445e8b81f6a7e4db72afc85ccba93d2014

Observation 323258c1-76b1-450c-9c1d-2d951894027c · outbound

This paper cites Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.251664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:f344047f6acf6de289b912ae3d3168fa9ced77d05fb9e669bb04eec4025b9525

Observation 891db1dc-2849-41a6-beab-792e3241c235 · outbound

This paper cites Unlocking the power of function vectors for characterizing and mitigating catastrophic forgetting in continual instruction tuning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Unlocking the power of function vectors for characterizing and mitigating catastrophic forgetting in continual instruction tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:c95d34062ab1125df70163e1e0cb254a0e0ba9c3f5ade53e9edddb43cf2a51ff

Observation 17048d55-71be-4e19-acdb-9a311510ee34 · outbound

This paper cites Tacler: Tailored curriculum reinforcement learning for efficient reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Tacler: Tailored curriculum reinforcement learning for efficient reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.282278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b567c9defa5b175e51c22ce1f8f664313c8b1fd7da7d06eb16603e464a1a96f2

Observation 17451c98-d139-4776-a6b5-bc6c2dd4e791 · outbound

This paper cites Language models can easily learn to reason from demonstrations.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Language models can easily learn to reason from demonstrations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b8c1a5ef9c1ae8f0d8e6128b2c3a128365cc2f6967177bfbdbea0f6dd873ce06

Observation decc794b-e1f0-437d-9233-7f06515cb487 · outbound

This paper cites Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:cdb67e2bb17738c0a908aea4b1520cf042689de6887719e483d7ae5438e889dd

Observation f8ad80ca-2fa3-494f-8906-9fb1e7df4371 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b420a3fc5bb6ab9b7d4601cd539e5278d843710bc2c61a29797c994173cfba8e

Observation 7482cacf-f7f2-4d09-a49e-852c9a44988d · outbound

This paper cites An empirical study of catastrophic forgetting in large language models during continual fine-tuning.IEEE Transactions on Audio, Speech and Language Processing, 33:3776–3786, 2025.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model An empirical study of catastrophic forgetting in large language models during continual fine-tuning.IEEE Transactions on Audio, Speech and Language Processing, 33:3776–3786, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:fd673d3ca06b1642b20357ed29d6a2e226671cd925fee3aa21bc96f765853723

Observation edaefd69-d03e-481a-8969-5ab596b96f69 · outbound

This paper cites s1: Simple test-time scaling.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model s1: Simple test-time scaling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.302931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:001936fd77b4264fb8e9eb5511f5d8777c7fa8fd03851d9fdcc1906e5ee9ddd7

Observation 723e8753-b778-4129-922a-d40edde2e06e · outbound

This paper cites Openai o1 system card, 2024.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Openai o1 system card, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:e1db443e030ce2a6caa50b79d4916baa13da44555cb7b582a78d64441b486c8d

Observation 016d0c7a-9751-40a5-8080-0e77cc613d27 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Curriculum reinforcement learning from easy to hard tasks improves llm reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:e08322c8a5d8d72e7675cf182cc91e251beb4d20b4e6027c71d0eb748f3baf3b

Observation 8362972f-329f-4e83-bec4-80781be32119 · outbound

This paper cites Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.305959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:03bef2019d953ed580c73959ad959bc7ee42ab86cfda8dcf729d82ca576f0bfc

Observation b8c90504-fcc3-44f6-8b25-198bfd340ed9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.296432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:349976a92e5e43fcd4cb08e6fcf2f480d4cbd3c3fd53c8517170e295f3c384ad

Observation 7ef419e3-2fbe-40db-8337-6893563729c0 · outbound

This paper cites Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.245665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:f3012f48fe03fc31284bca7943eed5c5cdfeb6408e9901912e752aacaec600ac

Observation ab311a9c-566d-43ae-b144-be5d30a34c1d · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.235813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:ba31ac155b05a22841294e3a7b2afe06ab226a0feb224502cdcd24832f333b7b

Observation a0133171-3f07-48a0-9401-bcdd1fa39746 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.240336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:d7f7714bc9820cc43ffc84fc0eca8453f53111a8c15c13487bcfa8f25583458e

Observation 8c149a40-a299-411c-8ba8-5a3bf370ad9c · outbound

This paper cites Continual gradient low-rank projec- tion fine-tuning for llms.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Continual gradient low-rank projec- tion fine-tuning for llms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:aa06bfbc1f2b9a96760a129b4398968d1ea1a6ebcd25ef68e43a2d33ef732272

Observation 06b8fae7-b780-4c1f-8dbe-e5be94315c3f · outbound

This paper cites Inscl: A data-efficient continual learning paradigm for fine-tuning large language models with instructions.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Inscl: A data-efficient continual learning paradigm for fine-tuning large language models with instructions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:734ac70c19af7393449a3b70da79acfc3dad8fad1b91f4025febf632df59ed13

Observation d7643db9-d66f-4523-901c-3a4510bdcdbb · outbound

This paper cites Reasoning scaffolding: Distilling the flow of thought from llms.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Reasoning scaffolding: Distilling the flow of thought from llms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:a770839f18d389f94a2e5a30673a20aa444b3ecd7f096c6562a002f011f6f8e8

Observation f5286da5-28a9-4042-aef8-60ae367d41e6 · outbound

This paper cites Reinforcement learning with verifiable rewards im- plicitly incentivizes correct reasoning in base llms.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Reinforcement learning with verifiable rewards im- plicitly incentivizes correct reasoning in base llms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:c0efdf393e53ad0b329d10e8d96e8f5f5e78490bb3a9d2d8d3ab3ad8e96d774e

Observation b972c2da-bbc1-4150-bcd6-8612b0d312db · outbound

This paper cites Enhancing long-chain reasoning distillation through error- aware self-reflection.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Enhancing long-chain reasoning distillation through error- aware self-reflection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.294055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:fcb51454ce466558dda5c473554320a84d348adc7c35af522e3d259d655ba828

Observation c5d14caa-9f01-4289-b406-a09d72b93085 · outbound

This paper cites Training large language models for reasoning through reverse curriculum reinforcement learning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Training large language models for reasoning through reverse curriculum reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:df695335e6d3a9ed67c6c484517c0a21f56611b3a1b641da4226d2480cd5bc29

Observation 839f41f2-2611-43a2-8774-b6641b528f98 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Learning to Reason under Off-Policy Guidance

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.270996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:9fe8f8b8aa6fc22d8b9ea0153451a597c08c8d6df668e3985888b28eb9f24326

Observation fbefa8f7-7a04-451b-93e7-c40665df5795 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.246400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:e72fc00b8a72e804db4b8e0c14c911b0de897a55cb9111bbabb27ad54fb59b86

Observation 718c7f29-425f-4845-8d62-542048da84d5 · outbound

This paper cites Qwen3 Technical Report.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Qwen3 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.291866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:18fb5ab6c808ed27089b0958a5a48fbd10ae3207c1b298fa78483b99fd708c14

Observation 95a29bb8-2dc5-4a77-bfbf-602c4f97dede · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.287324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:f0f7e6e139ac58dadb0c170ac080837716667a6c92e78fd4eedded24eaf570da

Observation a70a5b3d-ba88-4aea-8d5f-8f72a9783c5d · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? InThe Thirty-Ninth Annual Conference on Neural Information Processing Systems, 2025.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? InThe Thirty-Ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:9f2ff08f6c0b2091f96deca59a3a9eab0d4f5b90982497cd505bd77f1c6ec696

Observation f7a37b8b-b608-4975-b3e9-7ecac6a0816d · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:c89136bef124ed28367f4c0ff167c39fd02931f8e9c51c3a07cbe86b7922c687

Observation c27be518-5a20-411d-b4b7-0546b5c51438 · outbound

This paper cites Cures: From gradient analysis to efficient curriculum learning for reasoning llms.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Cures: From gradient analysis to efficient curriculum learning for reasoning llms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:dfc6bd6c32536bcec80eebe945e8dbbbf91a9a7d26526df01e9d05b0992ab490

Observation 18cc8457-84ee-4f9a-b06a-a2255195ebb4 · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:d43e5556138aa4144e8bed0e192ee0d8e9a89b0400d129ff7a77722df294f8b7

Observation 36f0bdf7-7d43-43d0-b75f-8fb4811e8bfd · outbound

This paper cites Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:b72c2cf2b2f79b459c0ae739d79a29335a2636807564e30ecf2a59c27adf7b69

Observation 98c9bfd3-6cf0-48e0-8cb6-b7c8f7025cd5 · outbound

This paper cites Automatic curricu- lum expert iteration for reliable llm reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Automatic curricu- lum expert iteration for reliable llm reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:1dbe47a7ec16b64f3dc3cb9b93e7fc106771d9c069f4e404342d37ce7c2c113a

Observation 25c971c0-9475-447f-a550-646d002e9c3e · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.309787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:4e728a1c00c195a11f2bc8235919f9b7e70675615d4f9667cbbb5fe874984f0a

Observation 9860daca-8e29-48b4-8ae9-2150d7e84b98 · outbound

This paper cites Group sequence policy optimization, 2025.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Group sequence policy optimization, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T11:10:56.755558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:6c0899983a4a7269cf03f4a27b0d4aa7660352f0b0b20306c570af5fcc9ada0b

Observation 9f250876-a5da-45b9-91d1-02e9001bbc6d · outbound

This paper cites Spurious Forgetting in Continual Learning of Language Models.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Spurious Forgetting in Continual Learning of Language Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:42.278601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:e33e37adc03be4488f3e31c3103ff81cac573f5e2d3802466a5c3bcd70a5a02d

Observation c0543458-410c-43b3-813c-6a4ba4d1f2cc · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model TTRL: Test-Time Reinforcement Learning

Reference 49

Resolution
malformed identifier
local_arxiv, observed 2026-07-04T08:39:42.296851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:bdf309d63d996efdbb8afbce332823147d8097e1bff425b209c148587bb9b5af

Pith citing papers

Observation 98747308-700e-46db-9866-9d4b7a981365 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:08:30.212990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T00:08:22.453766Z digest=sha256:eb521c2d665aaa355fc727b3202f487cd3f309eef299d80bc486502368183175