Pith. sign in

Paper Citation Record · LEDGER

Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2503.10460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10460 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:19:53.921525Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a06483f-e30c-4c86-a0d8-fc1f430aa4c1 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:29:57.071592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:c60fe29fb2e56db2ab477381d88847675aed3d9cefd714717f4134ba7209d518

Observation 616e83c4-3352-43ef-b983-289c2dbd6c2a · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:17:02.869340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:207bc471391a0c4ce369b7a87a648ff8df02946d1fe817cfc3300bf8424634ed

Observation 89e88595-6b82-4a51-ba4e-fabbdd4e0f50 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.002577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:48e52bd26d6097f11dbcbba42151e45640d652b60e8919705780446f62208931

Observation d906625e-03ee-4bc7-90cf-733794c7e472 · inbound

Skywork Open Reasoner 1 Technical Report cites this paper.

Skywork Open Reasoner 1 Technical Report Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:26:47.350497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:26:47.283983Z digest=sha256:c3ba10bb9f8142cbfe51ad554e296e7b5d3774f4f3770528ee7b040e33003ec0

Observation ba84593e-9ae9-4427-8d11-3e07adbfd0b4 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T22:34:23.988237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:a84c9eeb03b768b8d35427991b41124e8059e1f4afca45fe71b90ec2900e0ef4

Observation c5b6fc67-3e29-4c6a-bd9e-45e90dcbf817 · inbound

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models cites this paper.

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:53.921525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:19:53.921525Z digest=sha256:cf0fda3503db09555a01d5bdee9bf73931cd08fb681bf0c1a0b5cd88e6b242a4

Observation d205813b-567a-43ec-9793-c23f36ab7807 · inbound

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM cites this paper.

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:11:01.110314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:11:01.110314Z digest=sha256:3f703df1430c2bda905d9354c1c57d36ed24bcbb7347e9902cdcc1069a87ff8e

Observation 08fe2101-9ae3-4671-9a47-74cd828efa0a · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.041310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.041310Z digest=sha256:0e238f5ac0644d3f0aca06d6c07505a58425d502fa0684940a23aee1d2efc78f

Observation 2df28670-0e00-44f9-a3c3-8772a2db9d8a · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.085149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.085149Z digest=sha256:9100902eb238aa9c1fc70c0ff2019c464f7d2feb7f8514b120580132202741d8

Observation e1c883b7-520d-4b4a-97f1-d8b690a827bc · inbound

Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval cites this paper.

Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:41.779153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:41.779153Z digest=sha256:96851079a7782415ddcf50f834ec28f22922eae9405495f89bd33188b2d193a5

Observation 400fda6c-c35b-4b51-92e0-b5081c768edb · inbound

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training cites this paper.

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:31:24.978441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T13:28:32.093512Z digest=sha256:b21677f8529e15a76ed588caaf4ac2d0a4048c463669c1a356ae02655c8abd3d

Observation b7b7f24b-3eea-4265-9b27-09e63c6e8bb5 · inbound

The Signal is in the Steps: Local Scoring for Reasoning Data Selection cites this paper.

The Signal is in the Steps: Local Scoring for Reasoning Data Selection Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:16:14.172268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T10:14:27.739531Z digest=sha256:f51aac2118dac7aca7c0ffa14a4b0cf5ddd8a201619bf7673d63b57da2997a63

Observation 6eb96528-8a20-4626-b3e4-0e81d6459750 · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:06:13.802986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:4a3be252b1c8a8c4e0a3c027cd9aa44726aa6ceadcef49391a3749aeec70d122

Observation 15289bb2-e131-49f9-ae78-6e8beb41fd51 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:12:23.662328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T05:11:29.205366Z digest=sha256:0f782cbd0c55d0ea650ca6e6e022b4350852288eb57bd38b02a4fb8b0b9e4435

Observation be7e2720-b16e-4fdb-ae74-588e4da07b09 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T20:10:34.891829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:06:16.172916Z digest=sha256:1df9dba8c6af3416d3e7d1c641cfef1d6f05d698a47a3a5197b7939d162372ba

Observation a4abdab7-eff5-47e4-8935-3ee46b435881 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.747822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.747822Z digest=sha256:4d1b6057cdac738d4657e80f81aa7d7633e3795ff797a8ee19a612d7f82ace42

Observation d14952ea-21c3-4e8c-8ef4-34e0280d2194 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:50.926478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:9a4605af5cf6a64f9c33866f0fd6d11b4b2ce382f3d7070ab378223d21d244f3

Observation 1bf0c8de-d4d5-449c-998f-c000e8621a0d · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:10:13.173355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:c798a81c09be28d99a7a788d86669d89ef8e396c23c07cbf9576ef234d5d03f1

Observation 09a06dac-043b-4ecc-bfb4-0a1933af2e20 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.521433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:cd3659979ac009d2deb789353b5871715e2269572e5161f2097d9b1c2eea2eb5

Observation 0e9be8f2-6c0b-4de7-98bb-6803cffdb2a4 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.755531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:4d44d1a081ad6971b1656d7126236b6176c2ccc23c3e4f779b063ebb7139ea53

Observation 5a49ae54-1398-4a86-9a44-d95a03d6c9a1 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:46:28.509523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:7d7aa00aade192796fe32b90b7c8f517aafcc64e714383a2c9cbf01db82c9428

Observation e229cccd-8685-440f-a136-6c54e5289bd3 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.839590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:5f41ffc8ecadc13d7fe22295fb28c941c27ef3b0a0643e03a55d3c78f72ea6ae

Observation 1d6efd00-eb3d-4326-8317-629164fefbbe · inbound

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning cites this paper.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:02.865013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:14:28.256032Z digest=sha256:56681098f75bac1efd9d499001af6e050aef3d06fe23255bf7a5c942a2989537

Observation d4a5b6d9-9cbb-40e4-9b29-12e08523d53f · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:09.548667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:e90e5c5ca4181db4a814ac288f8c6e4496421da30d850dbc88187d6a778f1f81

Observation ae67e4f2-82f4-4c79-bc4f-db0a819cbd93 · inbound

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs cites this paper.

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:12:25.200187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:05:48.244094Z digest=sha256:2ce68f3ab965e7c054e70e8410e9104cd0e58b6fb6d6a4577aa8539a1ec3c52a

Observation 221004fe-29e5-487f-8700-2e681c0208f9 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.717007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:2a836686e29e31c0c2b652e27761625c32fe6271498c0d8a4a948a0d774a9917

Observation ec1f147d-143b-4586-bbec-816290c2e37b · inbound

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think cites this paper.

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:58:21.051491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T13:56:13.827493Z digest=sha256:5bec3e04c3c24bcfda406a7378f93c4b8600deca74c8e4fbdd9f74db09d75a31