Pith. sign in

Paper Citation Record · LEDGER

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2510.18814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18814 v4

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:53:08.348473Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T16:53:00.860162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T20:16:08.415773Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf39d836-2137-4f3b-84d0-0e6d6bf2a823 · outbound

This paper cites GPT-4 Technical Report.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.173699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.173699Z digest=sha256:cd6488f2687e4511e7e581b9a4c59c37cec6634b0e0c42a934b34a3aec71f1a9

Observation b4076495-9561-4ac2-a970-59b3cd42d405 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.547904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.547904Z digest=sha256:7e7266b20d13c434d2169ee9cf84643cb70d5668c0955d5588a7a143a0710b55

Observation c6342b43-de70-4242-ba2b-59ae70103976 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.655329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.655329Z digest=sha256:ae2d60f67e9159971137f5ff5ed80258f4425286fd617fdb049c34e2bd3b2d64

Observation da3e1e1d-bd81-450f-82b4-e482f3bb5800 · outbound

This paper cites The Llama 3 Herd of Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.768844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.768844Z digest=sha256:737b50f39f12e7fab072f64d6fc5ae17ffb3581b78ff0b05e09be75e10664032

Observation 491c0c6c-141d-4f76-8c3e-12a663d4e1fe · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning OpenThoughts: Data Recipes for Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.911880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.911880Z digest=sha256:20e268f8bab4f13ba5327df819294f72780d71b9abcdf972805b51e2406af976

Observation 98dabe1f-2958-439c-8e28-0816f7a67f9d · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.053247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.053247Z digest=sha256:2a2f3b949927d1f7a8a8f7b8eccb04b47f58cc894f355fa05de2e5d028483985

Observation b7b4f8c6-1178-4203-9f0b-22206d04bf26 · outbound

This paper cites GRPO) is similar under our experimental conditions.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning GRPO) is similar under our experimental conditions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:08.348473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:08.348473Z digest=sha256:c4325eea9c1a00592fc97cca3e35b82f1a4c382a5cfd64e82aa5d256b2c77c74

Observation 9243bc89-0d0c-449c-99c0-035e7851e51f · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.544194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.544194Z digest=sha256:3fab47e81a5e0fce1182a976e280f9d39a85fa2e393d79133810d4181916c3a4

Observation 27ceed12-5a91-4ccc-9808-3734ef39a896 · outbound

This paper cites s1: Simple test-time scaling.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.710900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.710900Z digest=sha256:e5cf089c97c4662708a8880261d426f60c17d498260f6233b1d3e6c427760a04

Observation 1ccc0185-6031-48d4-ad19-5cd3fa13cedf · outbound

This paper cites Accessed: 2025-01-24.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Accessed: 2025-01-24

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.881699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.881699Z digest=sha256:c3d4319155f98dab7fc3124a9d1402bbf2e7367e68f62824abe0aa6abeb7d16d

Observation a4ee002c-3ebc-47a7-a92a-24d383a522f7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.031456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.031456Z digest=sha256:27bbb91b8ffb98b4c8cfd3791a6f0f6300c2b87dd0b8379424465475ba9d95ee

Observation ae146751-903f-4e7d-9bc4-c3984c324908 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.327544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.327544Z digest=sha256:08ad2ef6d7b78f7c3c98e22485d67be8a9d717541a46247a4965adbc26515d19

Observation 5ec71e96-83f3-4424-80e2-8c52bf4bf450 · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.434034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.434034Z digest=sha256:351a60b62c26d8760dbe2c81c319363b396f0124c997ce471451590217b2d19d

Observation 9041d8d7-48e6-45f5-ab80-a72bc9f0c30a · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.536180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.536180Z digest=sha256:cbf9746ab453968c586c751375ec53051b2d416313c014127856454c86511ed7

Observation c4c81fab-bc18-4eac-aacc-45484db3baf2 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.648043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.648043Z digest=sha256:48b40f264351360c1c27a53f1191a34a322b3f3aeaa268cd0dc3958e2d4e5650

Observation a4abdab7-eff5-47e4-8935-3ee46b435881 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.747822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.747822Z digest=sha256:62fda5dda59f871d4c977b332c0b06f815458e0cde6655f0b5ab5497d6ff15de

Observation f83abaae-1da9-4cdc-a873-9fb1f7b67d01 · outbound

This paper cites Qwen2.5 Technical Report.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.799133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.799133Z digest=sha256:957549f6c5d2fe7429c4b76010c0fdd3e248c7148fd11b9ace82d2b81901e684

Observation 3b5f9f11-bbfc-48ce-bc9e-fd9644904b20 · outbound

This paper cites LIMO: Less is More for Reasoning.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning LIMO: Less is More for Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.877463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.877463Z digest=sha256:b7761234fc00795cda5e9c74122df0eb0e5e80f4e0447a6351dd4c39901b6a43

Observation 5c946287-8095-42a9-8015-c6cf3e942dce · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.926195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.926195Z digest=sha256:33e4aac55d502fcd1f3c5c60a7379f8d0047f96f4992d3dd00ba23e4bfcd2cb2

Observation ff0600e3-ab3d-47dd-90b6-b523632f6565 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.999980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.999980Z digest=sha256:6dbe82e4ef9c6c8f0167b0f7306b178ef8c1101a27fe89069f5027e03542308f

Observation a1cc2718-4317-4daa-aa01-d73bc473d594 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:08.054190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:08.054190Z digest=sha256:af9d903f7255ea2f29b97274862533024ebe62917c3f8bd4fc20113e73e71e48

Observation 44cfdba6-2e01-470f-bbd1-8d739dea1c34 · outbound

This paper cites Group Sequence Policy Optimization.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Group Sequence Policy Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:08.129407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:08.129407Z digest=sha256:c808d2a4f113fd914ee6cb9be3b09352e56b15a722938d5d2ae5d63c6855741a

Observation 8f3f9da3-053c-43dd-a9cb-1d46c9d7af85 · outbound

This paper cites CONTENTS 1 Introduction 1 2 Preliminaries 3 2.1 Language Models.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning CONTENTS 1 Introduction 1 2 Preliminaries 3 2.1 Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:08.208235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:08.208235Z digest=sha256:6b72bde93a46a74b5e08dd453909ef845051f901e5a733aef57321773a11948a

Observation c3c5c9e1-aa6e-4f4b-81d0-08eee520b591 · outbound

This paper cites We also provide ablation study forτ eval in Section 4.3.3.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning We also provide ablation study forτ eval in Section 4.3.3

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:08.264018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:08.264018Z digest=sha256:c938eb22dc3b1dd6f980cf4981857aa57e5f02bbf4c73eb57fc78eca494c3912

Observation 1d569e59-5604-43f3-ab7f-8d9ffc7f311f · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.162212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.162212Z digest=sha256:97f343c2eda7ef6d24322ea454d522a21eb704362ecf1f34f0b83fb3ce315aad

Observation c857a778-3de6-4291-857f-798d494a7e6c · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.305427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.305427Z digest=sha256:137d41f0b3721914f3437d3491e0e2b494f881ff43e4d28284585c4e66addb05

Observation 836a0a5e-7ca5-4377-ac97-b2b72f0e58b8 · outbound

This paper cites LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.449889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.449889Z digest=sha256:1018c76a94eff9cbaadff79e225f88061a27794eb43f2a87f55728632613b598

Observation d2895232-604d-418b-ba1e-5dcc43126f89 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.285310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.285310Z digest=sha256:adc2eb9d3816d556937fa54f9e3dadd9e9425a80e0bcb9aa9d512ef33221e853

Observation fde9b704-daeb-40ec-8a0e-f1086bca7d6d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:06.195668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:06.195668Z digest=sha256:761abbe27ed82c3c59fdeae43b4e97d53cc92fd1bda1349896301e11adcfadac

Observation 076618a5-094f-4da6-80c7-4231b5d28439 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.427930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.427930Z digest=sha256:fbf66e42295667088348cf1c9db3c2e473f02460e9bde8f19d9ece147a9e7ecb

Pith citing papers

Observation 85b1f823-e10a-4142-a696-91ed2d311549 · inbound

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation cites this paper.

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

Reference 198

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:56:08.204338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T16:53:00.860162Z digest=sha256:ff748b17b7b0a912cdade478fc102de6b5e03731fec60fcd5a63833be68ed4af

Observation 879b653f-a7ec-4d1d-96a0-6cccdbc684cc · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.418279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:5ec00785326418653cd7e954c7ebcc6c2ef9fea63481a6fd06f94bbcf85542c9