Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:53:08.348473Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2510.18814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:53:08.348473Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T16:53:00.860162Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-11T20:16:08.415773Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf39d836-2137-4f3b-84d0-0e6d6bf2a823 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4076495-9561-4ac2-a970-59b3cd42d405 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Reasoning Models Don't Always Say What They Think
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6342b43-de70-4242-ba2b-59ae70103976 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3e1e1d-bd81-450f-82b4-e482f3bb5800 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491c0c6c-141d-4f76-8c3e-12a663d4e1fe · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning OpenThoughts: Data Recipes for Reasoning Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98dabe1f-2958-439c-8e28-0816f7a67f9d · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b4f8c6-1178-4203-9f0b-22206d04bf26 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning GRPO) is similar under our experimental conditions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9243bc89-0d0c-449c-99c0-035e7851e51f · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ceed12-5a91-4ccc-9808-3734ef39a896 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning s1: Simple test-time scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ccc0185-6031-48d4-ad19-5cd3fa13cedf · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Accessed: 2025-01-24
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ee002c-3ebc-47a7-a92a-24d383a522f7 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae146751-903f-4e7d-9bc4-c3984c324908 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec71e96-83f3-4424-80e2-8c52bf4bf450 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9041d8d7-48e6-45f5-ab80-a72bc9f0c30a · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c81fab-bc18-4eac-aacc-45484db3baf2 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4abdab7-eff5-47e4-8935-3ee46b435881 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f83abaae-1da9-4cdc-a873-9fb1f7b67d01 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Qwen2.5 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5f9f11-bbfc-48ce-bc9e-fd9644904b20 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning LIMO: Less is More for Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c946287-8095-42a9-8015-c6cf3e942dce · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0600e3-ab3d-47dd-90b6-b523632f6565 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cc2718-4317-4daa-aa01-d73bc473d594 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cfdba6-2e01-470f-bbd1-8d739dea1c34 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Group Sequence Policy Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3f9da3-053c-43dd-a9cb-1d46c9d7af85 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning CONTENTS 1 Introduction 1 2 Preliminaries 3 2.1 Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c5c9e1-aa6e-4f4b-81d0-08eee520b591 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning We also provide ablation study forτ eval in Section 4.3.3
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d569e59-5604-43f3-ab7f-8d9ffc7f311f · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c857a778-3de6-4291-857f-798d494a7e6c · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836a0a5e-7ca5-4377-ac97-b2b72f0e58b8 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2895232-604d-418b-ba1e-5dcc43126f89 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde9b704-daeb-40ec-8a0e-f1086bca7d6d · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076618a5-094f-4da6-80c7-4231b5d28439 · outbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b1f823-e10a-4142-a696-91ed2d311549 · inbound
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 879b653f-a7ec-4d1d-96a0-6cccdbc684cc · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.