Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:48.296848Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 12 inbound Pith citation observations for arXiv:2507.02841.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:48.296848Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:53:10.631113Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:00:08.075019Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 53902c7f-ceaa-4cd8-ae04-df45ca4878b2 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657f7dfc-fbbc-40ff-9958-12be451463f3 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a0d590-15ee-4f88-9130-d17931502e63 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5086abbf-6a0c-47bf-89e1-13bc93dfb8a8 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806c79e2-5a96-4689-90ea-3f25cfc6655e · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Measuring Mathematical Problem Solving With the MATH Dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b383e2a9-380d-49a0-bf16-fe6abb213d44 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200c78f7-87fc-4304-90aa-fd1b3003eee1 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac67542-5f34-49c7-b98a-53504476c0d9 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e0f7d7-0436-40f6-9d38-34d50fec9880 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aaf876e-1be8-4460-aac9-ebda4c7a5c30 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Asynchronous methods for deep reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f6ef865-02fe-46fb-a6e1-f9f437debb17 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70addeb1-6245-49eb-9b54-4269a32733bd · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason HybridFlow: A Flexible and Efficient RLHF Framework
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f8688a-5a95-4f2c-8ade-b482b719f4a2 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fcf59b5-b3a2-4d95-ab77-06c710e09d8c · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7f977f-f13f-4f48-a2e8-ee4c534fe2b4 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Learning to Reason under Off-Policy Guidance
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 188137c9-c0d6-41d2-bfc0-301584891c98 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Qwen2 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3148fa1c-30e0-4b98-acfb-40c08c2a5224 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1551ff57-ec59-4a4b-b813-29883e73e219 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59db289d-0fdf-468d-b436-3c737347f318 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ad57c4-f146-45a3-9a69-58be11b90cb3 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26de7637-ffdd-4a6c-869b-47c3972b6a63 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1948
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c032893-f8dc-4e88-9ad2-bbd2e931f2cd · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Orca-Math: Unlocking the potential of SLMs in Grade School Math
Reference 2003
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f467dd-ab2f-4c2b-8753-a8a7af1b99b5 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b04e2f4d-a1b8-467c-8b58-6f70f84d6524 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Reasoning with Exploration: An Entropy Perspective
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e9def5-e837-48b2-b972-09dd171b7349 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason MathPile: A Billion-Token-Scale Pretraining Corpus for Math
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f84943-4e26-4301-82be-f15ef631ab84 · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Llama 3 Herd of Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f09d8f1-7c70-4020-8387-322b242215ba · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98de604d-3fd3-4c47-9df0-62ca1eb1dfdc · outbound
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Evaluating Large Language Models Trained on Code
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63ed6416-f211-42d3-8ae8-6792164f8643 · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a95a1a-5436-40a1-9875-db81b3a9ba99 · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da54ada4-7c1f-4b5b-8944-c9de425040ab · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da2a00c6-4e88-4cb1-998e-660119840990 · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50102c52-6333-46ce-867d-49c934cbcc78 · inbound
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48091707-a566-4411-ab25-00bddaffb930 · inbound
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 604f0662-faf6-49e7-87a6-41b31def1ce5 · inbound
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 102dbb83-68ec-464a-a929-a3df65e0e30e · inbound
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5093c845-1be9-4bd5-a621-1039b93d1e57 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49501d2a-a3eb-4771-bbb5-a115b9e2e2b4 · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e162964d-6e0c-4a53-96a2-b7f75a8d13c6 · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b0118de-0224-4485-af04-c1dad64f514f · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.