Pith. sign in

Paper Citation Record · LEDGER

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 12 inbound Pith citation observations for arXiv:2507.02841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02841 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:26:48.296848Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:53:10.631113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.075019Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 53902c7f-ceaa-4cd8-ae04-df45ca4878b2 · outbound

This paper cites Concrete Problems in AI Safety.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:45.550459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:45.550459Z digest=sha256:0c2112064b7897e7dec1066a817d78ef69fa29a598e4f5054143250c73b08582

Observation 657f7dfc-fbbc-40ff-9958-12be451463f3 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:45.878265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:45.878265Z digest=sha256:0871f096f9198f57eff9eedb480ca053b06ca86f400735b26e43180afd4aa63b

Observation 93a0d590-15ee-4f88-9130-d17931502e63 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.011935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.011935Z digest=sha256:2ae6a07fbdc63159366cf8e63f626a3b8d4699b1f8197cd7d8fb52d2fbf818b9

Observation 5086abbf-6a0c-47bf-89e1-13bc93dfb8a8 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.372789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.372789Z digest=sha256:49eafd123d7d7c9de35cf19d219f59055f2c518f926abfdf15de5227545a59b1

Observation 806c79e2-5a96-4689-90ea-3f25cfc6655e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.458931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.458931Z digest=sha256:856dbf080852cc19f4df317d027273959ac27f32e4a67d9bc42183cba24ba3eb

Observation b383e2a9-380d-49a0-bf16-fe6abb213d44 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.553671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.553671Z digest=sha256:e42ad1a46f3085ba2ebba8c96e930fc15b196e18f4eaaa22f4076241cb3615e3

Observation 200c78f7-87fc-4304-90aa-fd1b3003eee1 · outbound

This paper cites OpenAI o1 System Card.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.651273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.651273Z digest=sha256:942cabd9e2e224e8eb986f067a99fc7651e610e090816a37b77fc822677ccfdf

Observation 6ac67542-5f34-49c7-b98a-53504476c0d9 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Understanding R1-Zero-Like Training: A Critical Perspective

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.745830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.745830Z digest=sha256:f39046f700fb9b8c9c92e964edd8ee8b3008ce84d086d91703a58b628ee14082

Observation c1e0f7d7-0436-40f6-9d38-34d50fec9880 · outbound

This paper cites The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.857127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.857127Z digest=sha256:a0c8d1e80ecd77273ce5ead7b5caab170af4e55ee5eaad245ac6c32fbbab9392

Observation 7aaf876e-1be8-4460-aac9-ebda4c7a5c30 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Asynchronous methods for deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:26:48.607497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:26:47.065513Z digest=sha256:563036c87f6d331013eeb225aba9012aca79dbdb477d23f7b1e59d5bff43c8f0

Observation 1f6ef865-02fe-46fb-a6e1-f9f437debb17 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.156279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.156279Z digest=sha256:e768614605090241fed2e2d442034eb328bec5c78eb2a07a46ed43060594a5e6

Observation 70addeb1-6245-49eb-9b54-4269a32733bd · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason HybridFlow: A Flexible and Efficient RLHF Framework

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.348775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.348775Z digest=sha256:0ced39f280c135090a9c73c9c63d2f71a560a403279fb91f690ac89ba9243f8e

Observation 93f8688a-5a95-4f2c-8ade-b482b719f4a2 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.488611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.488611Z digest=sha256:3d97e124cc98aaea4b53ac54599fd4268af0493829627b85c5f85cdf52acd61c

Observation 1fcf59b5-b3a2-4d95-ab77-06c710e09d8c · outbound

This paper cites Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.740169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.740169Z digest=sha256:8b86d71157e29f0960992150ce5dfcc25a701cf2892082be5213553bd44a84bd

Observation 1a7f977f-f13f-4f48-a2e8-ee4c534fe2b4 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Learning to Reason under Off-Policy Guidance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.836925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.836925Z digest=sha256:c044a8b04fc62abc765ec0a4268204ee58eaf59daef27b395465b2df9d7c344d

Observation 188137c9-c0d6-41d2-bfc0-301584891c98 · outbound

This paper cites Qwen2 Technical Report.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.918660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.918660Z digest=sha256:ad85bb3f7128e9edaae48ef5dbefe542a8c5086848190b6790179fa4ebafe363

Observation 3148fa1c-30e0-4b98-acfb-40c08c2a5224 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:48.035446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:48.035446Z digest=sha256:b89fdf61156ab658c1b500e65fb427061669756171163adee58eadaebb51743a

Observation 1551ff57-ec59-4a4b-b813-29883e73e219 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:48.085634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:48.085634Z digest=sha256:29c24458ec3da34f7ba6aa3d39603fe47764c55b2b7abf17beca3593c88e0146

Observation 59db289d-0fdf-468d-b436-3c737347f318 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:48.151553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:48.151553Z digest=sha256:472bffa7142d01f3a0773028afadf482e70b4a74d54280c531886c3cf726a74c

Observation 23ad57c4-f146-45a3-9a69-58be11b90cb3 · outbound

This paper cites an unresolved cited work.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:26:48.597951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:26:48.296848Z digest=sha256:cd175661c2087c4b56fb0201e5a13460b0237eb74d851601a324b3600b955608

Observation 26de7637-ffdd-4a6c-869b-47c3972b6a63 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1948

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.242484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.242484Z digest=sha256:dcb17559ea983621f2337608a476002a9519b53ae30925b8f80c75dedbd72869

Observation 7c032893-f8dc-4e88-9ad2-bbd2e931f2cd · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.985350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.985350Z digest=sha256:3d301fc384177005c4fe97afce9559242ce5502dceac9812c8c09afddf696168

Observation 42f467dd-ab2f-4c2b-8753-a8a7af1b99b5 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:45.633252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:45.633252Z digest=sha256:905af0c92ec6dc4831c34e441ed5e500884d936c8641e5622d427523aa8576f5

Observation b04e2f4d-a1b8-467c-8b58-6f70f84d6524 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Reasoning with Exploration: An Entropy Perspective

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:45.833877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:45.833877Z digest=sha256:876ee4ae135df59c8267ee1d3c870730c9b5c92067f2dcc868b3db27a9012415

Observation 91e9def5-e837-48b2-b972-09dd171b7349 · outbound

This paper cites MathPile: A Billion-Token-Scale Pretraining Corpus for Math.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:47.613196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:47.613196Z digest=sha256:5d5602d9cc48979e80c8950f4479a8a2f3823e341c0c0334dd0321bf5af77b05

Observation 06f84943-4e26-4301-82be-f15ef631ab84 · outbound

This paper cites The Llama 3 Herd of Models.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.134022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.134022Z digest=sha256:67ed4e7b89987e98b9de36ab48981859a3568326787c3d0ac59496c8e3ff5f10

Observation 5f09d8f1-7c70-4020-8387-322b242215ba · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.255880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.255880Z digest=sha256:1f30b0c99b06c879323b65517a11f3c6bbbc5ee15ba599d268a262b1d1f4285f

Observation 98de604d-3fd3-4c47-9df0-62ca1eb1dfdc · outbound

This paper cites Evaluating Large Language Models Trained on Code.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason Evaluating Large Language Models Trained on Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:45.719770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:45.719770Z digest=sha256:97510e796e790688cb95f94aa4bdb06545c6a2d57a76c4b82ee8752da20cde1e

Pith citing papers

Observation 63ed6416-f211-42d3-8ae8-6792164f8643 · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.631113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.631113Z digest=sha256:0d92f66711e08cd121e0869d5d67b369fce9df9028c99fb7424dbe694f944133

Observation e5a95a1a-5436-40a1-9875-db81b3a9ba99 · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:55:42.926347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:58:38.234197Z digest=sha256:ffd3629161a6bde9615d3afb755a2405ac328725612aaf3d2b82b312e835dff5

Observation da54ada4-7c1f-4b5b-8944-c9de425040ab · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:54.508842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:09:39.645781Z digest=sha256:f4c69f8505c09346355de0b3ab1077b5e816493e565ceaf7c341d4dd16963ebd

Observation da2a00c6-4e88-4cb1-998e-660119840990 · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:32:41.453150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:32:34.585808Z digest=sha256:c8726338317019c162b1caf398e4cf09ad538321480b22ae372a6cc2d3d71acb

Observation 50102c52-6333-46ce-867d-49c934cbcc78 · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.670232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:20867b40a7d54881d0b556ced4a002b21b5ad03cbd27148c1b0656fa319bfdc2

Observation 48091707-a566-4411-ab25-00bddaffb930 · inbound

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards cites this paper.

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:29:40.154298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T05:24:46.545570Z digest=sha256:8d98adee32483fcd4d59c3e207cf4cee7704133c72f3d7963f9d906f00a6c6da

Observation 604f0662-faf6-49e7-87a6-41b31def1ce5 · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.335384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:4d1566f96ca10ce577cec1a441d7d72c3253233241d1818beb5c1ba2e6ffd4f4

Observation 102dbb83-68ec-464a-a929-a3df65e0e30e · inbound

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning cites this paper.

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.912680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T06:30:55.592334Z digest=sha256:2eeeef54ddc6105c5e67cb6947172a77cf434e885dab95c965f29bce3855066a

Observation 5093c845-1be9-4bd5-a621-1039b93d1e57 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.192377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:9e9c4f96ae3113aa37b875ad6e5d74bbf29d0128e87083292526ab34cb89572b

Observation 49501d2a-a3eb-4771-bbb5-a115b9e2e2b4 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.076804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:d7cabdeac7e5b017e64fbf2c190e7bb90795a2ad7cac38af818ccde0baf608ab

Observation e162964d-6e0c-4a53-96a2-b7f75a8d13c6 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.539506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:829b3b232a86bc1f1900d0f9fc0f49cb82d717fbe2115c0e02c951917c6d5863

Observation 7b0118de-0224-4485-af04-c1dad64f514f · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.127455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.127455Z digest=sha256:b0661fcbfab53243c0c6b540434bbea752a9ed84b17c9e1a56079d5cac4c2692