Pith. sign in

Paper Citation Record · LEDGER

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.07674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07674 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:20:40.229376Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact21
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7feb0db6-746c-465b-8429-ab56ebe1bf9e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.826547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:83cdde2b5f291551027278b22111d601dded2403dc305b39dd0e1fbe9ba282b5

Observation 549f727b-acea-40ef-95ad-c461e05d5f1d · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.802842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:0b43470ed2d0886e0d19d60fbe714c5169c4a4c9824cb897b1aa4174b7ce7eb1

Observation a99ee716-3ffe-489d-8b5b-29fa8a1a2dde · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.834564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:0bc8bb396c9309c8d232c8a139328e3c0e84765f2fd7fe9dbd80d3d7930b825d

Observation 5563cdc8-2fe4-4eea-9619-a90c0e6f1c8d · outbound

This paper cites Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.824146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:ea6019fccb270fa287cc186997a16aed0979dca3df5d9f04ec5d6ee431e44cc8

Observation af4623f5-24c8-43d1-b05b-d780020a0754 · outbound

This paper cites Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.791721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:7842b4c4362b7f443a3f92fe5c7145d71a1a6d6a00df92fe83e6b271d7db18b3

Observation cb77a87f-83cd-4b53-ac29-ce2a0c3b81e7 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.810399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:e8b538c699194f76c3916cd15e1831f4feac29ee80921ca41853102926166835

Observation 20b9b9b3-b7dd-4fc8-a989-404cbf300221 · outbound

This paper cites arXiv preprint arXiv:2510.01135 , year=.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems arXiv preprint arXiv:2510.01135 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.829414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:96482b950c82c5f67912806ea860c9873154ffcc69cda38e93b1c388a6b1075d

Observation 22afc2ab-00e6-4edd-aac9-940b8f276cd0 · outbound

This paper cites Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.816174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:f9fef2f63bba538e95a2a9d2bd22ad38c3cc887203f50b4960d6ba32699ef519

Observation 6de9f22c-4787-4d37-8836-fd4369b43e5b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Proximal Policy Optimization Algorithms

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.808011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:02f1be61bd6a63b2f52f55d0cb298415401249454803c898d9d293c559de5209

Observation a7fd83c7-f141-4b77-89ff-788be4ac41d2 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.836976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:aaf90e4abd6054346aa868da3ec2a30c078d24999517711e9c6f6a5c449c20e3

Observation cf77be72-9782-4428-ab86-11bc600eb237 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.821678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:02c313ca22c7bd869d2427f4ad29a64ae12d1ac9732bb3e354fa2544cbf030a0

Observation 06adb3ca-45e4-45eb-92a9-a4eb7e0b4b1e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.831866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:d28241c1d4146b8fdf1a7b82bdb45b0ef51ec243d94036bcebeb6baec1d16d61

Observation 9d2b0d31-fa6a-4f94-a299-5a38af249d85 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.839356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:ccb6fa003e8812e51d711b96f5257f1cd7689fd06aef1c961b93745dc9bbff7c

Observation 4a1dc497-6031-4de0-87ed-b1c7a839baa1 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.797458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:e0b0ae46ee0dc259d6390aea12999253e289b9c49ed748539242017b25989039

Observation 0da35ac9-16c6-4de0-8812-3750e063f7ba · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning to Reason under Off-Policy Guidance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.794724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:7aa503fb1e313c65031baacdfdd57f2ffa5b029f90b5699059674d5d4fab3087

Observation 737d1113-7763-432b-94e5-7e95e7e972a2 · outbound

This paper cites Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.813418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:d9a5af1b45ac4466b82e04ae5b0556acb8e995478aae3e51c430ad47ebe41627

Observation b7ccd8a8-742b-4b69-9b5e-a9c96dc71207 · outbound

This paper cites Qwen3 Technical Report.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Qwen3 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.819124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:ba378c227c88e3310dcb7b42230eedb3bac0e5f0ed2d277f84ae02a29c2f8cfe

Observation 5004e8d1-2b04-470e-9391-df500e37ae63 · outbound

This paper cites Gemma 4 Technical Report.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Gemma 4 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.800044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:dc13793e14fb63a3f45f05a12d75555254b4e9d2910132fb25743ff7044e6549

Observation 2dea5867-8b5f-49dc-9414-306428d24929 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.841864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:775879adb082622c8c128fc159ae66d1736b8183d4c09c84688bd8227631a14c

Observation 0c14db54-8009-4346-b6c2-c90fc5dd85db · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.805458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:a07dc964917b110c9f7debdb8f8a34c4077c1e5612c2b298b47649923db48508

Observation 1f612d8e-6e7c-4fc5-a97c-28a4d692d36f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Verifiers to Solve Math Word Problems

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.788893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:8ceb5162a39df6c3bff89932fe43f48281ff7b0ddb09c4471ab1a2110ba3f8d0

Pith citing papers

No inbound Pith citation observations are available.