Pith. sign in

Paper Citation Record · LEDGER

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.07674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07674 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:20:40.229376Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact21
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7feb0db6-746c-465b-8429-ab56ebe1bf9e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.826547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:7e01be8efe22c8d589963a6240f46a328ec05d26c39ca800d00bc7280a6bff27

Observation 549f727b-acea-40ef-95ad-c461e05d5f1d · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.802842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:b22df949311085f516e2e20c013a171afe3f13f6460a5b32a1e34f6dd60c8237

Observation a99ee716-3ffe-489d-8b5b-29fa8a1a2dde · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.834564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:e8b1ff3a6368abf1564fd40e01f8691e7c4eda81e89a9049c7d275389a045a30

Observation 5563cdc8-2fe4-4eea-9619-a90c0e6f1c8d · outbound

This paper cites Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.824146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:246e677fa4f3700498d2f40cde4cd74bb314785ac9101a217d622c9b79f1c691

Observation af4623f5-24c8-43d1-b05b-d780020a0754 · outbound

This paper cites Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.791721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:84b780f1e95bb3f5f092b10eb76b81f92628af7e4cb1add4b81387a4848d9d47

Observation cb77a87f-83cd-4b53-ac29-ce2a0c3b81e7 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.810399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:5a391577dafa55bff0e079aacacb3587bb4fb4c44ded71df22dd9af1e12e2cff

Observation 20b9b9b3-b7dd-4fc8-a989-404cbf300221 · outbound

This paper cites arXiv preprint arXiv:2510.01135 , year=.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems arXiv preprint arXiv:2510.01135 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.829414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:e015527fe5b247cf9f415f3e6d7e0a31bc93906b8e80145159d6fbe561476603

Observation 22afc2ab-00e6-4edd-aac9-940b8f276cd0 · outbound

This paper cites Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.816174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:bbd14674125e90b2b3e4afa79e0c84996e9ffdc5402b69f51cb736e70e9dc0ce

Observation 6de9f22c-4787-4d37-8836-fd4369b43e5b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Proximal Policy Optimization Algorithms

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.808011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:94af56033893ecb2da86d6a2c9d9b200c0ccc1ebaed71afcc11b49bf6a12244a

Observation a7fd83c7-f141-4b77-89ff-788be4ac41d2 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.836976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:fa341ff5cee5ddb46711135801ef1c47a9589aaa1b06ad43c16a1769c2e4b093

Observation cf77be72-9782-4428-ab86-11bc600eb237 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.821678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:5d7bd9c9e235b97cbc6f6f0e0b877e465f4f01d7ed2279b255f4e4aae3bd98cd

Observation 06adb3ca-45e4-45eb-92a9-a4eb7e0b4b1e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.831866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:0a7cf9707174cb12f315f653defe21b846da040de8c3328a5380dcb68d9f1316

Observation 9d2b0d31-fa6a-4f94-a299-5a38af249d85 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.839356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:7d31263c154e2a094bdc53466897279723ba96abeeebf8c9d950649fb60910b1

Observation 4a1dc497-6031-4de0-87ed-b1c7a839baa1 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.797458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:5e28c95aaef93dcb47bdc31b549d299e9ba772d47b770c53902e02173a89ba40

Observation 0da35ac9-16c6-4de0-8812-3750e063f7ba · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning to Reason under Off-Policy Guidance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.794724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:5095a77cd06f193858a82a1495051969726caade153488c1635f99b8b9ecfbb2

Observation 737d1113-7763-432b-94e5-7e95e7e972a2 · outbound

This paper cites Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:25:57.813418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:557af2513ff6bd808dd7d6af8646d0b0677e8326ed8fb4dad2abb42c0fb8acf0

Observation b7ccd8a8-742b-4b69-9b5e-a9c96dc71207 · outbound

This paper cites Qwen3 Technical Report.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Qwen3 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.819124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:57265fe9ad1157342a347ef0e5c067a7f1e9713b67cb9162af5166e5735fc183

Observation 5004e8d1-2b04-470e-9391-df500e37ae63 · outbound

This paper cites Gemma 4 Technical Report.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Gemma 4 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.800044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:7c6714b8a9f3b85305b55b7add11afe4c24d21451e62e21e01990282c8f7b625

Observation 2dea5867-8b5f-49dc-9414-306428d24929 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.841864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:92ae96dda6a330869f31a26a9c98f9bb2166c948fe40fdcd6753302becb36bbc

Observation 0c14db54-8009-4346-b6c2-c90fc5dd85db · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.805458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:988af6ea24e1d7a9e992576579f740457dfa336780f65767b2e9eee07f30652b

Observation 1f612d8e-6e7c-4fc5-a97c-28a4d692d36f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Verifiers to Solve Math Word Problems

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.788893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:ff50536ef6c404d53b1ab8c4f1383f6e2a2c3aa099001a4d3c01d95b59ac7aa7

Pith citing papers

No inbound Pith citation observations are available.