Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T03:20:40.229376Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.07674.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T03:20:40.229376Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7feb0db6-746c-465b-8429-ab56ebe1bf9e · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 549f727b-acea-40ef-95ad-c461e05d5f1d · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a99ee716-3ffe-489d-8b5b-29fa8a1a2dde · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Pope: Learning to reason on hard problems via privileged on-policy exploration
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5563cdc8-2fe4-4eea-9619-a90c0e6f1c8d · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation af4623f5-24c8-43d1-b05b-d780020a0754 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb77a87f-83cd-4b53-ac29-ce2a0c3b81e7 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20b9b9b3-b7dd-4fc8-a989-404cbf300221 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems arXiv preprint arXiv:2510.01135 , year=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 22afc2ab-00e6-4edd-aac9-940b8f276cd0 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6de9f22c-4787-4d37-8836-fd4369b43e5b · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7fd83c7-f141-4b77-89ff-788be4ac41d2 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cf77be72-9782-4428-ab86-11bc600eb237 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 06adb3ca-45e4-45eb-92a9-a4eb7e0b4b1e · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Understanding R1-Zero-Like Training: A Critical Perspective
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9d2b0d31-fa6a-4f94-a299-5a38af249d85 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4a1dc497-6031-4de0-87ed-b1c7a839baa1 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0da35ac9-16c6-4de0-8812-3750e063f7ba · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning to Reason under Off-Policy Guidance
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 737d1113-7763-432b-94e5-7e95e7e972a2 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b7ccd8a8-742b-4b69-9b5e-a9c96dc71207 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Qwen3 Technical Report
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5004e8d1-2b04-470e-9391-df500e37ae63 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Gemma 4 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2dea5867-8b5f-49dc-9414-306428d24929 · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0c14db54-8009-4346-b6c2-c90fc5dd85db · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f612d8e-6e7c-4fc5-a97c-28a4d692d36f · outbound
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Training Verifiers to Solve Math Word Problems
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.