Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:36.792201Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2607.19313.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:36.792201Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 76a670ab-125a-49ab-bcf2-60f6d7ef615a · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99017e16-c354-4bea-aba2-d2e838c5bc7e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the nineteenth international conference on machine learning , pages=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8705f917-7223-4f55-a128-6ad9eab14077 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information ICML , pages=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61627157-aa84-4dae-a66a-b432b96878ca · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f9204f-7363-468f-a164-9954ffa6dd87 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b274825a-23f3-474c-ae78-8246a3dc99f5 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information 2024 , journal =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a43daa9-f572-4af5-8892-15caad909072 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Qwen2.5: A Party of Foundation Models , url =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe26e55c-788f-49e7-b414-56eaf8a2b6c7 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information 2023 , publisher =
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5643572d-cba5-448b-bf0a-f02b61bfcf04 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ac987c-45d7-4838-ace6-36278d191624 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information NeurIPS , year=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7367ab7-f7a0-4192-bbac-2a4db28eda96 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information GRPO is Secretly a Process Reward Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d3459f2-de37-42f7-bf7f-244c99a16118 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information arXiv preprint arXiv:2510.02263 , year=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e78c49e-2660-4181-98bc-872084b33d59 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a5b6de-1ab0-41b3-90ec-3cb6462a186e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187488eb-2ccf-491e-a63c-b61a49664b14 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information International Conference on Machine Learning , pages=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8684b0c-36a8-467d-95c0-6630ede22555 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Advances in neural information processing systems , volume=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81fd7a86-975a-4723-baaf-c577a25adda2 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c3a91c-0595-4ec2-a3d4-7446c5a4d6be · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Advances in Neural Information Processing Systems , volume=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2dd0e6-3e08-434e-9947-053483f120ad · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b33592-04e5-4541-9343-2464ea75f564 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information The twelfth international conference on learning representations , year=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847d2d67-8a82-4edc-95fd-9fa9ada2624a · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information International conference on machine learning , pages=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca3c92c-467d-41f1-a650-3508eceb818a · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Problems and Projects , url =
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecff8e9a-7084-4922-a063-5c192f79c509 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information On-policy distillation of language models: Learning from self-generated mistakes
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752e4cd9-36ac-4349-9f16-d28225d4274a · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Rl for reasoning by adaptively revealing rationales
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0c8ca3-0dcb-4f57-8e20-c189d09fa0e9 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Hindsight experience replay
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a1a692-43c6-4559-b87e-27e4e2779401 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Matharena: Evaluating llms on uncontaminated math competitions, February 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4454bc61-3edf-4c3e-bebb-180c771a479e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating Large Language Models Trained on Code
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7137d105-8ddd-4a1f-a3a6-c8e245cb7a38 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Self-evolving curriculum for llm reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419b9932-4c23-4d10-ab26-9d6f049cfd37 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Off-Policy Actor-Critic
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf129b3-9172-4710-b54d-160be1d82385 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d6c613-43c4-4846-8694-4ff07d2545a9 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5610a2b9-97f2-43c0-a751-1c31f3f67122 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c748f7-a37a-483c-9dcc-c030f20fa417 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03087007-6f49-4b36-8a16-2aeaae6281b1 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Measuring mathematical problem solving with the math dataset
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d55ebfe2-6231-45b1-b371-fc8b049906a1 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Lo RA : Low-rank adaptation of large language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05672ce9-64c7-4aa3-a746-8cb9d08e8e8f · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Sample-efficient online learning in lm agents via hindsight trajectory rewriting
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf029e8a-32fa-4a6f-af6e-c50d37b0bcff · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5040be-b6de-47f0-9fa6-ba73798a1b03 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Reinforcement Learning via Self-Distillation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39bbffcd-d84d-4a37-a280-23fd6827cea0 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information OpenAI o1 System Card
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee3af15-8551-43dd-89d0-104a085bbd3e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Vcrl: Variance-based curriculum reinforcement learning for large language models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47053d8-2de4-4222-833d-7ec867bde426 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Gonzalez, Hao Zhang, and Ion Stoica
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b558264b-4770-498c-b288-d069b5ff1e56 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8df385-a953-4309-8779-b0c3f383755d · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae50384-9417-481b-b292-f156f396d70f · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Off-policy temporal-difference learning with function approximation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60edbbf-4f4d-4b4e-934e-62116143940b · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Pope: Learning to reason on hard problems via privileged on-policy exploration
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a631dee2-f408-4aa5-a223-957af15975d8 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Qwen2.5: A party of foundation models, September 2024
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb29fe9-27b0-47a6-9c07-b36ce72614eb · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proximal Policy Optimization Algorithms
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ab3a12-d993-4405-aab7-671825443b07 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612771ae-c2fb-4ddf-89c9-2d035ecfaa18 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375c48a9-cd49-42aa-955b-46bf68ffb8e9 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information HybridFlow: A Flexible and Efficient RLHF Framework
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8601a47d-6b23-4ca9-8ba9-1fc876d16795 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1abed22-93a8-4a9b-ad17-f379c24fafb9 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Trial and error: Exploration-based trajectory optimization of llm agents
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f5f7ff-1ce5-4d48-948b-0c8be5d470a1 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Gates: Self-distillation under privileged context with consensus gating
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c7df61-efa2-4389-9325-e52819274aed · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information AIME problem set 1983--2024
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d44f30-1f9f-4cb2-a9ac-d71eda467f99 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Dump: Automated distribution-level curriculum learning for rl-based llm post-training
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b78cf20-0067-4eaa-a8e2-c7b386ba9419 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Learning to Reason under Off-Policy Guidance
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c985210b-79e4-4f8f-bcfc-52f65016847d · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea2e8f1-d79f-4a16-92a5-ca54c10c1d7e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 164d67d3-89e1-49e6-913c-00530ca1fb4d · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Star: Bootstrapping reasoning with reasoning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a63bc5-dc4d-451b-8b07-b4d204c121f9 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Exgrpo: Learning to reason from experience
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0118de-0224-4485-af04-c1dad64f514f · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af17d960-ca6b-4a1c-9b65-28ccca6f4bf8 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f33e5cc-4fe0-474a-8b59-c835c36bf40e · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a40d002-ff71-456a-9dc0-5938a947c162 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4499d63a-74c1-46f9-9299-604148447797 · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc720a5a-df46-463f-a6d2-210078bbd77b · outbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.