Pith. sign in

Paper Citation Record · LEDGER

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

As of 21 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2607.19313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19313 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:36.792201Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76a670ab-125a-49ab-bcf2-60f6d7ef615a · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.027123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.027123Z digest=sha256:ad465a924718bbf251db2ec7df3a9baad43898a21d4e3d74db555f0c235d4d40

Observation 99017e16-c354-4bea-aba2-d2e838c5bc7e · outbound

This paper cites Proceedings of the nineteenth international conference on machine learning , pages=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the nineteenth international conference on machine learning , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.211148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.211148Z digest=sha256:0e709124e3649dacc5c3e0bb170daa270479c9bed4145ea3c54b1886025be2d1

Observation 8705f917-7223-4f55-a128-6ad9eab14077 · outbound

This paper cites ICML , pages=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information ICML , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.413310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.413310Z digest=sha256:01dbe791c83379cb2a0fbe4b587c1a1ea1f61e5c2d5319d74e98f09db24f0a3b

Observation 61627157-aa84-4dae-a66a-b432b96878ca · outbound

This paper cites an unresolved cited work.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.593189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.593189Z digest=sha256:0ae754781ac192d4d863734563a84a31c336fd73f3807ac2f6ba70bc24110802

Observation 93f9204f-7363-468f-a164-9954ffa6dd87 · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.764992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.764992Z digest=sha256:3600e0f08afdc592f8692e378a72cc779d7006c03c3158bd88c661b12d2858ef

Observation b274825a-23f3-474c-ae78-8246a3dc99f5 · outbound

This paper cites 2024 , journal =.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information 2024 , journal =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:25.910824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:25.910824Z digest=sha256:57c1d8a76247ed6a63f8e9a8c8ac7e637f1c45ecfebc82e58a2710a15c09ba4a

Observation 8a43daa9-f572-4af5-8892-15caad909072 · outbound

This paper cites Qwen2.5: A Party of Foundation Models , url =.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Qwen2.5: A Party of Foundation Models , url =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:26.110796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:26.110796Z digest=sha256:8cf364e549d22d4301332dec7539cb417aece76d4bc06db2974ccd15c56ae117

Observation fe26e55c-788f-49e7-b414-56eaf8a2b6c7 · outbound

This paper cites 2023 , publisher =.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information 2023 , publisher =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:26.237387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:26.237387Z digest=sha256:3df47384e8201cb7981c9d1ad08c8a99d7d300c68f541c7b695c287722ce9e4f

Observation 5643572d-cba5-448b-bf0a-f02b61bfcf04 · outbound

This paper cites an unresolved cited work.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:26.374603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:26.374603Z digest=sha256:78bdb8ceaf5126ad07031dbdb6d974f36bf82e0cbbe39a90a610bc1aa93f58bf

Observation f5ac987c-45d7-4838-ace6-36278d191624 · outbound

This paper cites NeurIPS , year=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information NeurIPS , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:26.821313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:26.821313Z digest=sha256:dc6c6821a3c777afb93536156439d6a6f15e61e6e2f99b65b132ee3094575a63

Observation c7367ab7-f7a0-4192-bbac-2a4db28eda96 · outbound

This paper cites GRPO is Secretly a Process Reward Model.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information GRPO is Secretly a Process Reward Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:27.252303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:27.252303Z digest=sha256:da41d4b773bd7028d2baa1bc25caf0dfa0631d84c8fce122209a2f132fb6cf7b

Observation 9d3459f2-de37-42f7-bf7f-244c99a16118 · outbound

This paper cites arXiv preprint arXiv:2510.02263 , year=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information arXiv preprint arXiv:2510.02263 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:27.376399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:27.376399Z digest=sha256:ac982d0f81447c59168ada09a26c48beadbc27f27c7c697e0cb18a37f3b36bae

Observation 4e78c49e-2660-4181-98bc-872084b33d59 · outbound

This paper cites Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:27.503850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:27.503850Z digest=sha256:695f8715e2c0ac67347f4a88219ce12fd2a5778ec1c7ba7ef45dda65fb111e46

Observation f8a5b6de-1ab0-41b3-90ec-3cb6462a186e · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:27.855935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:27.855935Z digest=sha256:a325b23b54b9682bbc40de54c8f03c82c0391d2353b3663e5fc496dbcbf572c7

Observation 187488eb-2ccf-491e-a63c-b61a49664b14 · outbound

This paper cites International Conference on Machine Learning , pages=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information International Conference on Machine Learning , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:28.288722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:28.288722Z digest=sha256:dc10afb91912331d45c8cbdad88759dc1c41e22e9a6c4688449d1b6e682f8e07

Observation a8684b0c-36a8-467d-95c0-6630ede22555 · outbound

This paper cites Advances in neural information processing systems , volume=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:28.860846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:28.860846Z digest=sha256:ee1c8e41239709f2f8d4d882cebd3c224fab7d88642396ef7e97716e88f3fa65

Observation 81fd7a86-975a-4723-baaf-c577a25adda2 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:29.129869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:29.129869Z digest=sha256:6a95db4760ec7538ad41d3c6a8d1080c31d561985e33d77aea34aed19848733f

Observation 96c3a91c-0595-4ec2-a3d4-7446c5a4d6be · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Advances in Neural Information Processing Systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:29.921610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:29.921610Z digest=sha256:0c935fb6beae144f73716dab0691d6c88c779cc5fdaa58ce9c4283c7dbbcffa7

Observation 9e2dd0e6-3e08-434e-9947-053483f120ad · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:30.683113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:30.683113Z digest=sha256:0719f0ae3c5a5e8a091d3ca70e97bf2d6b5132138ea661a350ddbfdedffabc89

Observation e0b33592-04e5-4541-9343-2464ea75f564 · outbound

This paper cites The twelfth international conference on learning representations , year=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information The twelfth international conference on learning representations , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:30.803160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:30.803160Z digest=sha256:d860170e14b8e19fb676b86694033d70edb2499f3df4aef562d5a77340d73098

Observation 847d2d67-8a82-4edc-95fd-9fa9ada2624a · outbound

This paper cites International conference on machine learning , pages=.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information International conference on machine learning , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:30.919231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:30.919231Z digest=sha256:34f27e34dd610881e7102d363be97f3f356ab195ed81be5c751fe4b4ffa51723

Observation 9ca3c92c-467d-41f1-a650-3508eceb818a · outbound

This paper cites Problems and Projects , url =.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Problems and Projects , url =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:31.588781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:31.588781Z digest=sha256:8b898d063944c98321253df10eef7ccc9aff51db509664e5f6271a9e638b0a3c

Observation ecff8e9a-7084-4922-a063-5c192f79c509 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information On-policy distillation of language models: Learning from self-generated mistakes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:31.783694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:31.783694Z digest=sha256:054416220b536ac95d412e0472dab1c11452baba2aef8e7ba58a3f882ff11941

Observation 752e4cd9-36ac-4349-9f16-d28225d4274a · outbound

This paper cites Rl for reasoning by adaptively revealing rationales.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Rl for reasoning by adaptively revealing rationales

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:31.938291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:31.938291Z digest=sha256:c9fe47a1cad94f5c4b27861cf173a811c2aae12f064b3d8107d311d720f1b79c

Observation 6b0c8ca3-0dcb-4f57-8e20-c189d09fa0e9 · outbound

This paper cites Hindsight experience replay.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Hindsight experience replay

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.110014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.110014Z digest=sha256:f70103a9f6e40eb31dd09673b2864c50b3487ece0cccca655d97056073f5ae6f

Observation 86a1a692-43c6-4559-b87e-27e4e2779401 · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.206324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.206324Z digest=sha256:761a38f7e988c462e1340380f327791ddc1f0026ce22eb35195c75e4bb00b2a0

Observation 4454bc61-3edf-4c3e-bebb-180c771a479e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating Large Language Models Trained on Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.306487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.306487Z digest=sha256:a7b3a5c058afb306dfd17b3584daf35b644ada5d2f86ec86cb575ed3a31c75ba

Observation 7137d105-8ddd-4a1f-a3a6-c8e245cb7a38 · outbound

This paper cites Self-evolving curriculum for llm reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Self-evolving curriculum for llm reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.380382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.380382Z digest=sha256:a54a9f058a93cb0d1af09c2f04c00104754c5991ec3d314204fca47cc213532c

Observation 419b9932-4c23-4d10-ab26-9d6f049cfd37 · outbound

This paper cites Off-Policy Actor-Critic.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Off-Policy Actor-Critic

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.478278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.478278Z digest=sha256:25a2dd9dd4178b141d26fb4ec4f1a172536b0c3eed45c1fe682a64ec47d3fea3

Observation bcf129b3-9172-4710-b54d-160be1d82385 · outbound

This paper cites AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.568310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.568310Z digest=sha256:0b8eec320109ff87ad0d9b750b1462c8a8405bb355369ff18d6a1ded11efe7ef

Observation 95d6c613-43c4-4846-8694-4ff07d2545a9 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.701608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.701608Z digest=sha256:67193d18e719eb3ab916350fb98ce0ddb2e72c1595de40bf54e013c1bcd4781f

Observation 5610a2b9-97f2-43c0-a751-1c31f3f67122 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.776683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.776683Z digest=sha256:81683b573d8b8019e5b974b9204333f5a11fec6fb245bb1bc5ea88c1f900e2fc

Observation a7c748f7-a37a-483c-9dcc-c030f20fa417 · outbound

This paper cites G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.886708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.886708Z digest=sha256:8a56d36ac759931742d8773fb4c23f4f4c85ecf044c1347b1a8ac22f6d2854bf

Observation 03087007-6f49-4b36-8a16-2aeaae6281b1 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Measuring mathematical problem solving with the math dataset

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.002235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.002235Z digest=sha256:0225d8364f7e393ed2bdacc9628a287530aa96e0bac30b89fcbbba0cf08a8ac9

Observation d55ebfe2-6231-45b1-b371-fc8b049906a1 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Lo RA : Low-rank adaptation of large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.124762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.124762Z digest=sha256:76ccbdc685d4f62f57052c1374874a09ca234c1552aa0248a69ab318dc59d5b5

Observation 05672ce9-64c7-4aa3-a746-8cb9d08e8e8f · outbound

This paper cites Sample-efficient online learning in lm agents via hindsight trajectory rewriting.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Sample-efficient online learning in lm agents via hindsight trajectory rewriting

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.248959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.248959Z digest=sha256:f448c06bb45ec404d3b5ec611fe2851d709f611f88026d27cd1962ab32b7bfd3

Observation bf029e8a-32fa-4a6f-af6e-c50d37b0bcff · outbound

This paper cites Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.390704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.390704Z digest=sha256:fccf630f3ee02097636b28d6fa032a191209c1d84531095bd2dbb5f211e2800b

Observation 4b5040be-b6de-47f0-9fa6-ba73798a1b03 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Reinforcement Learning via Self-Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.512106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.512106Z digest=sha256:20e23bfac17b2ce66f77a0c5b023822542107e0b98a2589eac83e316d5528e0d

Observation 39bbffcd-d84d-4a37-a280-23fd6827cea0 · outbound

This paper cites OpenAI o1 System Card.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information OpenAI o1 System Card

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.573130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.573130Z digest=sha256:e82f207e39231eb21949be5ff67f6fa9174950db259cde7a16afee2c6afb894b

Observation bee3af15-8551-43dd-89d0-104a085bbd3e · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.680429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.680429Z digest=sha256:d13effd38e8b0c9db1dd6f683186273cbfce11a8cfc2e887f45083857de4f527

Observation a47053d8-2de4-4222-833d-7ec867bde426 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Gonzalez, Hao Zhang, and Ion Stoica

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.776828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.776828Z digest=sha256:b6bd5a8bdda124e68a2a6e979cbf851dd1b9ec84db57ef85d65f000d3d0c5296

Observation b558264b-4770-498c-b288-d069b5ff1e56 · outbound

This paper cites Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.849756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.849756Z digest=sha256:d8df0b6ce08156f80a93dd8001a96e735e2ba15f03a5afc2be9b77ae7c75e97f

Observation 5c8df385-a953-4309-8779-b0c3f383755d · outbound

This paper cites Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.943411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.943411Z digest=sha256:6fd3a32458a869e03dbb31c9d3aa2c83c34a327fe0a5debc08b6c0d5137ecebc

Observation 6ae50384-9417-481b-b292-f156f396d70f · outbound

This paper cites Off-policy temporal-difference learning with function approximation.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Off-policy temporal-difference learning with function approximation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.055058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.055058Z digest=sha256:ec6dc19865c3b0f80d3e4d2d7e12274baf0e2e79709c937f7f66fccf1cf45f9c

Observation a60edbbf-4f4d-4b4e-934e-62116143940b · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.154806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.154806Z digest=sha256:2c8b428898d2eb293d4d804178c2d688a691511171b795bc964ce0f3b0c1d84a

Observation a631dee2-f408-4aa5-a223-957af15975d8 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Qwen2.5: A party of foundation models, September 2024

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.330274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.330274Z digest=sha256:6d71024a31018a693b439a38d1145663c88b7f31c772de5a9e1ea1e6b38b07d3

Observation 7bb29fe9-27b0-47a6-9c07-b36ce72614eb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Proximal Policy Optimization Algorithms

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.459150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.459150Z digest=sha256:ae5f7e7b5ad48eb67575be06a1331b86e97522258c97b2ceecad2962ae5262b6

Observation c1ab3a12-d993-4405-aab7-671825443b07 · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.538083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.538083Z digest=sha256:c26fb2cba18b28253b2d1f0a83fc9195cd85e296cf27822af4178613b572dfe6

Observation 612771ae-c2fb-4ddf-89c9-2d035ecfaa18 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.607496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.607496Z digest=sha256:82ff4b3dc37aa8982241b524fcee1c296d7727d001423a5c5d9022a27177513b

Observation 375c48a9-cd49-42aa-955b-46bf68ffb8e9 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information HybridFlow: A Flexible and Efficient RLHF Framework

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.702372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.702372Z digest=sha256:58bbfc2a0210ecaf998c1365cc8dbe733f2f0a220855f38d2f2f7e9d5a16e518

Observation 8601a47d-6b23-4ca9-8ba9-1fc876d16795 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.843879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.843879Z digest=sha256:296527ea110d33fc1eb7460ae58efdc70a50c51aa296a7141c27562d971b6535

Observation b1abed22-93a8-4a9b-ad17-f379c24fafb9 · outbound

This paper cites Trial and error: Exploration-based trajectory optimization of llm agents.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Trial and error: Exploration-based trajectory optimization of llm agents

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.973667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.973667Z digest=sha256:f50ea226ce2b17aa2b197e7b9c16f199f111b724797e975ba2b947ed44b3bd24

Observation c0f5f7ff-1ce5-4d48-948b-0c8be5d470a1 · outbound

This paper cites Gates: Self-distillation under privileged context with consensus gating.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Gates: Self-distillation under privileged context with consensus gating

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.136503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.136503Z digest=sha256:b6da23a79042a73d0c36dd39e14054fdfde0e99f7dde90e213f4b97a7e018ec4

Observation 75c7df61-efa2-4389-9325-e52819274aed · outbound

This paper cites AIME problem set 1983--2024.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information AIME problem set 1983--2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.300797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.300797Z digest=sha256:adbc8d3f5c3803fc58c7a31149484d507648ef7ec20ead636a15bc8f313e0fc0

Observation d8d44f30-1f9f-4cb2-a9ac-d71eda467f99 · outbound

This paper cites Dump: Automated distribution-level curriculum learning for rl-based llm post-training.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Dump: Automated distribution-level curriculum learning for rl-based llm post-training

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.433069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.433069Z digest=sha256:d78c2232db83811ae54878e28e4092dae44dee8046e9bb60b2b5ac01615cb6e2

Observation 4b78cf20-0067-4eaa-a8e2-c7b386ba9419 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Learning to Reason under Off-Policy Guidance

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.562246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.562246Z digest=sha256:230c95410a6ff33dc68da4a4d5fde93a2f1ba744c0bebc6750363e2205cb70d7

Observation c985210b-79e4-4f8f-bcfc-52f65016847d · outbound

This paper cites Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.696337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.696337Z digest=sha256:142ee6a088447c07fa99c6cc3de447bb2b36290b26fa6731ad47f490e218b061

Observation 7ea2e8f1-d79f-4a16-92a5-ca54c10c1d7e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.828678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.828678Z digest=sha256:1fa163105162b1bff7ec3ac6832b8d123b92b50c34df45d28b9001aaa7da932e

Observation 164d67d3-89e1-49e6-913c-00530ca1fb4d · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Star: Bootstrapping reasoning with reasoning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:35.977861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:35.977861Z digest=sha256:94423edcfcb9c6e6517bad5420b530e9c7a8990205d316dc04bca1cb3d7114a9

Observation 18a63bc5-dc4d-451b-8b07-b4d204c121f9 · outbound

This paper cites Exgrpo: Learning to reason from experience.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Exgrpo: Learning to reason from experience

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.041500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.041500Z digest=sha256:c7d739ccfae1d259d422e29e940681d1cf40c77bee7ab699c899607a377ad90b

Observation 7b0118de-0224-4485-af04-c1dad64f514f · outbound

This paper cites StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.127455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.127455Z digest=sha256:d3783fbc0a5e8c779aa0be26672db33cd03864b846c1c76d69f3fe9329d201e1

Observation af17d960-ca6b-4a1c-9b65-28ccca6f4bf8 · outbound

This paper cites CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.252133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.252133Z digest=sha256:e542888ec270a7707d3f7242666198b872e359c58374c6b18e135fcd950b07c2

Observation 2f33e5cc-4fe0-474a-8b59-c835c36bf40e · outbound

This paper cites On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.359853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.359853Z digest=sha256:6b7984a75d187928ddbafe94efdd119f6628eec3d100d17cab38dd20c6902c69

Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · outbound

This paper cites Evaluating the Performance of Large Language Models on GAOKAO Benchmark.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.403464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.403464Z digest=sha256:2ab2a8793d0cade09682978eb03aad721399b65c0fa473b78222dd45bc108569

Observation 4a40d002-ff71-456a-9dc0-5938a947c162 · outbound

This paper cites Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.575815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.575815Z digest=sha256:95d8758813d1c5401203b553f8228b6de306cc269ffa8f633bbe9eefe91eb97d

Observation 4499d63a-74c1-46f9-9299-604148447797 · outbound

This paper cites BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.704591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.704591Z digest=sha256:2ac218c6946df9764745dc011bd45719479145e168f65786c9ec7f998a5a8fbb

Observation fc720a5a-df46-463f-a6d2-210078bbd77b · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.792201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.792201Z digest=sha256:0a78480616da17ba67ed447a5fa8ec69bf5bb1bc985972b668dfe7aa358ca52c

Pith citing papers

No inbound Pith citation observations are available.