Pith. sign in

REVIEW 10 cited by

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.22257 v2 pith:BAOAZKMS submitted 2025-05-28 cs.LG stat.ML

classification cs.LGstat.ML
keywords off-policygrpoon-policyoptimizationpolicygroupobjectivesrecent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Optimization (PPO), which improves training stability, sampling efficiency, and memory usage. In addition, a recent analysis of GRPO suggests that estimating the advantage function with off-policy samples could be beneficial. Building on these observations, we adapt GRPO to the off-policy setting. We show that both on-policy and off-policy GRPO objectives yield an improvement in the reward. This result motivates the use of clipped surrogate objectives in the off-policy version of GRPO. We then compare the empirical performance of reinforcement learning with verifiable rewards in post-training using both GRPO variants. Our results show that off-policy GRPO either significantly outperforms or performs on par with its on-policy counterpart.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya

    cs.AI 2026-02 conditional novelty 7.0 of 10

    Fine-tuning LLMs on Navya-Nyaya's six-phase reasoning structure yields 100% semantic correctness on held-out logical problems despite only 40% strict format adherence.

  2. REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Using only the replay buffers from RL training of domain-expert LLMs, REGEN trains a multi-domain generalist via offline RL and matches online multi-teacher distillation accuracy at a fraction of the reported training cost.

  3. Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Intent Projection decomposes literal and pragmatic signals in multimodal memes via orthogonal projection and contrastive objectives in LVLMs, outperforming baselines on six benchmarks especially for high-divergence posts.

  4. Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Smaller models provide temporally correlated policy-level diversity that serves as structured exploration for training larger models in GRPO, yielding accuracy gains such as +8.8% on AIME 24 with reduced compute via t...

  5. From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    A group-revision paradigm for GRPO-based RL fine-tuning of VLMs converts failure responses into improvement signals that refine rewards and advantages, yielding gains on referring segmentation, REC, and counting benchmarks.

  6. Gradient Extrapolation-Based Policy Optimization

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    GXPO approximates longer local lookahead in GRPO training via gradient extrapolation from two optimizer steps using three backward passes total, improving pass@1 accuracy by 1.65-5.00 points over GRPO and delivering u...

  7. Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs

    cs.CL 2026-01 conditional novelty 6.0 of 10

    SOUP mixes off-policy historical prefixes with on-policy continuations at token level and reports small but consistent math-reasoning gains over on-policy GRPO/DAPO baselines in selected configurations.

  8. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.

  9. PubSwap: Public-Data Off-Policy Coordination for Federated RLVR

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    PubSwap uses a small public dataset for selective off-policy response swapping in federated RLVR to improve coordination and performance over standard baselines on math and medical reasoning tasks.

  10. POPI: Personalizing LLMs via Optimized Natural Language Preference Inference

    cs.CL 2025-10 unverdicted novelty 5.0 of 10

    POPI distills user preferences into reusable natural-language summaries via a shared inference model and conditions a generator on them, trained jointly with RL to improve personalization quality while cutting context...

Pith tools