Pith. sign in

arXiv preprint arXiv:2505.16265 , year =

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

years

2026 2 2025 3

roles

background 2

polarities

background 2

representative citing papers

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

VRPRM: Process Reward Modeling via Visual Reasoning

cs.LG · 2025-08-05 · conditional · novelty 5.0

VRPRM combines 3.6K CoT-PRM SFT data with 50K non-CoT PRM RL data to train a visual PRM that beats a 400K-data non-thinking PRM and boosts best-of-N accuracy.

citing papers explorer

Showing 5 of 5 citing papers.

  • Leveraging Verifier-Based Reinforcement Learning in Image Editing cs.CV · 2026-04-30 · unverdicted · none · ref 26 · 2 links

    Edit-R1 builds a CoT-based reasoning reward model (RRM) via SFT and GCPO, then applies it with GRPO to improve image editing models such as FLUX.1-kontext.

  • The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cs.AI · 2025-09-02 · accept · none · ref 266

    Survey that defines agentic RL for LLMs via POMDPs, introduces a taxonomy of planning/tool-use/memory/reasoning capabilities and domains, and compiles open environments from over 500 papers.

  • Trust Region On-Policy Distillation cs.LG · 2026-05-31 · unverdicted · none · ref 153

    TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

  • VRPRM: Process Reward Modeling via Visual Reasoning cs.LG · 2025-08-05 · conditional · none · ref 4

    VRPRM combines 3.6K CoT-PRM SFT data with 50K non-CoT PRM RL data to train a visual PRM that beats a 400K-data non-thinking PRM and boosts best-of-N accuracy.

  • A Survey of Reinforcement Learning for Large Reasoning Models cs.CL · 2025-09-10 · accept · none · ref 196

    A survey compiling RL methods, challenges, data resources, and applications for enhancing reasoning in large language models and large reasoning models since DeepSeek-R1.