Pith. sign in

Self-guided process reward optimization with redefined step-wise advantage for process reinforcement learning

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

years

2026 5 2025 3

roles

background 2

polarities

background 2

representative citing papers

Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models

cs.AI · 2026-01-08 · conditional · novelty 7.0

Miner uses intrinsic policy uncertainty with token-level focal credit assignment and adaptive advantage calibration as a self-supervised reward to enable efficient RL training on positive homogeneous prompts, yielding up to 4.58 Pass@1 gains over GRPO on Qwen3 models.

CARL: Criticality-Aware Agentic Reinforcement Learning

cs.LG · 2025-12-04 · unverdicted · novelty 6.0

CARL improves long-horizon agentic RL by using entropy as a proxy to identify critical states and selectively updating only on high-criticality actions, yielding stronger performance and efficiency.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

citing papers explorer

Showing 8 of 8 citing papers.