Pith. sign in

A technical survey of reinforcement learning techniques for large language models

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it

citation-role summary

background 3

citation-polarity summary

years

2026 9 2025 3

roles

background 3

polarities

background 2 unclear 1

representative citing papers

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF

cs.LG · 2026-06-29 · conditional · novelty 7.0

PS-PPO samples a per-trajectory cutoff and importance-weights truncated gradients, preserving the full critic-free update in expectation while cutting RLHF training compute and memory.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

Self-Distilled Policy Gradient

cs.LG · 2026-06-02 · unverdicted · novelty 4.0

SDPG combines group-relative verifier advantages, normalized standard deviation, full-vocabulary on-policy self-distillation, and reference-policy KL regularization to improve stability and performance over RLVR and self-distillation baselines in language model RL.

citing papers explorer

Showing 12 of 12 citing papers.