Pith. sign in

hub

Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a

19 Pith papers cite this work. Polarity classification is still indexing.

19 Pith papers citing it

hub tools

citation-role summary

background 3 method 1

citation-polarity summary

years

2026 17 2025 2

polarities

background 4

representative citing papers

Task-Focused Memorization for Multimodal Agents

cs.CV · 2026-05-29 · unverdicted · novelty 6.0

TaskMem uses RL in two phases to learn a task-focused memorization policy for multimodal agents, yielding 5.3-7.0% VQA accuracy gains on reformulated streaming benchmarks from VideoMME, EgoLife, and EgoTempo.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

cs.AI · 2026-05-13 · unverdicted · novelty 6.0

ICRL uses joint RL training of solver and critic with distribution-calibration re-weighting and role-wise advantage estimation to internalize critique into unassisted LLM performance, yielding 6.4-point gains on agentic tasks and 7.0 on math reasoning with Qwen3 models.

Rethinking the Divergence Regularization in LLM RL

cs.LG · 2026-06-08 · unverdicted · novelty 5.0

DRPO introduces a smooth quadratic regularizer on policy divergence that preserves DPPO's trust-region geometry while providing continuous corrective gradients instead of hard masking.

citing papers explorer

Showing 19 of 19 citing papers.