Pith. sign in

Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms

11 Pith papers cite this work. Polarity classification is still indexing.

11 Pith papers citing it

citation-role summary

method 1

citation-polarity summary

years

2026 11

roles

method 1

polarities

use method 1

representative citing papers

Tandem Reinforcement Learning with Verifiable Rewards

cs.AI · 2026-06-26 · unverdicted · novelty 7.0

TRL extends tandem training to RLVR pipelines, matching GRPO solo reasoning on Qwen3-4B math tasks while improving handoff robustness, reducing distributional drift, and increasing CoT legibility for the junior.

Not only where, But when: Temporal Scheduling for RLVR

cs.LG · 2026-05-25 · unverdicted · novelty 7.0

Temporal scheduling of credit allocation criteria over RLVR training, using trajectory percentiles to target heterogeneous behaviors, yields more stable policy entropy and better reasoning benchmark results than static allocation.

APPO: Agentic Procedural Policy Optimization

cs.LG · 2026-06-10 · conditional · novelty 5.0

APPO improves LLM agent training by branching at tokens selected for both uncertainty and future impact, then scaling credit for consequential reasoning procedures.

One-Way Policy Optimization for Self-Evolving LLMs

cs.LG · 2026-05-21 · unverdicted · novelty 5.0

OWPO decouples optimization direction from magnitude via asymmetric reweighting (Accelerated Alignment for inferior deviations, Gain Locking for superior) plus iterative references to create a ratchet effect for continuous LLM improvement.

citing papers explorer

Showing 11 of 11 citing papers.