pith. sign in

hub Canonical reference

Self-evolving curriculum for LLM reasoning

Canonical reference. 88% of citing Pith papers cite this work as background.

16 Pith papers citing it
Background 88% of classified citations

hub tools

citation-role summary

background 7 baseline 1

citation-polarity summary

years

2026 13 2025 3

representative citing papers

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning

cs.LG · 2026-05-11 · unverdicted · novelty 6.0

METIS internalizes curriculum judgment in LLM reinforcement fine-tuning by predicting within-prompt reward variance via in-context learning and jointly optimizing with a self-judgment reward, yielding superior performance and up to 67% faster convergence across math, code, and agent benchmarks.

Policy Improvement Reinforcement Learning

cs.LG · 2026-04-01 · unverdicted · novelty 6.0

PIRL maximizes cumulative policy improvement across iterations instead of surrogate rewards and is proven aligned with final performance; PIPO implements it via retrospective verification for stable closed-loop optimization.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds

cs.LG · 2025-10-09 · unverdicted · novelty 6.0

The paper defines a Gradient Gap for RLVR policy gradients and proves a sharp step-size threshold below which training converges and above which it collapses, with predictions for length and success-rate scaling validated in simulations and on Qwen2.5-Math-7B.

citing papers explorer

Showing 16 of 16 citing papers.