TCOD stabilizes on-policy distillation for multi-turn agents via temporal curriculum on trajectory depth, improving performance up to 18 points over vanilla OPD and sometimes surpassing the teacher.
Trinity-rft: A general-purpose and unified framework for reinforcement fine-tuning of large language models
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 3years
2026 3verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
Geoalign curates rollouts in online LLM RL by learning a projector on hidden states to flag and replace directionally inconsistent examples, yielding higher final performance and less oscillation than baselines on dialogue and math tasks.
FEST improves RLVR sample efficiency on math and coding benchmarks by combining supervised signals, on-policy signals, and decaying weights on just 128 randomly chosen demonstrations, matching full-dataset baselines.
citing papers explorer
-
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
TCOD stabilizes on-policy distillation for multi-turn agents via temporal curriculum on trajectory depth, improving performance up to 18 points over vanilla OPD and sometimes surpassing the teacher.
-
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning
Geoalign curates rollouts in online LLM RL by learning a projector on hidden states to flag and replace directionally inconsistent examples, yielding higher final performance and less oscillation than baselines on dialogue and math tasks.
-
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
FEST improves RLVR sample efficiency on math and coding benchmarks by combining supervised signals, on-policy signals, and decaying weights on just 128 randomly chosen demonstrations, matching full-dataset baselines.