RSTG selectively distills a teacher on negative zero-variance prompts, with confidence weighting, token-level selection, and auxiliary SFT, improving math and code RL post-training over naive GRPO+OPD.
CoRR , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
RSTG selectively distills a teacher on negative zero-variance prompts, with confidence weighting, token-level selection, and auxiliary SFT, improving math and code RL post-training over naive GRPO+OPD.