S-GRPO and T-SPMO, which update only a subset of output tokens, beat full-token GRPO and the base model on arithmetic reasoning under LoRA fine-tuning.
Buy 4 reinforce samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction Workshop at ICLR 2019, 2019
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Token-Efficient RL for LLM Reasoning
S-GRPO and T-SPMO, which update only a subset of output tokens, beat full-token GRPO and the base model on arithmetic reasoning under LoRA fine-tuning.