Pith. sign in

Buy 4 reinforce samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction Workshop at ICLR 2019, 2019

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Token-Efficient RL for LLM Reasoning

cs.LG · 2025-04-29 · conditional · novelty 4.0

S-GRPO and T-SPMO, which update only a subset of output tokens, beat full-token GRPO and the base model on arithmetic reasoning under LoRA fine-tuning.

citing papers explorer

Showing 1 of 1 citing paper.

  • Token-Efficient RL for LLM Reasoning cs.LG · 2025-04-29 · conditional · none · ref 3

    S-GRPO and T-SPMO, which update only a subset of output tokens, beat full-token GRPO and the base model on arithmetic reasoning under LoRA fine-tuning.