A GRPO-based method that rewards only self-reflection tokens, not answer tokens, improves LLM accuracy on function calling and Countdown math tasks using only binary success/failure feedback.
Title resolution pending
1 Pith paper cite this work, alongside 17 external citations. Polarity classification is still indexing.
1
Pith paper citing it
17
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
A GRPO-based method that rewards only self-reflection tokens, not answer tokens, improves LLM accuracy on function calling and Countdown math tasks using only binary success/failure feedback.