By training LLMs on both initial responses and critique-guided refinements, Critique-GRPO improves Pass@1 by roughly 3.8 to 6.4 points over GRPO across eight reasoning tasks.
needle in a haystack
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
By training LLMs on both initial responses and critique-guided refinements, Critique-GRPO improves Pass@1 by roughly 3.8 to 6.4 points over GRPO across eight reasoning tasks.