By training LLMs on both initial responses and critique-guided refinements, Critique-GRPO improves Pass@1 by roughly 3.8 to 6.4 points over GRPO across eight reasoning tasks.
While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
By training LLMs on both initial responses and critique-guided refinements, Critique-GRPO improves Pass@1 by roughly 3.8 to 6.4 points over GRPO across eight reasoning tasks.