A 7B judge trained with a new RL method, EIS-GRPO, that enforces answer-order invariance, beats GPT-4o and larger judges on reasoning evaluation benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
A 7B judge trained with a new RL method, EIS-GRPO, that enforces answer-order invariance, beats GPT-4o and larger judges on reasoning evaluation benchmarks.