CoRE replaces majority-vote rewards in test-time RL with a graph-based consensus extracted by replicator dynamics, improving accuracy and convergence speed on math and science benchmarks.
Let’s verify step by step
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
CoRE replaces majority-vote rewards in test-time RL with a graph-based consensus extracted by replicator dynamics, improving accuracy and convergence speed on math and science benchmarks.