SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.
• From the second equation, we get: t = 540 s+2 − 144
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.