SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.
• The total time spent walking is 2 + 24 60 = 2.4 hours when the speed is s + 2 km/h
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.