• From the second equation, we get: t = 540 s+2 − 144

Solve the Equations: • From the first equation, we get: t = 540 s − 240

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning

cs.CL · 2025-12-17 · unverdicted · novelty 7.0

SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.

citing papers explorer

Showing 1 of 1 citing paper.

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning cs.CL · 2025-12-17 · unverdicted · none · ref 13
SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.

• From the second equation, we get: t = 540 s+2 − 144

fields

years

verdicts

representative citing papers

citing papers explorer