Pith. sign in

Rewarding progress: Scaling automated process verifiers for LLM reasoning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Intra-Trajectory Consistency for Reward Modeling

cs.LG · 2025-06-10 · conditional · novelty 6.0

Intra-trajectory consistency regularization, weighted by next-token generation probabilities, improves outcome reward models on RewardBench and in downstream DPO and best-of-N evaluations.

citing papers explorer

Showing 1 of 1 citing paper.

  • Intra-Trajectory Consistency for Reward Modeling cs.LG · 2025-06-10 · conditional · none · ref 6

    Intra-trajectory consistency regularization, weighted by next-token generation probabilities, improves outcome reward models on RewardBench and in downstream DPO and best-of-N evaluations.