Pith. sign in

REVIEW 4 cited by

StepWiser: Stepwise Generative Judges for Wiser Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.19229 v2 pith:2JFOB32W submitted 2025-08-26 cs.AI cs.CL

classification cs.AIcs.CL
keywords reasoningmodelstepsgenerativeintermediatemodelspolicyproviding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As models increasingly leverage multi-step reasoning strategies to solve complex problems, supervising the logical validity of these intermediate steps has become a critical research challenge. Process reward models address this by providing step-by-step feedback, but current approaches have two major drawbacks: they typically function as classifiers without providing explanations, and their reliance on supervised fine-tuning with static datasets limits generalization. Inspired by recent advances, we reframe stepwise reward modeling from a classification task to a reasoning task itself. We thus propose a generative judge that reasons about the policy model's reasoning steps (i.e., meta-reasons), outputting thinking tokens before delivering a final verdict. Our model, StepWiser, is trained by reinforcement learning using relative outcomes of rollouts. We show it provides (i) better judgment accuracy on intermediate steps than existing methods; (ii) can be used to improve the policy model at training time; and (iii) improves inference-time search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Agent-declared causal-hypothesis boundaries yield variable-length semantic phases that stay attributable after declaration scrubbing, but the resulting DPO preference signal is construction-bound and does not transfer...

  2. P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

    cs.CL 2026-01 unverdicted novelty 7.0 of 10

    P-Check trains a checklist generator that produces query-specific, user-weighted evaluation criteria, improving LLM-judge reward accuracy on personalization benchmarks.

  3. Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

    cs.LG 2025-10 conditional novelty 6.0 of 10

    MAHALO aligns LLMs to multiple objectives in one model via per-objective action heads and PRM-guided decoding, improving math, value, and tutoring metrics jointly.

  4. Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation

    cs.LG 2026-02 conditional novelty 4.0 of 10

    A taxonomy-driven survey arguing that reward design is the central mechanism shaping reliable LLM reasoning, with maps of reward paradigms, reward-hacking failure modes, and benchmark pitfalls.

Pith tools