Pith. sign in

REVIEW 2 cited by

Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02658 v2 pith:YSP72YZT submitted 2024-02-05 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords verifiersupervisionaccuracydataintermediatemipsprocesscompletions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Process supervision, using a trained verifier to evaluate the intermediate steps generated by a reasoner, has demonstrated significant improvements in multi-step problem solving. In this paper, to avoid the expensive effort of human annotation on the verifier training data, we introduce Model-induced Process Supervision (MiPS), a novel method for automating data curation. MiPS annotates an intermediate step by sampling completions of this solution through the reasoning model, and obtaining an accuracy defined as the proportion of correct completions. Inaccuracies of the reasoner would cause MiPS underestimating the accuracy of intermediate steps, therefore, we suggest and empirically show that verification focusing on high predicted scores of the verifier shall be preferred over that of low predicted scores, contrary to prior observations on human curated data. Our approach significantly improves the performance of PaLM 2 on math and coding tasks (accuracy +0.67% on GSM8K, +4.16% on MATH, +0.92% on MBPP compared with an output supervision trained verifier). Additionally, our study demonstrates that the verifier exhibits strong generalization ability across different reasoning models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RRO: LLM Agent Optimization Through Rising Reward Trajectories

    cs.AI 2025-05 reject novelty 5.0 of 10

    RRO samples LLM agent actions until a step shows a rising process reward, converts those steps into DPO preference pairs, and reports modest gains on WebShop and InterCode-SQL with fewer samples.

  2. Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A process reward model trained with Error Propagation and Error Cessation labels from an o1 judge outperforms existing PRMs on math selection and step-level scoring.

Pith tools