Pith. sign in

REVIEW 2 cited by

Dyve: Thinking Fast and Slow for Dynamic Process Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11157 v1 pith:5NFA33AX submitted 2025-02-16 cs.AI

classification cs.AI
keywords dyveprocessdynamicfastslowsupervisionsystemthinking
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present Dyve, a dynamic process verifier that enhances reasoning error detection in large language models by integrating fast and slow thinking, inspired by Kahneman's Systems Theory. Dyve adaptively applies immediate token-level confirmation System 1 for straightforward steps and comprehensive analysis System 2 for complex ones. Leveraging a novel step-wise consensus-filtered process supervision technique, combining Monte Carlo estimation with LLM based evaluation, Dyve curates high-quality supervision signals from noisy data. Experimental results on ProcessBench and the MATH dataset confirm that Dyve significantly outperforms existing process-based verifiers and boosts performance in Best-of-N settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier

    cs.AI 2025-05 conditional novelty 6.0 of 10

    FlexiVe, a GRPO-trained generative verifier with fast and slow modes, plus an early-detection pipeline, improves AIME math accuracy while reducing tokens versus self-consistency.

  2. Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A two-model collaborative reasoning framework (Long⊗Short) scores reasoning thoughts by rollout accuracy, trains one model for important thoughts and one for the rest, and reports over 80% token savings with small acc...

Pith tools