Pith. sign in

REVIEW 5 cited by

Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12521 v1 pith:PJ2ZM3SB submitted 2025-02-18 cs.AI cs.LG

classification cs.AIcs.LG
keywords reasoninginference-timeplanningtaskstechniquesacrossbenchmarkperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Notably, OpenAI's o1 model shows promising performance through its novel use of multi-step reasoning and verification. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as Sys2Bench, and perform extensive experiments evaluating existing inference-time techniques on eleven diverse tasks across five categories, including arithmetic reasoning, logical reasoning, common sense reasoning, algorithmic reasoning, and planning. Our findings indicate that simply scaling inference-time computation has limitations, as no single inference-time technique consistently performs well across all reasoning and planning tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

    cs.RO 2026-03 accept novelty 6.0 of 10

    SCOUT matches LLM planners on open-world interactive object search by scoring 3D scene-graph nodes with lightweight models distilled from LLM relational priors, at far lower compute cost.

  2. Empirical Modeling of Therapist-Client Dynamics in Psychotherapy Using LLM-Based Assessments

    cs.CY 2026-02 reject novelty 6.0 of 10

    LLM-based scoring of 1,610 therapy sessions finds therapist empathy and exploration are followed by more client disclosure, while prior-session rapport is associated with less self-directed negative emotion—but the cl...

  3. OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A private, contamination-resistant benchmark of 250 olympiad-level programming problems shows top reasoning models reaching about 36% solve rates, far above conventional models.

  4. Quantum Circuit Generation via test-time learning with large language models

    quant-ph 2026-02 conditional novelty 5.0 of 10

    An LLM with memory, score feedback, and restart-from-best finds high-entanglement quantum circuits, reaching Meyer-Wallach 1.0 on 25 qubits within 45 queries.

  5. Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.

Pith tools