Pith. sign in

REVIEW 7 cited by

Critique Ability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.04815 v1 pith:3OFY22UQ submitted 2023-10-07 cs.LG

classification cs.LG
keywords llmsmodelscritiquemodelabilitycritiqueslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Critical thinking is essential for rational decision-making and problem-solving. This skill hinges on the ability to provide precise and reasoned critiques and is a hallmark of human intelligence. In the era of large language models (LLMs), this study explores the ability of LLMs to deliver accurate critiques across various tasks. We are interested in this topic as a capable critic model could not only serve as a reliable evaluator, but also as a source of supervised signals for model tuning. Particularly, if a model can self-critique, it has the potential for autonomous self-improvement. To examine this, we introduce a unified evaluation framework for assessing the critique abilities of LLMs. We develop a benchmark called CriticBench, which comprises 3K high-quality natural language queries and corresponding model responses; and annotate the correctness of these responses. The benchmark cover tasks such as math problem-solving, code completion, and question answering. We evaluate multiple LLMs on the collected dataset and our analysis reveals several noteworthy insights: (1) Critique is generally challenging for most LLMs, and this capability often emerges only when models are sufficiently large. (2) In particular, self-critique is especially difficult. Even top-performing LLMs struggle to achieve satisfactory performance. (3) Models tend to have lower critique accuracy on problems where they are most uncertain. To this end, we introduce a simple yet effective baseline named self-check, which leverages self-critique to improve task performance for various models. We hope this study serves as an initial exploration into understanding the critique abilities of LLMs, and aims to inform future research, including the development of more proficient critic models and the application of critiques across diverse tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Large multimodal models mostly fail to proactively detect flawed textual premises, and their performance depends on error type and on how they weight text versus images.

  2. DeepCritic: Deliberate Critique with Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A two-stage SFT and RL pipeline turns a 7B instruct model into a deliberate math critic that outperforms GPT-4o and same-size R1-distill models at locating the first erroneous step.

  3. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  4. Can LLMs and humans be friends? Uncovering factors affecting human-AI intimacy formation

    cs.HC 2025-05 reject novelty 5.0 of 10

    Gradual self-disclosure raises perceived intimacy with an LLM chatbot, but the study's design cannot separate this effect from time or question order.

  5. Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving

    cs.AI 2025-05 conditional novelty 5.0 of 10

    MATH-VF formalizes LLM math solutions into SimpleMath and uses a tool-augmented critic to verify each reasoning step and offer corrective feedback.

  6. Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

    cs.SE 2025-05 conditional novelty 4.0 of 10

    Guided by a learned value critic, 1-step lookahead and trajectory selection lift an open-weights SWE agent to 40.8% on SWE-bench Verified, roughly doubling its success rate.

  7. MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree

    cs.LG 2024-11 reject novelty 2.0 of 10

    MC-NEST adds a constant probability term to MCTSr's node selection and reports improved AIME pass@1 for GPT-4o, but the numbers are weakened by test-set rollout tuning and internal inconsistencies.

Pith tools