Pith. sign in

REVIEW 9 cited by

Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12205 v1 pith:UVBPH4AC submitted 2024-05-20 cs.AI cs.LG

classification cs.AIcs.LG
keywords skilllabelsmathquestionsreasoningknowledgellmsmetacognitive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Metacognitive knowledge refers to humans' intuitive knowledge of their own thinking and reasoning processes. Today's best LLMs clearly possess some reasoning processes. The paper gives evidence that they also have metacognitive knowledge, including ability to name skills and procedures to apply given a task. We explore this primarily in context of math reasoning, developing a prompt-guided interaction procedure to get a powerful LLM to assign sensible skill labels to math questions, followed by having it perform semantic clustering to obtain coarser families of skill labels. These coarse skill labels look interpretable to humans. To validate that these skill labels are meaningful and relevant to the LLM's reasoning processes we perform the following experiments. (a) We ask GPT-4 to assign skill labels to training questions in math datasets GSM8K and MATH. (b) When using an LLM to solve the test questions, we present it with the full list of skill labels and ask it to identify the skill needed. Then it is presented with randomly selected exemplar solved questions associated with that skill label. This improves accuracy on GSM8k and MATH for several strong LLMs, including code-assisted models. The methodology presented is domain-agnostic, even though this article applies it to math problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation

    stat.ML 2025-09 conditional novelty 7.0 of 10

    A Fisher random walk weighted residual estimator achieves semiparametric efficient confidence intervals for contextual Bradley-Terry-Luce preference comparisons with flexible score estimators.

  2. From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

    cs.CL 2026-07 accept novelty 6.0 of 10

    LLM formal provers must shift from competition solvers to research agents that handle open-ended, under-specified frontier mathematics under machine-checked rigor.

  3. The Role of Diversity in In-Context Learning for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Diversity-aware selection of in-context examples improves performance on complex and out-of-distribution tasks, though effect sizes are often modest.

  4. Truly Self-Improving Agents Require Intrinsic Metacognitive Learning

    cs.AI 2025-06 conditional novelty 5.0 of 10

    The paper proposes that self-improving agents must learn to manage their own learning processes, framing this as intrinsic metacognitive learning, and argues it is necessary for sustained and generalized improvement.

  5. Does It Make Sense to Speak of Introspection in Large Language Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.

  6. Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets

    cs.AI 2025-05 conditional novelty 5.0 of 10

    AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.

  7. Factored space models: Towards causality between levels of abstraction

    cs.AI 2024-12 conditional novelty 5.0 of 10

    Structural independence in a factored space is equivalent to conditional independence in all product distributions, generalizing d-separation to deterministic functions.

  8. LemmaHead: RAG Assisted Proof Generation Using Large Language Models

    cs.LG 2025-01 reject novelty 4.0 of 10

    RAG with five rounds of iterative hint retrieval lets GPT-4 prove 40% of MiniF2F formal problems, up from a 9.4% single-pass baseline.

  9. A Survey on Large Language Models for Mathematical Reasoning

    cs.AI 2025-06 conditional novelty 1.0 of 10

    Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.

Pith tools