Pith. sign in

REVIEW 5 cited by

Metacognitive Prompting Improves Understanding in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.05342 v4 pith:C3KBBJX3 submitted 2023-08-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmspromptingtasksmodelsunderstandinglanguagereasoningabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Large Language Models (LLMs), there have been consistent advancements in task-specific performance, largely influenced by effective prompt design. Recent advancements in prompting have enhanced reasoning in logic-intensive tasks for LLMs, yet the nuanced understanding abilities of these models, crucial for processing and interpreting complex information, remain underexplored. In this study, we introduce Metacognitive Prompting (MP), a strategy inspired by human introspective reasoning processes. Using MP, LLMs undergo a systematic series of structured, self-aware evaluations, drawing on both their vast inherent knowledge and new insights. We conduct extensive experiments on four prevalent LLMs: Llama2, PaLM2, GPT-3.5, and GPT-4, across ten natural language understanding (NLU) datasets from GLUE, SuperGLUE, BLUE, and LexGLUE benchmarks. Additionally, we compare our method with chain-of-thought prompting and its advanced versions. The results show that GPT-4 consistently excels across all tasks, while other models have shown significant progress in some tasks when used in conjunction with MP. Furthermore, MP consistently outperforms existing prompting methods in both general and domain-specific NLU tasks. This study underscores the potential to amplify the understanding abilities of LLMs and highlights the benefits of mirroring human introspective reasoning in NLU tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Large vision-language models that appear to recognize visual illusions often answer fake-illusion questions from prior knowledge, not from actually seeing the images.

  2. AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code

    cs.SE 2025-07 conditional novelty 5.0 of 10

    AccessGuru combines accessibility testing tools and LLM prompting to correct syntactic, semantic, and layout HTML accessibility violations, reporting up to 84% average violation score decrease on a new benchmark.

  3. Referential ambiguity and clarification requests: comparing human and LLM behaviour

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Humans seldom ask clarification questions for referential ambiguity, while LLMs ask them more often, and reasoning prompts increase LLM question frequency and relevance.

  4. Prompt Engineering for Requirements Engineering: A Literature Review and Roadmap

    cs.SE 2025-07 conditional novelty 5.0 of 10

    The first roadmap-oriented systematic literature review of prompt engineering for requirements engineering analyzes 35 studies and proposes a hybrid taxonomy and research roadmap.

  5. Bhatt Conjectures: On Necessary-But-Not-Sufficient Benchmark Tautology for Human Like Reasoning

    cs.CR 2025-06 reject novelty 2.0 of 10

    A position paper restates known AI evaluation criteria as 'tautological' definitions of reasoning and understanding, with proposed diagnostic tests but no empirical evidence.

Pith tools