Frontier LLMs released since early 2024 show limited, context-dependent evidence of metacognition in two non-linguistic games, with effects that are stronger in newer models but far from human-level.
Large Language Models lack essential metacognition for reliable medical reasoning
6 Pith papers cite this work, alongside 86 external citations. Polarity classification is still indexing.
representative citing papers
Fine-tuning LLMs on a clinician-curated ICU reasoning dataset improved their scores on five clinical reasoning benchmarks, including benchmarks outside critical care.
The paper defines warranted reliance on GenAI as requiring three non-fungible conditions: epistemic humility, epistemic access, and resistance to epistemic injustice.
RLMF uses quality of model self-judgments to refine RL rankings and select training data, achieving SOTA faithful calibration while preserving accuracy and outperforming standard RL by up to 63%.
Medical disclaimers in LLM and VLM outputs declined sharply from 2022 to 2025, dropping from 26.3% to 0.97% for text questions and from 19.6% to 1.05% for images.
IndicBERT-HPA with language-aware adapters and verification-guided deferral outperforms baselines on multilingual orthopedic note classification, reaching 0.8792 Macro-F1 overall and 84.4% selective accuracy at 72.3% coverage.
citing papers explorer
-
Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains
Fine-tuning LLMs on a clinician-curated ICU reasoning dataset improved their scores on five clinical reasoning benchmarks, including benchmarks outside critical care.
-
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
RLMF uses quality of model self-judgments to refine RL rankings and select training data, achieving SOTA faithful calibration while preserving accuracy and outperforming standard RL by up to 63%.
-
A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models
Medical disclaimers in LLM and VLM outputs declined sharply from 2022 to 2025, dropping from 26.3% to 0.97% for text questions and from 19.6% to 1.05% for images.
-
Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral
IndicBERT-HPA with language-aware adapters and verification-guided deferral outperforms baselines on multilingual orthopedic note classification, reaching 0.8792 Macro-F1 overall and 84.4% selective accuracy at 72.3% coverage.