A training-free, model-agnostic dashboard that exposes per-block context state (tokens, age, budget) with lossless archive/recovery improves long-horizon tool-agent performance on LOCA-Bench, BrowseComp-Plus, and GAIA.
arXiv preprint arXiv:2509.21545 (2025)
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5roles
background 2representative citing papers
The Metacognitive Monitoring Battery applied to 20 LLMs identifies three self-monitoring profiles, shows inverted accuracy and sensitivity ranks, and finds retrospective and prospective regulation largely dissociable.
Elicited preferences in LLMs do not function as effective incentives for higher-quality outputs on realistic writing tasks.
A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.
Meta-d′ and signal detection theory give comparable measures of AI metacognitive sensitivity and risk-sensitive decision regulation, demonstrated on three LLMs.
citing papers explorer
-
LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception
A training-free, model-agnostic dashboard that exposes per-block context state (tokens, age, budget) with lossless archive/recovery improves long-horizon tool-agent performance on LOCA-Bench, BrowseComp-Plus, and GAIA.
-
The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring
The Metacognitive Monitoring Battery applied to 20 LLMs identifies three self-monitoring profiles, shows inverted accuracy and sensitivity ranks, and finds retrospective and prospective regulation largely dissociable.
-
When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
Elicited preferences in LLMs do not function as effective incentives for higher-quality outputs on realistic writing tasks.
-
Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models
A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.
-
Measuring the metacognition of AI
Meta-d′ and signal detection theory give comparable measures of AI metacognitive sensitivity and risk-sensitive decision regulation, demonstrated on three LLMs.