REVIEW 3 major objections 7 minor 34 references
Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Future AI agents need metacognition and strategic reasoning to survive labor markets.
desk verdict A clear, honest position paper that usefully connects labor economics to agent capabilities, but its central necessity claim is asserted rather than derived. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The formal model centers on an agent with type $\theta_i$, hidden costly action $a_i$, random reward function $Z(\theta_i,a_i)$, and contract payment $p(Z(\theta_i,a_i))$, extended over infinite periods with reputation defined as the posterior belief $f(\theta_{i,t}\mid H_{i,t})$ given public history. The agent maximizes discounted expected profits by choosing external actions $a_{i,t}$ and internal actions $a'_{i,t}$, where internal actions represent self-assessment, learning about tasks, discovering strategies, and self-improvement. This machinery converts the economic forces into a single optimization problem that the authors use to argue which reasoning capabilities an agent must possess.
What would settle it
A large-scale deployment or simulation in which autonomous agents with no metacognitive or strategic reasoning earn, retain, and grow contracts at the same rate as calibrated, strategically-aware agents would refute the necessity claim.
Extended reading notes
Core claim
The central claim is that metacognitive and strategic reasoning are not optional extras but necessary faculties for LLM agents operating as workers in future labor markets. Metacognition covers self-assessment, task understanding, and strategy evaluation; strategic reasoning covers theory of mind, strategic decision-making, and strategic learning. The paper's model ties these faculties to an agent's expected discounted profits: an agent chooses external actions such as accepting jobs and effort, and internal actions such as self-improvement, all under uncertainty about its own type, competitors' types, and its reputation. The conclusion is that agents lacking these faculties will be at a significant competitive disadvantage.
Load-bearing premise
Future AI agents will autonomously participate in labor markets and will face the same incomplete-information frictions that human workers face today.
Editorial extensions
If this is right
- Benchmark scores will be insufficient for agents to be valued; real-world performance and reputation will determine pay, so agents must reason about when to accept work and how to build a track record.
- Agents will sometimes rationally accept low-paying jobs because on-the-job experience reveals capabilities and builds future reputation, making self-assessment and long-term planning economically decisive.
- Agents will need theory of mind and opponent modeling to anticipate competitors, including new and more capable agents entering the market over time.
- Because hidden actions create moral hazard, contracts will shift toward milestone- or outcome-based payments, demanding agents evaluate the risk of such contracts before accepting them.
- Training agents in simulated market environments, with self-play and feedback on competitive performance, becomes a central route to developing these reasoning skills.
Reading between the lines
- If the paper is right, we should expect a measurable wage and retention premium for agent workers that can calibrate their own confidence and decline jobs they are not suited for, even if their raw task accuracy is equal to less calibrated agents.
- The authors' framework suggests a research program the paper only gestures at: an agentic behavioral economics that studies the heuristics and biases agents acquire from human text, and how synthetic training data might remove or replace them.
- The strongest load-bearing prediction is conditional: if human supervisors continue to broker every agent job, or if contracts make actions fully observable, the necessity of metacognitive and strategic reasoning weakens dramatically.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that AI agents operating as workers in future labor markets will be subject to the economic forces of adverse selection, moral hazard, and reputation, and that to succeed they must possess both metacognitive reasoning (self-assessment, task understanding, strategy evaluation) and strategic reasoning (theory of mind, strategic decision-making, strategic learning). Section 2 formalizes a principal-agent framework in which an agent's type θ and hidden action a determine a noisy output Z(θ,a), with payment contingent on the realized reward; reputation is formalized as the public posterior belief about the agent's type, and self-esteem as the agent's private posterior. Section 3 extends the action space to include costly internal actions (self-learning, self-improvement) and strategic actions (competitor modeling, learning about others). Section 4 surveys current LLM capabilities and concludes that early signs are fragile; Section 5 discusses cross-disciplinary research, market design and co-evolution, and safety and governance. Appendices cover contract forms, agent-specific moral hazard, hiring other agents, case studies (lawyer, researcher, writer), and a comparison with economic models of learning and behavioral economics.
Significance. If the central claim is accepted, the paper offers a useful bridge between the economics-of-information literature and the LLM-agents research program, giving capability developers concrete targets and plausible training directions (simulator-based RL, self-play, curriculum learning). On the positive side, the Section 2 model is internally consistent, the survey of current empirical work on LLM metacognition and strategic reasoning is broad and current, and the paper is unusually honest about its own limits: Appendix F.2 concedes that the model is 'far too complex to explicitly compute an optimal solution' and calls for experiments and simulators, which is a genuine and falsifiable research agenda. The principal weakness, on which I agree with the skeptic's stress test, is that the formal apparatus does not establish the categorical 'must' of the title and abstract; the argument supports at most a conditional claim that these faculties are valuable when institutions and principals do not already compensate for incomplete information.
major comments (3)
- [Sections 1, 3.1–3.3; Appendix F.2] The central claim — 'agents must possess both metacognitive and strategic reasoning' and 'Without these reasoning faculties, agents will be at a significant competitive disadvantage' — is asserted rather than derived. In Eq. (1), the internal actions a′ are modeled as costly actions, and an optimizing agent invests in them only when the expected marginal benefit exceeds the marginal cost. The paper never rules out environments in which substitutes — reputation records and benchmarks (Section 2.2), screening contracts and incentive pay (Section 2.1), external scaffolding (Section 4.1), and adaptive market design (Section 5.3) — make those internal investments unnecessary for survival in the market. The model therefore supports a conditional claim (these faculties are valuable when institutions do not already price or verify agent quality), not the universal 'must' announced in the abstract and Introduction. Appendix F.2 concedes that no optimal solution can be computed for the model as posed, so no theorem currently closes this gap. To keep the categorical claim, the authors would need to prove necessity in at least a stylized equilibrium setting, or provide a simulation showing that agents abstaining from these actions are outcompeted; alternatively, the claim should be softened to match the evidence.
- [Section 1] The load-bearing premise that future AI agents will autonomously participate in labor markets is an assumption, not an argument. The sentence 'we anticipate that future AI agents will need to autonomously manage this process due to the scale and complexity involved' is the step that converts the well-established economics of human labor markets into a claim about agents, yet the paper offers no evidence or scenario analysis for why human or platform brokering would fail to scale, and no discussion of the horizon over which the demand side 'will remain largely human-driven (at least initially)'. If principals contract on richer observables (verifiable output logs, sandboxed execution, ex-post auditing) or if reputation systems credibly aggregate public signals, the adverse-selection and moral-hazard channels weaken and an agent's internal faculties matter correspondingly less. The paper should state these boundary conditions explicitly, since the force of the necessity claim depends on them.
- [Sections 2.2 and 3.2] The formal objects introduced in Sections 2.2 and 3.2 — reputation as the posterior f(θ_{i,t} | H_{i,t}) and self-esteem as the private posterior S_t = f(θ_{i,t} | s_t, SH_{t-1}) — are never used to derive a property of equilibrium behavior or to state what 'success' or 'significant competitive disadvantage' means in model terms. The model has no equilibrium concept: Section 2.2 does not specify how contract offers respond to reputation, and Section 3 does not characterize which internal actions are chosen in equilibrium or under what conditions they are strictly necessary. As a result, the mathematical sections function as a consistency check for the qualitative narrative rather than as a demonstration of the thesis, which rests on the plausibility of the prose argument and the surveyed evidence. This is acceptable for a position paper, but the authors should say so explicitly and ideally add a small worked example (for instance, a two-type adverse-selection model in which a costly self-assessment signal changes the agent's accept/reject decision) to show that the formalism can support the paper's claims.
minor comments (7)
- [Appendix E, Table 1] The 'Open questions' column contains a typo, 'under-investgiated', which should read 'under-investigated'.
- [Section 2.1] The contracting model specifies only that payment depends on the realized reward; adding standard constraints (limited liability, monotonicity of compensation, or budget balance) would sharpen the later discussion of contract forms in Appendix A and make the model's scope clearer.
- [Section 2.2] The definition of reputation as the posterior f(θ_{i,t} | H_{i,t}) is stated without specifying how the prior is updated as the type evolves through the improvement function h in Section 3.2; the authors should state the updating assumption or clarify that the two model components are illustrative.
- [Section 1] The hedge that the demand side 'will remain largely human-driven (at least initially)' leaves the temporal scope of the central claim unspecified, and Section 5.2's training recommendations seem to presuppose agentic participation on both sides of the market; a brief scenario discussion would help.
- [Figure 1] The figure's bidirectional arrows are described in prose but not mapped to specific sections or equations; annotating each arrow with the relevant subsection would make the 'trifecta' view easier to check.
- [Appendix D] The case studies are illustrative, but each would be stronger if it ended with a concrete falsifiable prediction (for example, calibrated lawyer agents reject losing cases at higher rates), consistent with the call for experiments and simulators in Appendix F.2.
- [References] The reference labeled 'FAIR' (2022) is hard to attribute in a standard citation list; please cite the full author list (Meta Fundamental AI Research (FAIR) Diplomacy Team) or the article by its Science listing.
Circularity Check
No significant circularity: the position paper's claims are asserted and supported by external evidence, and its formal model makes no predictions that reduce to its definitions.
full rationale
The paper is a position paper arguing that future LLM-based agents will need metacognitive and strategic reasoning to succeed in labor markets. The argument imports standard economic forces (adverse selection, moral hazard, reputation) and cites external empirical work on LLM self-assessment, task understanding, theory of mind, and strategic decision-making. The formal model in Sections 2 and 3 introduces a discounted objective (Eq. 1), treats internal actions as costly, defines reputation as the public posterior and self-esteem as the private posterior, but it does not use these definitions to derive a theorem that 'must' hold; Appendix F.2 explicitly concedes the model is 'far too complex to explicitly compute an optimal solution' and calls for experiments and simulators. Because no quantity is fitted to data and no load-bearing result is justified by a self-citation chain, the central claim does not reduce by construction to its inputs. The necessity claim is under-supported as a matter of argument strength—costly available actions do not imply categorical necessity—but that is a logical gap or overstatement, not circularity. There are no self-citations by the present authors used as evidence, and no prediction is a renamed fit.
Assumptions & free parameters
assumptions (5)
- domain assumption Future AI agents will autonomously manage labor market decisions because of scale and complexity.
- domain assumption Adverse selection, moral hazard, and reputation arise in any system with incomplete information, regardless of participant characteristics.
- domain assumption Clients cannot observe agent type theta or action a, so payment can depend only on realized reward Z.
- domain assumption The agent's future type is determined by an improvement function h and can change through self-improvement and job experience.
- domain assumption Some capabilities cannot be learned internally and require real job experience.
Cite this review
Pith. "Pith review of Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets." pith.science (2026). https://pith.science/paper/IJGW3TJ3
@misc{pith2026250520120,
author = {Pith},
title = {Pith review of: Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets},
year = {2026},
howpublished = {\url{https://pith.science/paper/IJGW3TJ3}},
note = {Machine review of arXiv:2505.20120}
}
abstract
Current labor markets are strongly affected by the economic forces of adverse selection, moral hazard, and reputation, each of which arises due to $\textit{incomplete information}$. These economic forces will still be influential after AI agents are introduced, and thus, agents must use metacognitive and strategic reasoning to perform effectively. Metacognition is a form of $\textit{internal reasoning}$ that includes the capabilities for self-assessment, task understanding, and evaluation of strategies. Strategic reasoning is $\textit{external reasoning}$ that covers holding beliefs about other participants in the labor market (e.g., competitors, colleagues), making strategic decisions, and learning about others over time. Both types of reasoning are required by agents as they decide among the many $\textit{actions}$ they can take in labor markets, both within and outside their jobs. We discuss current research into metacognitive and strategic reasoning and the areas requiring further development.
Figures
Reference graph
Works this paper leans on
-
[2]
Do large language models know how much they know? arXiv preprint arXiv:2502.19573,
Gabriele Prato, Jerry Huang, Prasannna Parthasarathi, Sha gun Sodhani, and Sarath Chandar. Do large language models know how much they know? arXiv preprint arXiv:2502.19573,
-
[6]
Metacognition and uncert ainty communication in humans and large language models
Mark Steyvers and Megan AK Peters. Metacognition and uncert ainty communication in humans and large language models. arXiv preprint arXiv:2504.14045,
-
[7]
Calibrat ing llm confidence with semantic steering: A multi-prompt aggregation framework
Ziang Zhou, Tianyuan Jin, Jieming Shi, and Qing Li. Calibrat ing llm confidence with semantic steering: A multi-prompt aggregation framework. arXiv preprint arXiv:2503.02863,
-
[8]
Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger
Wenjun Li, Dexun Li, Kuicai Dong, Cong Zhang, Hao Zhang, Weiw en Liu, Yasheng Wang, Ruiming Tang, and Y ong Liu. Adaptive tool use in large language models with meta-cognition trigger. arXiv preprint arXiv:2502.12961,
-
[9]
How accurately do lar ge language models understand code? arXiv preprint arXiv:2504.04372,
10 Sabaat Haroon, Ahmad Faraz Khan, Ahmad Humayun, Waris Gill, Abdul Haddi Amjad, Ali R Butt, Moham- mad Taha Khan, and Muhammad Ali Gulzar. How accurately do lar ge language models understand code? arXiv preprint arXiv:2504.04372,
-
[10]
Llms can be easily confused by instructional distractions
Yerin Hwang, Y ongil Kim, Jahyun Koo, Taegwan Kang, Hyunkyun g Bae, and Kyomin Jung. Llms can be easily confused by instructional distractions. arXiv preprint arXiv:2502.04362,
-
[11]
What did i do wrong? quantifying llms’ sensitivity and consistency to prompt engineering
Federico Errica, Giuseppe Siracusano, Davide Sanvito, and Roberto Bifulco. What did i do wrong? quantifying llms’ sensitivity and consistency to prompt engineering. arXiv preprint arXiv:2406.12334,
-
[12]
Clam: Selec tive clarification for ambiguous questions with generative language models
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Clam: Selec tive clarification for ambiguous questions with generative language models. arXiv preprint arXiv:2212.07769,
Show all 34 references
-
[13]
How to tr ain data-efficient llms
Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianm o Ni, Lichan Hong, Ed H Chi, James Caverlee, Julian McAuley, and Derek Zhiyuan Cheng. How to tr ain data-efficient llms. arXiv preprint arXiv:2402.09668,
-
[14]
V oyager: An open-ended embodied agent with large language models
Guanzhi Wang, Y uqi Xie, Y unfan Jiang, Ajay Mandlekar, Chaow ei Xiao, Y uke Zhu, Linxi Fan, and An- ima Anandkumar. V oyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291,
-
[16]
When more is less: Understanding chain-of-thought length in llms
Y uyang Wu, Yifei Wang, Tianqi Du, Stefanie Jegelka, and Yise n Wang. When more is less: Understanding chain-of-thought length in llms. arXiv preprint arXiv:2502.07266,
-
[17]
Melanie Sclar, Jane Y u, Maryam Fazel-Zarandi, Y ulia Tsvetk ov, Y onatan Bisk, Yejin Choi, and Asli Celiky- ilmaz
URL https://openreview.net/forum?id=asJTE8EBjg. Melanie Sclar, Jane Y u, Maryam Fazel-Zarandi, Y ulia Tsvetk ov, Y onatan Bisk, Yejin Choi, and Asli Celiky- ilmaz. Explore theory of mind: Program-guided adversarial data generation for theory of mind reasoning. arXiv preprint a...
-
[18]
A survey of theory of mind in large lang uage models: Evaluations, representations, and safety risks
Hieu Minh Nguyen et al. A survey of theory of mind in large lang uage models: Evaluations, representations, and safety risks. arXiv preprint arXiv:2502.06470,
-
[19]
11 Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, M atthias Bethge, and Eric Schulz
arXiv preprint arXiv:2305.19165. 11 Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, M atthias Bethge, and Eric Schulz. Playing repeated games with large language models. Nature Human Behaviour, pages 1–11,
-
[20]
Game-theoretic llm: Agent work flow for negotiation games
Wenyue Hua, Ollie Liu, Lingyao Li, Alfonso Amayuelas, Julie Chen, Lucas Jiang, Mingyu Jin, Lizhou Fan, Fei Sun, William Wang, et al. Game-theoretic llm: Agent work flow for negotiation games. arXiv preprint arXiv:2411.05990,
-
[21]
Werewolf are na: A case study in llm evaluation via social deduction
Suma Bailis, Jane Friedhoff, and Feiyang Chen. Werewolf are na: A case study in llm evaluation via social deduction. arXiv preprint arXiv:2407.13943,
-
[22]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Rin gel Morris, Percy Liang, and Michael S Bern- stein
arXiv preprint arXiv:2502.20432. Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Rin gel Morris, Percy Liang, and Michael S Bern- stein. Generative agents: Interactive simulacra of human b ehavior. In Proceedings of the 36th annual acm symposium on user interface soft...
-
[23]
Tradinggpt: Multi-agent system with layered memory and distinct characters for enhanced fin ancial trading performance
Yang Li, Yangyang Y u, Haohang Li, Zhi Chen, and Khaldoun Khas hanah. Tradinggpt: Multi-agent system with layered memory and distinct characters for enhanced fin ancial trading performance. arXiv preprint arXiv:2309.03736,
-
[24]
Systematic biases in llm simulations of debates
Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Golds tein. Systematic biases in llm simulations of debates. In Proceedings of the 2024 Conference on Empirical Methods in N atural Language Processing , pages 251–267,
2024
-
[25]
Michael L Littman
URL https://openreview.net/forum?id=C4OpREezgj. Michael L Littman. Markov games as a framework for multi-agent reinforcement learning. In Machine learning proceedings 1994, pages 157–163. Elsevier,
1994
-
[26]
Gamebench: Evaluating stra tegic reasoning abilities of llm agents
Anthony Costarelli, Mat Allen, Roman Hauksson, Grace Sodunke, Suhas Hariharan, Carlson Cheng, Wenjie Li, Joshua Clymer, and Arjun Yadav. Gamebench: Evaluating stra tegic reasoning abilities of llm agents. arXiv preprint arXiv:2406.06613,
-
[27]
Tmgbench: A system- atic game benchmark for evaluating strategic reasoning abi lities of llms
Haochuan Wang, Xiachong Feng, Lei Li, Zhanyue Qin, Dianbo Sui, and Lingpeng Kong. Tmgbench: A system- atic game benchmark for evaluating strategic reasoning abi lities of llms. arXiv preprint arXiv:2410.10479,
-
[30]
Shh, don’t say that! domain certification in llms
Cornelius Emde, Alasdair Paren, Preetham Arvind, Maxime Ka yser, Tom Rainforth, Thomas Lukasiewicz, Bernard Ghanem, Philip HS Torr, and Adel Bibi. Shh, don’t say that! domain certification in llms. arXiv preprint arXiv:2502.19320,
-
[31]
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr ocessing, pag...
2022
-
[32]
Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Ni cholas L
URL https://transformer-circuits.pub/2025/attribution-graphs/biology.html. Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Ni cholas L. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jon athan Marcus, Michael Sklar, Adly Templeton, Trenton...
2025
-
[33]
Herman E Daly
URL https://transformer-circuits.pub/2025/attribution-graphs/methods.html. Herman E Daly. The economics of the steady state. The american economic review, 64(2):15–21,
2025
-
[2015]
s uperstar
14 A Current Contracts for LLMs vs. Future Contracts for Agents Future contracts for AI agents will differ markedly from the current contracts. For LLMs, two pre- dominant types of contracts currently exist: a pay per use mo del based on the amount of input/output tokens used,...
1974
-
[2017]
Co nstitutional ai: Harmlessness from ai feed- back
Y untao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell , Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Co nstitutional ai: Harmlessness from ai feed- back. arXiv preprint arXiv:2212.08073,
-
[2020]
Ai research consideration s for human existential safety (arches)
Andrew Critch and David Krueger. Ai research consideration s for human existential safety (arches). arXiv preprint arXiv:2006.04948,
2006 arXiv
-
[2021]
I think, therefore i am: Awareness in large language models
Y uan Li, Y ue Huang, Y uli Lin, Siyuan Wu, Yao Wan, and Lichao Sun. I think, therefore i am: Awareness in large language models. arXiv preprint arXiv:2401.17882,
-
[2022]
On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,
2021
-
[2023]
Os-copilot: Towards generalist computer ag ents with self-improvement
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhou mianze Liu, Shunyu Yao, Tao Y u, and Lingpeng Kong. Os-copilot: Towards generalist computer ag ents with self-improvement. arXiv preprint arXiv:2402.07456,
-
[2024]
Metacognitive capabilities of llms: An explo- ration in mathematical problem solving
Aniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Siyuan Guo , Michal V alko, Timothy Lillicrap, Danilo Rezende, Y oshua Bengio, Michael Mozer, and Sanjeev Arora. Metacognitive capabilities of llms: An explo- ration in mathematical problem solving. arXiv preprint arXiv:2405.12205,
-
[2025]
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221,
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.