REVIEW 4 major objections 5 minor 111 references
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read A five-part taxonomy organizes the cognitive capability gaps that keep generative and agentic AI from sustained, reliable cognition.
desk verdict A useful organizing frame for a fragmented literature, but the taxonomy is a plausibility argument, not a discovered fact. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The taxonomy itself is the central organizing identity: five cognitive components with their subcapability gaps, derived via a narrative synthesis of the literature. The second piece of machinery is ACIA, a closed-loop cognitive architecture whose six components map one-to-one onto the five taxonomy dimensions (perception/attention and action both feed environment interaction). The paper's third mechanism is cognition-centric evaluation, illustrated by three named metrics—Cognitive Persistence Index, Cognitive Adaptation Rate, and Cognitive Consistency Score—that measure whether a system retains, adapts, and stays coherent over time.
What would settle it
A systematic literature review that codes the same research areas with explicit, prespecified inclusion criteria and finds additional fundamental dimensions beyond these five—or fails to reproduce them—would directly refute the taxonomy's claim to completeness.
Extended reading notes
Core claim
The central claim is that five recurring dimensions—persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation—organize the major cognitive capability gaps in modern generative and agentic AI systems. Each dimension is decomposed into concrete subcomponents (e.g., persistent memory, state revision, latent state modeling; goal formulation, planning, goal persistence; metacognitive monitoring, uncertainty detection, abstention; world modeling, tool-augmented reasoning, feedback loops; policy adaptation, continual learning, safe knowledge updates). The paper argues that current systems fail not on any single benchmark but
Load-bearing premise
The taxonomy and ACIA stand on the assumption that a narrative synthesis of the authors' chosen literature reveals the five fundamental, recurring dimensions—if that literature selection is not representative of the field, the unified framework may be an artifact of the selection rather than a property of the field.
Editorial extensions
If this is right
- If the taxonomy is right, fragmented research threads in memory, planning, metacognition, grounding, and continual learning can be compared and combined under a single set of capability headings.
- ACIA provides a blueprint for what a cognitive AI system must contain—persistent state, goal management, metacognitive control, environment coupling, and learning—and implies that scaling models alone cannot close the gaps.
- The proposed evaluation metrics shift assessment from one-shot accuracy to longitudinal behavior: does the system keep memory, adapt without forgetting, and avoid contradictions over extended interaction?
- The paper's framing suggests that agentic AI systems currently excel at short-horizon tasks while systematically missing the persistent internal state and self-regulation needed for long-horizon autonomy.
Reading between the lines
- The five dimensions may imply a developmental ordering—persistent state before goal stability, goal stability before meaningful metacognition—that the paper does not explicitly assert but that could be tested empirically.
- Because the taxonomy rests on a narrative literature synthesis, a larger systematic review with explicit inclusion criteria might confirm, revise, or extend the dimension set; that test is a natural next step.
- The concrete metrics the paper sketches (CPI, CAR, CCS) are presented as illustrative functions; a testable extension would be to instantiate them with specific formulas and validate them against existing benchmark failures.
- ACIA is conceptual and unbuilt; its strongest testable claim is that each component addresses a specific gap, so a partial implementation could directly examine whether adding metacognitive control or persistent memory actually reduces observed goal drift and inconsistency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a taxonomy of cognitive capability gaps in modern generative and agentic AI, organizing the literature around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension it reviews representative approaches, identifies recurring limitations, and distills open problems. The paper then outlines a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) whose components are mapped to the five gaps, and closes with a discussion of evaluation strategies, including illustrative cognition-centric metrics. The intended contribution is a unified organizing framework and roadmap for research toward Cognitive AI.
Significance. If the taxonomy is accepted as a faithful organization of the field, it would provide a useful shared vocabulary for comparing work on memory, planning, metacognition, grounding, and continual learning. The survey is broad and generally accurate in its characterization of current limitations, and it is explicit that ACIA is conceptual rather than implemented. The main value lies in the synthesis and in the articulation of evaluation challenges. However, the paper's central claims — that the five dimensions are the recurring fundamental gaps, that ACIA addresses them, and that the proposed metrics are useful — are currently not supported at the level the contributions claim. The strengths are the breadth of coverage, the clear graphical summaries (Figures 1 and 3, Table II), and the candid acknowledgment that many proposed mechanisms remain partial.
major comments (4)
- [§III, 'Survey Methodology and Taxonomy Overview'] The load-bearing step is the sentence 'Through this synthesis, five recurring dimensions of cognitive capability emerged across the surveyed literature.' No search strategy, inclusion criteria, coding protocol, inter-rater reliability, or comparison with competing taxonomies is provided. The paper calls the method a 'narrative synthesis approach,' which permits the selection of literature to be shaped by the authors' prior frame; notably, reference [27] is the same group's earlier 'cognitive autonomy' paper that already names similar dimensions. This creates a risk that the five dimensions are an artifact of the selection rather than a property of the field. Please either specify the methodology (corpus, inclusion criteria, coding, inter-rater checks) or reframe the contribution as a perspective/proposal rather than a survey-derived taxonomy. A comparison with at least one existing cogni
- [§V, Table III (ACIA mapping)] The claim that ACIA 'addresses' the five identified gaps is weakened by the fact that the mapping in Table III is one-to-one by construction: Memory is assigned to Persistent State Modeling, Metacognition to Self-Monitoring and Control, etc., using the same names as the gaps. Since no architectural details, formal specification, or empirical evidence are provided, the statement 'These components address the five cognitive capability gaps identified throughout this survey' is close to a restatement of the taxonomy rather than an independently motivated design. If the goal is a conceptual architecture, please state explicitly what distinguishes ACIA from a restatement of the taxonomy and provide at least a high-level mechanism for each component (e.g., how memory is updated, how metacognition triggers replanning). If the goal is a proposal, soften the 'addresses' claim to 'is designed to a
- [§VI-C, Table IV] The proposed metrics are not operationalizable as written. CPI is defined as f(T_ret, P_mem, R_mem), CAR as f(A_succ, A_opp), and CCS as f(1 - N_contra/N_reason), with the note that f(·) 'may vary across applications.' This means the paper provides no concrete metric, and the table is an illustration rather than a proposed evaluation framework. Since 'examining emerging directions in cognition-centric evaluation' and 'discuss cognition-centric evaluation strategies' are listed as contributions, please either provide concrete formulations (with at least one worked example or reference to an existing benchmark that instantiates them) or explicitly label this section as an open research agenda, not a metric proposal.
- [§II-A, Cognitive Artificial Intelligence] The paper defines Cognitive AI and its capability requirements largely through reference [27], which is authored by overlapping authors. This is a circularity risk: the survey's organizing frame is anchored to the same group's prior work without independent grounding. For example, the statement 'Cognitive AI seeks to integrate perception, memory, reasoning, learning, action, and metacognitive control within a unified feedback-driven framework [27], [26]' relies on [27] for the definition while [27] itself is the source of the paradigm. Please ground the definition in a broader set of sources (e.g., cognitive architectures, AGI safety, neuroscience-inspired AI) and explicitly discuss how [27] relates to the present survey.
minor comments (5)
- [§IV-D, heading] The heading 'T ool-Augmented Reasoning' contains a typographical spacing error ('T ool'); should be 'Tool-Augmented Reasoning.'
- [References] Reference [22] lists 'P. N. Russell' as the author of 'Artificial Intelligence: A Modern Approach'; the authors are Stuart Russell and Peter Norvig. Please correct.
- [§VI-C, Table IV] Notation is inconsistent: the table uses 'Rmem' in the notation column but 'R_mem' in the formula. Also 'Pmem' and 'Rmem' are not defined as to whether they are rates, counts, or scores.
- [Figure 2] The labels 'Rules Data Network Create Act Cognition' in Figure 2 are telegraphic and may confuse readers; consider adding short descriptors or a legend.
- [§IV, Table II] The 'Key Limitation' column entries are format-inconsistent (some are sentence fragments, some start with lowercase). A quick copyedit would improve readability.
Circularity Check
No equation-level circularity; the main caveat is a minor self-referential framing via the authors' prior work [27].
-
self citation load bearing
[Sec. II-A (Background: Cognitive Artificial Intelligence); Table I; Sec. V; reference [27]]
"Within this context, Cognitive Artificial Intelligence has emerged as a research direction focused on developing systems that exhibit persistent, adaptive, and self-regulating behavior. Cognitive AI seeks to integrate perception, memory, reasoning, learning, action, and metacognitive control within a unified feedback-driven framework [27], [26]."
The definition of the paper's central subject (Cognitive AI) and the capability checklist that motivates the five-dimension taxonomy are anchored to [27], a prior paper by three of the present authors (Golilarz, Penchala, Rahimi). The same reference is reused in Table I ('Cognitive AI[27]') and in the ACIA design. Thus the frame of the survey is partly self-referential: the field and its required capabilities are not independently established by an external consensus or a machine-checked result. However, the five-gap taxonomy itself is subsequently supported by a broad, mostly external literature, so this citation is a framing device rather than the load-bearing proof of the taxonomy.
full rationale
This is a survey/taxonomy paper, not an equation-level derivation: there are no fitted parameters renamed as predictions, no uniqueness theorems imported from the authors, and no ansatz smuggled in via citation. The five dimensions in Sec. III are an organizational claim derived from a narrative synthesis; the absence of a systematic search protocol is a selection-bias/external-validity concern, not circularity. Similarly, the ACIA 'addresses' claim in Sec. V and Table III maps components to gaps by design; because the paper does not use that mapping to validate the taxonomy or to make a quantitative prediction, it is not a circular derivation. The only noteworthy point is the self-referential definition of Cognitive AI via the same group's [27], which colors the framing but does not reduce the central taxonomy or the individual gap reviews to that citation. Hence a low score of 2.
Assumptions & free parameters
free parameters (1)
- Unspecified aggregation function f(·) in CPI/CAR/CCS metrics
assumptions (4)
- ad hoc to paper Narrative synthesis of the selected literature yields a representative and complete set of fundamental cognitive gaps
- ad hoc to paper Cognitive AI is a distinct and meaningful paradigm, as defined in authors' prior work [27]
- domain assumption Classical cognitive architectures (ACT-R, Soar, CLARION, etc.) provide the relevant foundation for modern cognitive AI
- domain assumption Current LLMs operate as reactive next-token predictors without stable latent state across interactions
invented entities (2)
-
Adaptive Cognitive Intelligence Architecture (ACIA)
-
Cognitive Persistence Index (CPI), Cognitive Adaptation Rate (CAR), Cognitive Consistency Score (CCS)
Cite this review
Pith. "Pith review of A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI." pith.science (2026). https://pith.science/paper/HDP3RQU7
@misc{pith2026260802553,
author = {Pith},
title = {Pith review of: A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/HDP3RQU7}},
note = {Machine review of arXiv:2608.02553}
}
read the original abstract
Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents a taxonomy-driven survey of the major cognitive capability gaps that continue to constrain the development of Cognitive AI. The literature is organized around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension, we review recent advances, identify recurring limitations, and discuss open research challenges. Building on these insights, we outline a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) and examine emerging directions in cognition-centric evaluation. The proposed taxonomy provides a unified framework for organizing existing research, identifying unresolved challenges, and guiding the design of future cognitively capable systems. Together, the taxonomy, architectural perspective, and evaluation framework offer a roadmap for advancing AI systems that exhibit more reliable long-term reasoning, adaptive decision-making, and continual learning. The survey highlights key research opportunities toward more adaptive, reliable, and cognitively capable AI systems, providing a foundation for future progress toward Cognitive AI and, ultimately, Artificial General Intelligence (AGI).
Figures
Reference graph
Works this paper leans on
-
[27]
Bridging the gap: Toward cognitive autonomy in artificial intelligence,
N. A. Golilarz, S. Penchala, and S. Rahimi, “Bridging the gap: Toward cognitive autonomy in artificial intelligence,”arXiv preprint arXiv:2512.02280, 2025
arXiv 2025
-
[26]
Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,
D. B. Acharya, K. Kuppan, and B. Divya, “Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,”IEEE Access, vol. 13, pp. 18 912–18 936, 2025
2025
-
[1]
Knowledge editing for large language models: A survey,
S. Wang, Y . Zhu, H. Liu, Z. Zheng, C. Chen, and J. Li, “Knowledge editing for large language models: A survey,”ACM Computing Surveys, vol. 57, no. 3, pp. 1–37, 2024
2024
-
[2]
Agentic large language models, a survey,
A. Plaat, M. Van Duijn, N. Van Stein, M. Preuss, P. Van der Putten, and K. J. Batenburg, “Agentic large language models, a survey,”Journal of Artificial Intelligence Research, vol. 84, Dec
-
[3]
Goal-driven auton- omy for cognitive systems,
M. Paisner, M. Cox, M. Maynord, and D. Perlis, “Goal-driven auton- omy for cognitive systems,” inProceedings of the Annual Meeting of the Cognitive Science Society, vol. 36, no. 36, 2014
2014
-
[4]
Belief revision: The adaptability of large language models reasoning,
B. Wilie, S. Cahyawijaya, E. Ishii, J. He, and P. Fung, “Belief revision: The adaptability of large language models reasoning,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 10 480–10 4...
2024
-
[5]
On the failure of latent state persistence in large language models,
J. tse Huang, K. Sun, W. Wang, and M. Dredze, “On the failure of latent state persistence in large language models,” 2026. [Online]. Available: https://arxiv.org/abs/2505.10571
arXiv 2026
-
[6]
Continual learning of large language models: A comprehensive survey,
H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y . Wang, Z. Wang, S. Ebrahimi, and H. Wang, “Continual learning of large language models: A comprehensive survey,”ACM Computing Surveys, vol. 58, no. 5, pp. 1–42, 2025
2025
Show all 111 references
-
[7]
Continual learning: overcoming catastrophic forgetting for adaptive ai systems,
H. Salwa, N. Burhan, and E. Rahel, “Continual learning: overcoming catastrophic forgetting for adaptive ai systems,”Authorea Preprints, 2025
2025
-
[8]
Alas: A stateful multi-llm agent framework for disruption-aware planning,
E. Y . Chang and L. Geng, “Alas: A stateful multi-llm agent framework for disruption-aware planning,” 2025. [Online]. Available: https://arxiv.org/abs/2505.12501
2025 arXiv
-
[9]
Adaptation and learning in ai agents,
S. K. K. Parimi, “Adaptation and learning in ai agents,”Authorea Preprints, 2025
2025
-
[10]
Con- tinual learning for large language models: A survey,
T. Wu, L. Luo, Y .-F. Li, S. Pan, T.-T. Vu, and G. Haffari, “Con- tinual learning for large language models: A survey,”arXiv preprint arXiv:2402.01364, 2024
2024 arXiv
-
[11]
Knowledgesmith: Uncovering knowledge updating in llms with model editing and unlearning,
Y . Luo, Z. Zhou, H. Chen, K. Qiu, M. Savvides, S. Li, and J. Wang, “Knowledgesmith: Uncovering knowledge updating in llms with model editing and unlearning,” 2025. [Online]. Available: https://arxiv.org/abs/2510.02392
2025
-
[12]
The illusion of insight in reasoning models,
L. G. d’Aliberti and M. H. Ribeiro, “The illusion of insight in reasoning models,”arXiv preprint arXiv:2601.00514, 2026
2026 arXiv
-
[13]
When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,
R. Kamoi, Y . Zhang, N. Zhang, J. Han, and R. Zhang, “When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 1417–1440, 2024. [Online]. Available: https://acla...
2024
-
[14]
When do llms admit their mistakes? understanding the role of model belief in retraction,
Y . Yang and R. Jia, “When do llms admit their mistakes? understanding the role of model belief in retraction,” 2026. [Online]. Available: https://arxiv.org/abs/2505.16170
2026
-
[15]
Arigraph: Learning knowledge graph world models with episodic memory for llm agents,
P. Anokhinet al., “Arigraph: Learning knowledge graph world models with episodic memory for llm agents,”arXiv preprint arXiv:2407.04363, 2025
2025 arXiv
-
[16]
Technical report: Evaluating goal drift in language model agents,
R. Arike, E. Donoway, H. Bartsch, and M. Hobbhahn, “Technical report: Evaluating goal drift in language model agents,” 2025. [Online]. Available: https://arxiv.org/abs/2505.02709
2025 arXiv
-
[17]
Towards mitigating LLM hallucination via self reflection,
Z. Ji, T. Yu, Y . Xu, N. Lee, E. Ishii, and P. Fung, “Towards mitigating LLM hallucination via self reflection,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics...
2023
-
[18]
Unlocking efficient, scalable, and continual knowledge editing with basis-level representation fine- tuning,
T. Liu, R. Li, Y . Qi, H. Liu, X. Tang, T. Zheng, Q. Yin, M. X. Cheng, J. Huan, H. Wang, and J. Gao, “Unlocking efficient, scalable, and continual knowledge editing with basis-level representation fine- tuning,” 2025. [Online]. Available: https://arxiv.org/abs/2503.00306
2025 arXiv
-
[19]
Learning to edit: Aligning LLMs with knowledge editing,
Y . Jiang, Y . Wang, C. Wu, W. Zhong, X. Zeng, J. Gao, L. Li, X. Jiang, L. Shang, R. Tang, Q. Liu, and W. Wang, “Learning to edit: Aligning LLMs with knowledge editing,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2024
-
[20]
Alleviating hallucinations from knowledge misalignment in large language models via selective abstention learning,
L. Huang, X. Feng, W. Ma, Y . Fan, X. Feng, Y . Gu, Y . Ye, L. Zhao, W. Zhong, B. Wang, D. Wu, G. Hu, L. Kong, T. Xiao, T. Liu, and B. Qin, “Alleviating hallucinations from knowledge misalignment in large language models via selective abstention learning,” inProceedings of the...
2025
-
[21]
A survey on uncertainty quantification methods for deep learning,
W. He, Z. Jiang, T. Xiao, Z. Xu, and Y . Li, “A survey on uncertainty quantification methods for deep learning,”ACM Comput. Surv., vol. 58, no. 7, Feb. 2026. [Online]. Available: https://doi.org/10.1145/3786319
2026 doi
-
[22]
Artificial intelligence: a modern approach by stuart,
P. N. Russell, “Artificial intelligence: a modern approach by stuart,” Russell and Peter Norvig contributing writers, Ernest Davis...[et al.], p. 22, 2010
2010
-
[23]
Early history of machine learning,
A. L. Fradkov, “Early history of machine learning,”IFAC- PapersOnLine, vol. 53, no. 2, pp. 1385–1390, 2020
2020
-
[24]
A survey on deep learning: Algorithms, techniques, and applications,
S. Pouyanfar, S. Sadiq, Y . Yan, H. Tian, Y . Tao, M. P. Reyes, M.- L. Shyu, S.-C. Chen, and S. S. Iyengar, “A survey on deep learning: Algorithms, techniques, and applications,”ACM computing surveys (CSUR), vol. 51, no. 5, pp. 1–36, 2018
2018
-
[25]
Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,
S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberget al., “Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,” 2023
2023
-
[28]
Artificial general intelligence: concept, state of the art, and future prospects,
B. Goertzel, “Artificial general intelligence: concept, state of the art, and future prospects,”Journal of Artificial General Intelligence, vol. 5, no. 1, pp. 1–48, 2014
2014
-
[29]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[30]
Generative agents: Interactive simulacra of human behavior,
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” 2023. [Online]. Available: https://arxiv.org/abs/2304.03442
2023 arXiv
-
[31]
Memory in the age of ai agents,
Y . Hu, S. Liu, Y . Yue, G. Zhang, B. Liu, F. Zhu, J. Lin, H. Guo, S. Dou, Z. Xiet al., “Memory in the age of ai agents,”arXiv preprint arXiv:2512.13564, 2025
2025 arXiv
-
[32]
Transformer-XL: Attentive language models beyond a fixed- length context,
Z. Dai, Z. Yang, Y . Yang, J. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed- length context,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L. M `arquez,...
2019
-
[33]
Compressive transformers for long-range sequence modelling,
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap, “Compressive transformers for long-range sequence modelling,” 2019. [Online]. Available: https://arxiv.org/abs/1911.05507
2019 arXiv
-
[34]
Memgpt: Towards llms as operating systems,
C. Packer, S. Wooders, K. Lin, V . Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez, “Memgpt: Towards llms as operating systems,” 2024. [Online]. Available: https://arxiv.org/abs/2310.08560
2024 arXiv
-
[35]
Memorybank: Enhancing large language models with long-term memory,
W. Zhong, L. Guo, Q. Gao, H. Ye, and Y . Wang, “Memorybank: Enhancing large language models with long-term memory,” inProceed- ings of the AAAI conference on artificial intelligence, vol. 38, no. 17, 2024, pp. 19 724–19 731
2024
-
[36]
Synapse: Empowering llm agents with episodic-semantic memory via spreading activation,
H. Jiang, J. Chen, Y . Pan, L. Chen, W. You, Y . Zhou, R. Zhang, A. Sikora, L. Zhao, Y . Abate, and T. Liu, “Synapse: Empowering llm agents with episodic-semantic memory via spreading activation,”
-
[37]
On the curse of memory in recurrent neural networks: Approximation and optimization analysis,
Z. Li, J. Han, W. E, and Q. Li, “On the curse of memory in recurrent neural networks: Approximation and optimization analysis,” 2024. [Online]. Available: https://arxiv.org/abs/2009.07799
2024 arXiv
-
[38]
To backtrack or not to backtrack: When sequential search limits model reasoning,
T. Qin, D. Alvarez-Melis, S. Jelassi, and E. Malach, “To backtrack or not to backtrack: When sequential search limits model reasoning,”
-
[39]
Large language models cannot self-correct reasoning yet,
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou, “Large language models cannot self-correct reasoning yet,”
-
[40]
Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning,
X. Chen, A. Zhao, H. Xia, X. Lu, H. Wang, Y . Chen, W. Zhang, J. Wang, W. Li, and X. Shen, “Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning,” 2025. [Online]. Available: https://arxiv.org/abs/2505.16782
2025
-
[41]
Available: https://arxiv.org/abs/2504.07052
[Online]. Available: https://arxiv.org/abs/2504.07052
-
[42]
Ctrls: Chain-of-thought reasoning via latent state-transition,
J. Wu, Y . Xiong, X. Li, S. Yu, Z. Hu, T. Yu, R. Wang, X. Chen, J. Shang, and J. McAuley, “Ctrls: Chain-of-thought reasoning via latent state-transition,” 2026. [Online]. Available: https://arxiv.org/abs/2507.08182
2026
-
[43]
On the relation of state space models and hidden markov models,
A. Ghojogh, M. H. Sepanj, and B. Ghojogh, “On the relation of state space models and hidden markov models,” 2026. [Online]. Available: https://arxiv.org/abs/2601.13357
2026
-
[44]
Unsupervised methods for subgoal discov- ery during intrinsic motivation in model-free hierarchical reinforcement learning
J. Rafati and D. C. Noelle, “Unsupervised methods for subgoal discov- ery during intrinsic motivation in model-free hierarchical reinforcement learning.” inKEG@ AAAI, 2019, pp. 17–25
2019
-
[45]
Training large language models to reason in a continuous latent space,
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y . Tian, “Training large language models to reason in a continuous latent space,” 2025. [Online]. Available: https://arxiv.org/abs/2412.06769
2025 arXiv
-
[46]
Plangenllms: A modern survey of llm planning capabilities,
H. Wei, Z. Zhang, S. He, T. Xia, S. Pan, and F. Liu, “Plangenllms: A modern survey of llm planning capabilities,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 19 497–19 521
2025
-
[47]
Task planning and decision-making methods for intelligent agents based on large language models,
X. Li, “Task planning and decision-making methods for intelligent agents based on large language models,” ser. AIIIP ’25. New York, NY , USA: Association for Computing Machinery, 2026, p. 817–822. [Online]. Available: https://doi.org/10.1145/3778534.3778661
2026
-
[48]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2201.11903
2023 arXiv
-
[49]
A brain-inspired agentic architecture to improve planning with LLMs,
T. Webb, S. S. Mondal, and I. Momennejad, “A brain-inspired agentic architecture to improve planning with LLMs,”Nat. Commun., vol. 16, no. 1, p. 8633, Sep. 2025
2025
-
[50]
Graph of thoughts: Solving elaborate problems with large language models,
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk et al., “Graph of thoughts: Solving elaborate problems with large language models,” inProceedings of the AAAI conference on artificial intelligence,...
2024
-
[51]
Hiplan: Hierarchical planning for llm-based agents with adaptive global-local guidance,
Z. Li, Y . Chang, G. Yu, and X. Le, “Hiplan: Hierarchical planning for llm-based agents with adaptive global-local guidance,”arXiv preprint arXiv:2508.19076, 2025
2025 arXiv
-
[52]
Adaplanner: Adaptive planning from feedback with language models,
H. Sun, Y . Zhuang, L. Kong, B. Dai, and C. Zhang, “Adaplanner: Adaptive planning from feedback with language models,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, In...
2023
-
[53]
Tree of thoughts: Deliberate problem solving with large language models,
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y . Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2305.10601
2023 arXiv
-
[54]
Self-refine: Iterative refinement with self-feedback,
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-refine: Iterative refinement with self-feedback,” 2023. [Online]. Available: ht...
2023 arXiv
-
[55]
Coherence-based alignment: A structural architecture for preventing goal drift in agentic ai systems
A. Abdi, “Coherence-based alignment: A structural architecture for preventing goal drift in agentic ai systems.”
-
[56]
Pepa: a persistently autonomous embodied agent with personalities,
K. Liu, Y . Li, L. Zhu, and W. Zhang, “Pepa: a persistently autonomous embodied agent with personalities,” 2026. [Online]. Available: https://arxiv.org/abs/2603.00117
2026 arXiv
-
[57]
Reflexion: Language agents with verbal reinforcement learning,
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023. [Online]. Available: https://arxiv.org/abs/2303.11366
2023 arXiv
-
[58]
Inherited goal drift: Contextual pressure can undermine agentic goals,
A. Menon, M. Saebo, T. Crosse, S. Gibson, E. Jang, and D. Cruz, “Inherited goal drift: Contextual pressure can undermine agentic goals,” 2026. [Online]. Available: https://arxiv.org/abs/2603.03258
2026
-
[59]
The metacognitive demands and opportunities of generative ai,
L. Tankelevitch, V . Kewenig, A. Simkute, A. E. Scott, A. Sarkar, A. Sellen, and S. Rintel, “The metacognitive demands and opportunities of generative ai,” inProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, ser. CHI ’24. New York, NY , USA: Associa...
2024
-
[60]
From human to model overconfidence: Evaluating confidence dynamics in large language models,
B. Wen, C. Xu, B. HAN, R. Wolfe, L. L. Wang, and B. Howe, “From human to model overconfidence: Evaluating confidence dynamics in large language models,” inNeurIPS 2024 Workshop on Behavioral Machine Learning, 2024
2024
-
[61]
Measuring goal-directedness,
M. MacDermott, J. Fox, F. Belardinelli, and T. Everitt, “Measuring goal-directedness,” inAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 11 4...
2024
-
[62]
Do i really know? learning factual self-verification for hallucination reduction,
E. Altinisik, M. Fatehkia, F. Deniz, N. Durrani, M. Hawasly, M. Raza, and H. T. Sencar, “Do i really know? learning factual self-verification for hallucination reduction,” 2026. [Online]. Available: https://arxiv.org/abs/2602.02018 14
2026
-
[63]
ProgCo: Program helps self-correction of large language models,
X. Song, Y . Wu, W. Wang, J. Liu, W. Su, and B. Zheng, “ProgCo: Program helps self-correction of large language models,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), W. Che, J. Nabende, E. Shutova, and M. T...
2025
-
[64]
Uncertainty quantification and confidence calibration in large language models: A survey,
X. Liu, T. Chen, L. Da, C. Chen, Z. Lin, and H. Wei, “Uncertainty quantification and confidence calibration in large language models: A survey,” ser. KDD ’25. New York, NY , USA: Association for Computing Machinery, 2025, p. 6107–6117. [Online]. Available: https://doi.org/10.1...
2025
-
[65]
Large language models can self-correct with minimal effort,
Z. Wu, Q. Zeng, Z. Zhang, Z. Tan, C. Shen, and M. Jiang, “Large language models can self-correct with minimal effort,” inAI for Math Workshop@ ICML 2024, 2024
2024
-
[66]
Out-of-distribution detection with positive and negative prompt super- vision using large language models,
Z. He, C. Zhao, M. Shao, X. Wu, X. Zhao, D. Li, Q. Tian, and L. Yu, “Out-of-distribution detection with positive and negative prompt super- vision using large language models,”arXiv preprint arXiv:2511.10923, 2025
2025
-
[67]
Revisiting epistemic markers in confidence estimation: Can markers accurately reflect large language models’ uncertainty?
J. Liu, Q. Zong, W. Wang, and Y . Song, “Revisiting epistemic markers in confidence estimation: Can markers accurately reflect large language models’ uncertainty?” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers...
2025
-
[68]
Know your limits: A survey of abstention in large language models,
B. Wen, J. Yao, S. Feng, C. Xu, Y . Tsvetkov, B. Howe, and L. L. Wang, “Know your limits: A survey of abstention in large language models,” Transactions of the Association for Computational Linguistics, vol. 13, pp. 529–556, 2025
2025
-
[69]
Envisioning outlier exposure by large language models for out-of-distribution detection,
C. Cao, Z. Zhong, Z. Zhou, Y . Liu, T. Liu, and B. Han, “Envisioning outlier exposure by large language models for out-of-distribution detection,” 2024. [Online]. Available: https://arxiv.org/abs/2406.00806
2024 arXiv
-
[70]
Risan: Robust instance specific abstention network,
B. Kalra, K. Shah, and N. Manwani, “Risan: Robust instance specific abstention network,”arXiv preprint arXiv:2107.03090, 2021
2021 arXiv
-
[71]
R-tuning: Instructing large language models to say ‘i don’t know’,
H. Zhang, S. Diao, Y . Lin, Y . Fung, Q. Lian, X. Wang, Y . Chen, H. Ji, and T. Zhang, “R-tuning: Instructing large language models to say ‘i don’t know’,” inProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Huma...
2024
-
[72]
Halluci- nate less by thinking more: Aspect-based causal abstention for large language models,
V . Nguyen, Z. Xu, J. Chan, E. He, F. Xia, and X. Zhang, “Halluci- nate less by thinking more: Aspect-based causal abstention for large language models,”arXiv preprint arXiv:2511.17170, 2025
2025
-
[73]
Selectivenet: A deep neural network with an integrated reject option,
Y . Geifman and R. El-Yaniv, “Selectivenet: A deep neural network with an integrated reject option,” inInternational conference on machine learning. PMLR, 2019, pp. 2151–2159
2019
-
[74]
Language agents meet causality – bridging llms and causal world models,
J. Gkountouras, M. Lindemann, P. Lippe, E. Gavves, and I. Titov, “Language agents meet causality – bridging llms and causal world models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.19923
2024 arXiv
-
[75]
Language-guided world models: A model-based approach to AI control,
A. Zhang, K. Nguyen, J. Tuyls, A. Lin, and K. Narasimhan, “Language-guided world models: A model-based approach to AI control,” inProceedings of the 4th Workshop on Spatial Language Understanding and Grounded Communication for Robotics (SpLU-RoboNLP 2024), P. Kordjamshidi, X. ...
2024
-
[76]
Causalarc: Abstract reasoning with causal world models,
J. Maasch, J. Kalantari, and K. Khezeli, “Causalarc: Abstract reasoning with causal world models,” 2026. [Online]. Available: https://arxiv.org/abs/2509.03636
2026
-
[77]
Beyond world models: Rethinking understanding in ai models,
T. Gupta and D. Pruthi, “Beyond world models: Rethinking understanding in ai models,” 2025. [Online]. Available: https://arxiv.org/abs/2511.12239
2025
-
[78]
Recthinker: An agentic framework for tool-augmented reasoning in recommendation,
H. Zhang, Y . Zhu, K. Mao, T. Li, and Z. Dou, “Recthinker: An agentic framework for tool-augmented reasoning in recommendation,”
-
[79]
Evaluating and improving tool-augmented computation-intensive math reasoning,
B. Zhang, K. Zhou, X. Wei, X. Zhao, J. Sha, S. Wang, and J.-R. Wen, “Evaluating and improving tool-augmented computation-intensive math reasoning,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., v...
2023
-
[80]
Grounding large language models for robot task planning using closed-loop state feedback,
V . Bhat, A. U. Kaypak, P. Krishnamurthy, R. Karri, and F. Khorrami, “Grounding large language models for robot task planning using closed-loop state feedback,”Advanced Robotics Research, Nov. 2025. [Online]. Available: http://dx.doi.org/10.1002/adrr.202500072
2025 doi
-
[81]
ChatCoT: Tool-augmented chain-of-thought reasoning on chat- based large language models,
Z. Chen, K. Zhou, B. Zhang, Z. Gong, X. Zhao, and J.-R. Wen, “ChatCoT: Tool-augmented chain-of-thought reasoning on chat- based large language models,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: As...
2023
-
[82]
Ai for closed- loop control systems: New opportunities for modeling, designing, and tuning control systems,
J. Sch ¨oning, A. Riechmann, and H.-J. Pfisterer, “Ai for closed- loop control systems: New opportunities for modeling, designing, and tuning control systems,” inProceedings of the 2022 14th International Conference on Machine Learning and Computing, ser. ICMLC ’22. New York, ...
2022
-
[83]
Available: https://arxiv.org/abs/2603.09843
[Online]. Available: https://arxiv.org/abs/2603.09843
-
[84]
Building self-evolving agents via experience- driven lifelong learning: A framework and benchmark,
Y . Cai, Y . Hao, J. Zhou, H. Yan, Z. Lei, R. Zhen, Z. Han, Y . Yang, J. Li, Q. Panet al., “Building self-evolving agents via experience- driven lifelong learning: A framework and benchmark,”arXiv preprint arXiv:2508.19005, 2025
2025
-
[85]
Deep reinforcement learning for robotics: A survey of real- world successes,
C. Tang, B. Abbatematteo, J. Hu, R. Chandra, R. Mart ´ın-Mart´ın, and P. Stone, “Deep reinforcement learning for robotics: A survey of real- world successes,”Annual Review of Control, Robotics, and Autonomous Systems, vol. 8, no. 1, pp. 153–188, 2025
2025
-
[86]
A metacognitive architecture for cor- recting llm errors in ai agents,
J. Kim, M. Islam, and A. Goel, “A metacognitive architecture for cor- recting llm errors in ai agents,” inProceedings of the AAAI Conference on Artificial Intelligence, 2026
2026
-
[87]
Contrastive test-time adaptation,
D. Chen, D. Wang, T. Darrell, and S. Ebrahimi, “Contrastive test-time adaptation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 295–305
2022
-
[88]
How human–ai feedback loops alter human perceptual, emotional and social judgements,
M. Glickman and T. Sharot, “How human–ai feedback loops alter human perceptual, emotional and social judgements,”Nature Human Behaviour, vol. 9, no. 2, pp. 345–359, 2025
2025
-
[89]
Mitigating the alignment tax of rlhf,
Y . Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y . Yao, and T. Zhang, “Mitigating the alignment tax of rlhf,” 2024. [Online]. Available: https://arxiv.org/abs/2309.06256
2024 arXiv
-
[90]
Analyzing and reducing catastrophic forgetting in parameter efficient tuning,
W. Ren, X. Li, L. Wang, T. Zhao, and W. Qin, “Analyzing and reducing catastrophic forgetting in parameter efficient tuning,” 2024. [Online]. Available: https://arxiv.org/abs/2402.18865
2024 arXiv
-
[91]
A proximal policy optimization-based reinforcement learning framework for real-time personalized endurance training,
C. Wu, D. Liang, B. Yang, L. Xu, and Y . Li, “A proximal policy optimization-based reinforcement learning framework for real-time personalized endurance training,”Informatica, vol. 50, no. 8, 2026
2026
-
[92]
Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation,
M. Fawi, “Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation,”arXiv preprint arXiv:2408.14572, 2024
2024 arXiv
-
[93]
Reasoning as meta-learning: An optimization perspective to decipher long cot reasoning in llms
J. Liu, H. Liu, L. Xiao, S. Liu, T. Zhang, Z. Ma, S. Zhang, and K. Chen, “Reasoning as meta-learning: An optimization perspective to decipher long cot reasoning in llms.”
-
[94]
Locating and editing factual associations in gpt,
K. Meng, D. Bau, A. Andonian, and Y . Belinkov, “Locating and editing factual associations in gpt,”Advances in neural information processing systems, vol. 35, pp. 17 359–17 372, 2022
2022
-
[95]
Mass-editing memory in a transformer,
K. Meng, A. S. Sharma, A. Andonian, Y . Belinkov, and D. Bau, “Mass-editing memory in a transformer,” 2023. [Online]. Available: https://arxiv.org/abs/2210.07229
2023 arXiv
-
[96]
Overcoming catastrophic forgetting in neural networks,
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al., “Overcoming catastrophic forgetting in neural networks,”Pro- ceedings of the national academy of sciences, vol. 114, no. 13, pp. 3521–3526, 2017
2017
-
[97]
Unveiling the pitfalls of knowledge editing for large language models,
Z. Li, N. Zhang, Y . Yao, M. Wang, X. Chen, and H. Chen, “Unveiling the pitfalls of knowledge editing for large language models,”arXiv preprint arXiv:2310.02129, 2023
2023 arXiv
-
[98]
Eval- uating the ripple effects of knowledge editing in language models,
R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva, “Eval- uating the ripple effects of knowledge editing in language models,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 283–298, 2024
2024
-
[99]
Think-in-memory: Recalling and post-thinking enable llms with long-term memory,
L. Liu, X. Yang, Y . Shen, B. Hu, Z. Zhang, J. Gu, and G. Zhang, “Think-in-memory: Recalling and post-thinking enable llms with long-term memory,” 2023. [Online]. Available: https://arxiv.org/abs/2311.08719
2023 arXiv
-
[100]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Rayet al., “Training language models to follow instructions with human feedback,”Advances in neural information processing systems, vol. 35, pp. 27 730–27 744, 2022. 15
2022
-
[101]
Alphaedit: Null-space constrained knowledge editing for language models,
J. Fang, H. Jiang, K. Wang, Y . Ma, S. Jie, X. Wang, X. He, and T.-S. Chua, “Alphaedit: Null-space constrained knowledge editing for language models,”arXiv preprint arXiv:2410.02355, 2024
2024 arXiv
-
[102]
Holistic evaluation of language models,
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,”arXiv preprint arXiv:2211.09110, 2022
2022 arXiv
-
[103]
40 years of cognitive architectures: core cognitive abilities and practical applications,
I. Kotseruba and J. K. Tsotsos, “40 years of cognitive architectures: core cognitive abilities and practical applications,”Artificial Intelli- gence Review, vol. 53, no. 1, pp. 17–94, 2020
2020
-
[104]
Survey of hallucination in natural language generation,
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y . Xu, E. Ishii, Y . J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”ACM computing surveys, vol. 55, no. 12, pp. 1–38, 2023
2023
-
[105]
Truthfulqa: Measuring how models mimic human falsehoods,
S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,” inProceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), 2022, pp. 3214–3252
2022
-
[106]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”Transactions on machine learning research, 2023
2023
-
[108]
Judging llm-as-a-judge with mt- bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt- bench and chatbot arena,”Advances in neural information processing systems, vol. 36, pp. 46 595–46 623, 2023
2023
-
[111]
Red teaming language models with language models,
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red teaming language models with language models,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3419–3448
2022
-
[2024]
Available: https://arxiv.org/abs/2310.01798
[Online]. Available: https://arxiv.org/abs/2310.01798
-
[2025]
Available: http://dx.doi.org/10.1613/jair.1.18675
[Online]. Available: http://dx.doi.org/10.1613/jair.1.18675
-
[2026]
Available: https://arxiv.org/abs/2601.02744
[Online]. Available: https://arxiv.org/abs/2601.02744
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.