Pith. sign in

REVIEW 4 major objections 5 minor 111 references

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read A five-part taxonomy organizes the cognitive capability gaps that keep generative and agentic AI from sustained, reliable cognition.

desk verdict A useful organizing frame for a fragmented literature, but the taxonomy is a plausibility argument, not a discovered fact. read the letter →

arxiv 2608.02553 v1 pith:HDP3RQU7 submitted 2026-08-03 cs.AI

classification cs.AI
keywords CognitiveAItaxonomycapabilitygapspersistentstatemodelinggoal-directedautonomyself-monitoringenvironmentinteractionlearningandadaptation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that the scattered limitations of generative and agentic AI—weak memory, goal drift, brittle self-correction, poor environmental grounding, and catastrophic forgetting—are not separate problems but symptoms of five underlying cognitive capability gaps: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. If the taxonomy holds, researchers gain a shared map for comparing work across AI, cognitive science, and robotics, and a checklist for what a system needs before it can be trusted over long horizons. The paper also proposes a closed-loop architecture, ACIA, that pairs each component with one of the gaps, and argues that evaluation should shift from one-shot benchmarks to longitudinal measures of memory, consistency, and adaptation.

What carries the argument

The taxonomy itself is the central organizing identity: five cognitive components with their subcapability gaps, derived via a narrative synthesis of the literature. The second piece of machinery is ACIA, a closed-loop cognitive architecture whose six components map one-to-one onto the five taxonomy dimensions (perception/attention and action both feed environment interaction). The paper's third mechanism is cognition-centric evaluation, illustrated by three named metrics—Cognitive Persistence Index, Cognitive Adaptation Rate, and Cognitive Consistency Score—that measure whether a system retains, adapts, and stays coherent over time.

What would settle it

A systematic literature review that codes the same research areas with explicit, prespecified inclusion criteria and finds additional fundamental dimensions beyond these five—or fails to reproduce them—would directly refute the taxonomy's claim to completeness.

Watch

Extended reading notes

Core claim

The central claim is that five recurring dimensions—persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation—organize the major cognitive capability gaps in modern generative and agentic AI systems. Each dimension is decomposed into concrete subcomponents (e.g., persistent memory, state revision, latent state modeling; goal formulation, planning, goal persistence; metacognitive monitoring, uncertainty detection, abstention; world modeling, tool-augmented reasoning, feedback loops; policy adaptation, continual learning, safe knowledge updates). The paper argues that current systems fail not on any single benchmark but

Load-bearing premise

The taxonomy and ACIA stand on the assumption that a narrative synthesis of the authors' chosen literature reveals the five fundamental, recurring dimensions—if that literature selection is not representative of the field, the unified framework may be an artifact of the selection rather than a property of the field.

Editorial extensions

If this is right

  • If the taxonomy is right, fragmented research threads in memory, planning, metacognition, grounding, and continual learning can be compared and combined under a single set of capability headings.
  • ACIA provides a blueprint for what a cognitive AI system must contain—persistent state, goal management, metacognitive control, environment coupling, and learning—and implies that scaling models alone cannot close the gaps.
  • The proposed evaluation metrics shift assessment from one-shot accuracy to longitudinal behavior: does the system keep memory, adapt without forgetting, and avoid contradictions over extended interaction?
  • The paper's framing suggests that agentic AI systems currently excel at short-horizon tasks while systematically missing the persistent internal state and self-regulation needed for long-horizon autonomy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The five dimensions may imply a developmental ordering—persistent state before goal stability, goal stability before meaningful metacognition—that the paper does not explicitly assert but that could be tested empirically.
  • Because the taxonomy rests on a narrative literature synthesis, a larger systematic review with explicit inclusion criteria might confirm, revise, or extend the dimension set; that test is a natural next step.
  • The concrete metrics the paper sketches (CPI, CAR, CCS) are presented as illustrative functions; a testable extension would be to instantiate them with specific formulas and validate them against existing benchmark failures.
  • ACIA is conceptual and unbuilt; its strongest testable claim is that each component addresses a specific gap, so a partial implementation could directly examine whether adding metacognitive control or persistent memory actually reduces observed goal drift and inconsistency.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a taxonomy of cognitive capability gaps in modern generative and agentic AI, organizing the literature around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension it reviews representative approaches, identifies recurring limitations, and distills open problems. The paper then outlines a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) whose components are mapped to the five gaps, and closes with a discussion of evaluation strategies, including illustrative cognition-centric metrics. The intended contribution is a unified organizing framework and roadmap for research toward Cognitive AI.

Significance. If the taxonomy is accepted as a faithful organization of the field, it would provide a useful shared vocabulary for comparing work on memory, planning, metacognition, grounding, and continual learning. The survey is broad and generally accurate in its characterization of current limitations, and it is explicit that ACIA is conceptual rather than implemented. The main value lies in the synthesis and in the articulation of evaluation challenges. However, the paper's central claims — that the five dimensions are the recurring fundamental gaps, that ACIA addresses them, and that the proposed metrics are useful — are currently not supported at the level the contributions claim. The strengths are the breadth of coverage, the clear graphical summaries (Figures 1 and 3, Table II), and the candid acknowledgment that many proposed mechanisms remain partial.

major comments (4)
  1. [§III, 'Survey Methodology and Taxonomy Overview'] The load-bearing step is the sentence 'Through this synthesis, five recurring dimensions of cognitive capability emerged across the surveyed literature.' No search strategy, inclusion criteria, coding protocol, inter-rater reliability, or comparison with competing taxonomies is provided. The paper calls the method a 'narrative synthesis approach,' which permits the selection of literature to be shaped by the authors' prior frame; notably, reference [27] is the same group's earlier 'cognitive autonomy' paper that already names similar dimensions. This creates a risk that the five dimensions are an artifact of the selection rather than a property of the field. Please either specify the methodology (corpus, inclusion criteria, coding, inter-rater checks) or reframe the contribution as a perspective/proposal rather than a survey-derived taxonomy. A comparison with at least one existing cogni
  2. [§V, Table III (ACIA mapping)] The claim that ACIA 'addresses' the five identified gaps is weakened by the fact that the mapping in Table III is one-to-one by construction: Memory is assigned to Persistent State Modeling, Metacognition to Self-Monitoring and Control, etc., using the same names as the gaps. Since no architectural details, formal specification, or empirical evidence are provided, the statement 'These components address the five cognitive capability gaps identified throughout this survey' is close to a restatement of the taxonomy rather than an independently motivated design. If the goal is a conceptual architecture, please state explicitly what distinguishes ACIA from a restatement of the taxonomy and provide at least a high-level mechanism for each component (e.g., how memory is updated, how metacognition triggers replanning). If the goal is a proposal, soften the 'addresses' claim to 'is designed to a
  3. [§VI-C, Table IV] The proposed metrics are not operationalizable as written. CPI is defined as f(T_ret, P_mem, R_mem), CAR as f(A_succ, A_opp), and CCS as f(1 - N_contra/N_reason), with the note that f(·) 'may vary across applications.' This means the paper provides no concrete metric, and the table is an illustration rather than a proposed evaluation framework. Since 'examining emerging directions in cognition-centric evaluation' and 'discuss cognition-centric evaluation strategies' are listed as contributions, please either provide concrete formulations (with at least one worked example or reference to an existing benchmark that instantiates them) or explicitly label this section as an open research agenda, not a metric proposal.
  4. [§II-A, Cognitive Artificial Intelligence] The paper defines Cognitive AI and its capability requirements largely through reference [27], which is authored by overlapping authors. This is a circularity risk: the survey's organizing frame is anchored to the same group's prior work without independent grounding. For example, the statement 'Cognitive AI seeks to integrate perception, memory, reasoning, learning, action, and metacognitive control within a unified feedback-driven framework [27], [26]' relies on [27] for the definition while [27] itself is the source of the paradigm. Please ground the definition in a broader set of sources (e.g., cognitive architectures, AGI safety, neuroscience-inspired AI) and explicitly discuss how [27] relates to the present survey.
minor comments (5)
  1. [§IV-D, heading] The heading 'T ool-Augmented Reasoning' contains a typographical spacing error ('T ool'); should be 'Tool-Augmented Reasoning.'
  2. [References] Reference [22] lists 'P. N. Russell' as the author of 'Artificial Intelligence: A Modern Approach'; the authors are Stuart Russell and Peter Norvig. Please correct.
  3. [§VI-C, Table IV] Notation is inconsistent: the table uses 'Rmem' in the notation column but 'R_mem' in the formula. Also 'Pmem' and 'Rmem' are not defined as to whether they are rates, counts, or scores.
  4. [Figure 2] The labels 'Rules Data Network Create Act Cognition' in Figure 2 are telegraphic and may confuse readers; consider adding short descriptors or a legend.
  5. [§IV, Table II] The 'Key Limitation' column entries are format-inconsistent (some are sentence fragments, some start with lowercase). A quick copyedit would improve readability.

Circularity Check

1 steps flagged · score 2.0 of 10

No equation-level circularity; the main caveat is a minor self-referential framing via the authors' prior work [27].

  1. self citation load bearing [Sec. II-A (Background: Cognitive Artificial Intelligence); Table I; Sec. V; reference [27]]
    "Within this context, Cognitive Artificial Intelligence has emerged as a research direction focused on developing systems that exhibit persistent, adaptive, and self-regulating behavior. Cognitive AI seeks to integrate perception, memory, reasoning, learning, action, and metacognitive control within a unified feedback-driven framework [27], [26]."

    The definition of the paper's central subject (Cognitive AI) and the capability checklist that motivates the five-dimension taxonomy are anchored to [27], a prior paper by three of the present authors (Golilarz, Penchala, Rahimi). The same reference is reused in Table I ('Cognitive AI[27]') and in the ACIA design. Thus the frame of the survey is partly self-referential: the field and its required capabilities are not independently established by an external consensus or a machine-checked result. However, the five-gap taxonomy itself is subsequently supported by a broad, mostly external literature, so this citation is a framing device rather than the load-bearing proof of the taxonomy.

full rationale

This is a survey/taxonomy paper, not an equation-level derivation: there are no fitted parameters renamed as predictions, no uniqueness theorems imported from the authors, and no ansatz smuggled in via citation. The five dimensions in Sec. III are an organizational claim derived from a narrative synthesis; the absence of a systematic search protocol is a selection-bias/external-validity concern, not circularity. Similarly, the ACIA 'addresses' claim in Sec. V and Table III maps components to gaps by design; because the paper does not use that mapping to validate the taxonomy or to make a quantitative prediction, it is not a circular derivation. The only noteworthy point is the self-referential definition of Cognitive AI via the same group's [27], which colors the framing but does not reduce the central taxonomy or the individual gap reviews to that citation. Hence a low score of 2.

Assumptions & free parameters 1 free parameters · 4 assumptions · 2 invented entities

The paper's central claims rest on the choice of five dimensions and on the ACIA blueprint, both of which are asserted through narrative synthesis rather than demonstrated. The only numeric objects—CPI, CAR, CCS—are placeholders forms of unspecified f(·). The cognitive-AI framing is imported from the authors' own prior work.

free parameters (1)
  • Unspecified aggregation function f(·) in CPI/CAR/CCS metrics
    Table IV defines CPI=f(Tret,Pmem,Rmem), CAR=f(Asucc,Aopp), CCS=f(1−Ncontra/Nreason), with f left unspecified; any numerical evaluation of these metrics depends on an unstated choice.
assumptions (4)
  • ad hoc to paper Narrative synthesis of the selected literature yields a representative and complete set of fundamental cognitive gaps
    Section III describes narrative synthesis with no search protocol or inclusion criteria; the five dimensions are asserted to be the recurring fundamental limitations.
  • ad hoc to paper Cognitive AI is a distinct and meaningful paradigm, as defined in authors' prior work [27]
    Section II-A defines Cognitive AI by citing [27] (overlapping authors) and treats it as the next stage; this is a stipulated definition rather than an externally established consensus.
  • domain assumption Classical cognitive architectures (ACT-R, Soar, CLARION, etc.) provide the relevant foundation for modern cognitive AI
    Section V relies on [98] to justify ACIA's components; no argument shows these architectures are the right starting point for LLM-era systems.
  • domain assumption Current LLMs operate as reactive next-token predictors without stable latent state across interactions
    Used throughout Sections IV-A and IV-B to motivate the gaps; supported by cited works but treated as established fact.
invented entities (2)
  • Adaptive Cognitive Intelligence Architecture (ACIA)
    purpose: Conceptual closed-loop framework integrating perception, memory, reasoning, metacognition, action, and learning to address the five gaps
    Proposed in Section V/Fig. 4; no implementation, system, or falsifiable prediction is given; mapping to gaps is asserted in Table III.
  • Cognitive Persistence Index (CPI), Cognitive Adaptation Rate (CAR), Cognitive Consistency Score (CCS)
    purpose: Illustrative metrics intended to evaluate long-horizon cognitive behavior
    Table IV defines them only up to unspecified f(·); no validation, benchmark, or data demonstrate they measure the intended constructs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI." pith.science (2026). https://pith.science/paper/HDP3RQU7

@misc{pith2026260802553,
  author       = {Pith},
  title        = {Pith review of: A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HDP3RQU7}},
  note         = {Machine review of arXiv:2608.02553}
}
read the original abstract

Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents a taxonomy-driven survey of the major cognitive capability gaps that continue to constrain the development of Cognitive AI. The literature is organized around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension, we review recent advances, identify recurring limitations, and discuss open research challenges. Building on these insights, we outline a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) and examine emerging directions in cognition-centric evaluation. The proposed taxonomy provides a unified framework for organizing existing research, identifying unresolved challenges, and guiding the design of future cognitively capable systems. Together, the taxonomy, architectural perspective, and evaluation framework offer a roadmap for advancing AI systems that exhibit more reliable long-term reasoning, adaptive decision-making, and continual learning. The survey highlights key research opportunities toward more adaptive, reliable, and cognitively capable AI systems, providing a foundation for future progress toward Cognitive AI and, ultimately, Artificial General Intelligence (AGI).

Figures

Figures reproduced from arXiv: 2608.02553 by the authors.

Figure 1
Figure 1. Conceptual evolution from Generative AI to Agentic AI and Cognitive AI, highlighting the additional cognitive capabilities introduced at each stage [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Evolution of artificial intelligence paradigms from rule-based systems to Cognitive AI. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy of core cognitive capability gaps in modern AI systems, organized into five cognitive components: persistent state modeling, goal-directed [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Proposed Adaptive Cognitive Intelligence Architecture (ACIA). The framework models a closed-loop cognitive system consisting of Perception and [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

111 extracted references · 1 canonical work pages

  1. [27]

    Bridging the gap: Toward cognitive autonomy in artificial intelligence,

    N. A. Golilarz, S. Penchala, and S. Rahimi, “Bridging the gap: Toward cognitive autonomy in artificial intelligence,”arXiv preprint arXiv:2512.02280, 2025

  2. [26]

    Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,

    D. B. Acharya, K. Kuppan, and B. Divya, “Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,”IEEE Access, vol. 13, pp. 18 912–18 936, 2025

  3. [1]

    Knowledge editing for large language models: A survey,

    S. Wang, Y . Zhu, H. Liu, Z. Zheng, C. Chen, and J. Li, “Knowledge editing for large language models: A survey,”ACM Computing Surveys, vol. 57, no. 3, pp. 1–37, 2024

  4. [2]

    Agentic large language models, a survey,

    A. Plaat, M. Van Duijn, N. Van Stein, M. Preuss, P. Van der Putten, and K. J. Batenburg, “Agentic large language models, a survey,”Journal of Artificial Intelligence Research, vol. 84, Dec

  5. [3]

    Goal-driven auton- omy for cognitive systems,

    M. Paisner, M. Cox, M. Maynord, and D. Perlis, “Goal-driven auton- omy for cognitive systems,” inProceedings of the Annual Meeting of the Cognitive Science Society, vol. 36, no. 36, 2014

  6. [4]

    Belief revision: The adaptability of large language models reasoning,

    B. Wilie, S. Cahyawijaya, E. Ishii, J. He, and P. Fung, “Belief revision: The adaptability of large language models reasoning,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 10 480–10 4...

  7. [5]

    On the failure of latent state persistence in large language models,

    J. tse Huang, K. Sun, W. Wang, and M. Dredze, “On the failure of latent state persistence in large language models,” 2026. [Online]. Available: https://arxiv.org/abs/2505.10571

  8. [6]

    Continual learning of large language models: A comprehensive survey,

    H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y . Wang, Z. Wang, S. Ebrahimi, and H. Wang, “Continual learning of large language models: A comprehensive survey,”ACM Computing Surveys, vol. 58, no. 5, pp. 1–42, 2025

Show all 111 references
  1. [7]

    Continual learning: overcoming catastrophic forgetting for adaptive ai systems,

    H. Salwa, N. Burhan, and E. Rahel, “Continual learning: overcoming catastrophic forgetting for adaptive ai systems,”Authorea Preprints, 2025

  2. [8]

    Alas: A stateful multi-llm agent framework for disruption-aware planning,

    E. Y . Chang and L. Geng, “Alas: A stateful multi-llm agent framework for disruption-aware planning,” 2025. [Online]. Available: https://arxiv.org/abs/2505.12501

  3. [9]

    Adaptation and learning in ai agents,

    S. K. K. Parimi, “Adaptation and learning in ai agents,”Authorea Preprints, 2025

  4. [10]

    Con- tinual learning for large language models: A survey,

    T. Wu, L. Luo, Y .-F. Li, S. Pan, T.-T. Vu, and G. Haffari, “Con- tinual learning for large language models: A survey,”arXiv preprint arXiv:2402.01364, 2024

  5. [11]

    Knowledgesmith: Uncovering knowledge updating in llms with model editing and unlearning,

    Y . Luo, Z. Zhou, H. Chen, K. Qiu, M. Savvides, S. Li, and J. Wang, “Knowledgesmith: Uncovering knowledge updating in llms with model editing and unlearning,” 2025. [Online]. Available: https://arxiv.org/abs/2510.02392

  6. [12]

    The illusion of insight in reasoning models,

    L. G. d’Aliberti and M. H. Ribeiro, “The illusion of insight in reasoning models,”arXiv preprint arXiv:2601.00514, 2026

  7. [13]

    When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,

    R. Kamoi, Y . Zhang, N. Zhang, J. Han, and R. Zhang, “When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 1417–1440, 2024. [Online]. Available: https://acla...

  8. [14]

    When do llms admit their mistakes? understanding the role of model belief in retraction,

    Y . Yang and R. Jia, “When do llms admit their mistakes? understanding the role of model belief in retraction,” 2026. [Online]. Available: https://arxiv.org/abs/2505.16170

  9. [15]

    Arigraph: Learning knowledge graph world models with episodic memory for llm agents,

    P. Anokhinet al., “Arigraph: Learning knowledge graph world models with episodic memory for llm agents,”arXiv preprint arXiv:2407.04363, 2025

  10. [16]

    Technical report: Evaluating goal drift in language model agents,

    R. Arike, E. Donoway, H. Bartsch, and M. Hobbhahn, “Technical report: Evaluating goal drift in language model agents,” 2025. [Online]. Available: https://arxiv.org/abs/2505.02709

  11. [17]

    Towards mitigating LLM hallucination via self reflection,

    Z. Ji, T. Yu, Y . Xu, N. Lee, E. Ishii, and P. Fung, “Towards mitigating LLM hallucination via self reflection,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics...

  12. [18]

    Unlocking efficient, scalable, and continual knowledge editing with basis-level representation fine- tuning,

    T. Liu, R. Li, Y . Qi, H. Liu, X. Tang, T. Zheng, Q. Yin, M. X. Cheng, J. Huan, H. Wang, and J. Gao, “Unlocking efficient, scalable, and continual knowledge editing with basis-level representation fine- tuning,” 2025. [Online]. Available: https://arxiv.org/abs/2503.00306

  13. [19]

    Learning to edit: Aligning LLMs with knowledge editing,

    Y . Jiang, Y . Wang, C. Wu, W. Zhong, X. Zeng, J. Gao, L. Li, X. Jiang, L. Shang, R. Tang, Q. Liu, and W. Wang, “Learning to edit: Aligning LLMs with knowledge editing,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  14. [20]

    Alleviating hallucinations from knowledge misalignment in large language models via selective abstention learning,

    L. Huang, X. Feng, W. Ma, Y . Fan, X. Feng, Y . Gu, Y . Ye, L. Zhao, W. Zhong, B. Wang, D. Wu, G. Hu, L. Kong, T. Xiao, T. Liu, and B. Qin, “Alleviating hallucinations from knowledge misalignment in large language models via selective abstention learning,” inProceedings of the...

  15. [21]

    A survey on uncertainty quantification methods for deep learning,

    W. He, Z. Jiang, T. Xiao, Z. Xu, and Y . Li, “A survey on uncertainty quantification methods for deep learning,”ACM Comput. Surv., vol. 58, no. 7, Feb. 2026. [Online]. Available: https://doi.org/10.1145/3786319

  16. [22]

    Artificial intelligence: a modern approach by stuart,

    P. N. Russell, “Artificial intelligence: a modern approach by stuart,” Russell and Peter Norvig contributing writers, Ernest Davis...[et al.], p. 22, 2010

  17. [23]

    Early history of machine learning,

    A. L. Fradkov, “Early history of machine learning,”IFAC- PapersOnLine, vol. 53, no. 2, pp. 1385–1390, 2020

  18. [24]

    A survey on deep learning: Algorithms, techniques, and applications,

    S. Pouyanfar, S. Sadiq, Y . Yan, H. Tian, Y . Tao, M. P. Reyes, M.- L. Shyu, S.-C. Chen, and S. S. Iyengar, “A survey on deep learning: Algorithms, techniques, and applications,”ACM computing surveys (CSUR), vol. 51, no. 5, pp. 1–36, 2018

  19. [25]

    Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberget al., “Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,” 2023

  20. [28]

    Artificial general intelligence: concept, state of the art, and future prospects,

    B. Goertzel, “Artificial general intelligence: concept, state of the art, and future prospects,”Journal of Artificial General Intelligence, vol. 5, no. 1, pp. 1–48, 2014

  21. [29]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021

  22. [30]

    Generative agents: Interactive simulacra of human behavior,

    J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” 2023. [Online]. Available: https://arxiv.org/abs/2304.03442

  23. [31]

    Memory in the age of ai agents,

    Y . Hu, S. Liu, Y . Yue, G. Zhang, B. Liu, F. Zhu, J. Lin, H. Guo, S. Dou, Z. Xiet al., “Memory in the age of ai agents,”arXiv preprint arXiv:2512.13564, 2025

  24. [32]

    Transformer-XL: Attentive language models beyond a fixed- length context,

    Z. Dai, Z. Yang, Y . Yang, J. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed- length context,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L. M `arquez,...

  25. [33]

    Compressive transformers for long-range sequence modelling,

    J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap, “Compressive transformers for long-range sequence modelling,” 2019. [Online]. Available: https://arxiv.org/abs/1911.05507

  26. [34]

    Memgpt: Towards llms as operating systems,

    C. Packer, S. Wooders, K. Lin, V . Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez, “Memgpt: Towards llms as operating systems,” 2024. [Online]. Available: https://arxiv.org/abs/2310.08560

  27. [35]

    Memorybank: Enhancing large language models with long-term memory,

    W. Zhong, L. Guo, Q. Gao, H. Ye, and Y . Wang, “Memorybank: Enhancing large language models with long-term memory,” inProceed- ings of the AAAI conference on artificial intelligence, vol. 38, no. 17, 2024, pp. 19 724–19 731

  28. [36]

    Synapse: Empowering llm agents with episodic-semantic memory via spreading activation,

    H. Jiang, J. Chen, Y . Pan, L. Chen, W. You, Y . Zhou, R. Zhang, A. Sikora, L. Zhao, Y . Abate, and T. Liu, “Synapse: Empowering llm agents with episodic-semantic memory via spreading activation,”

  29. [37]

    On the curse of memory in recurrent neural networks: Approximation and optimization analysis,

    Z. Li, J. Han, W. E, and Q. Li, “On the curse of memory in recurrent neural networks: Approximation and optimization analysis,” 2024. [Online]. Available: https://arxiv.org/abs/2009.07799

  30. [38]

    To backtrack or not to backtrack: When sequential search limits model reasoning,

    T. Qin, D. Alvarez-Melis, S. Jelassi, and E. Malach, “To backtrack or not to backtrack: When sequential search limits model reasoning,”

  31. [39]

    Large language models cannot self-correct reasoning yet,

    J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou, “Large language models cannot self-correct reasoning yet,”

  32. [40]

    Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning,

    X. Chen, A. Zhao, H. Xia, X. Lu, H. Wang, Y . Chen, W. Zhang, J. Wang, W. Li, and X. Shen, “Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning,” 2025. [Online]. Available: https://arxiv.org/abs/2505.16782

  33. [41]

    Available: https://arxiv.org/abs/2504.07052

    [Online]. Available: https://arxiv.org/abs/2504.07052

  34. [42]

    Ctrls: Chain-of-thought reasoning via latent state-transition,

    J. Wu, Y . Xiong, X. Li, S. Yu, Z. Hu, T. Yu, R. Wang, X. Chen, J. Shang, and J. McAuley, “Ctrls: Chain-of-thought reasoning via latent state-transition,” 2026. [Online]. Available: https://arxiv.org/abs/2507.08182

  35. [43]

    On the relation of state space models and hidden markov models,

    A. Ghojogh, M. H. Sepanj, and B. Ghojogh, “On the relation of state space models and hidden markov models,” 2026. [Online]. Available: https://arxiv.org/abs/2601.13357

  36. [44]

    Unsupervised methods for subgoal discov- ery during intrinsic motivation in model-free hierarchical reinforcement learning

    J. Rafati and D. C. Noelle, “Unsupervised methods for subgoal discov- ery during intrinsic motivation in model-free hierarchical reinforcement learning.” inKEG@ AAAI, 2019, pp. 17–25

  37. [45]

    Training large language models to reason in a continuous latent space,

    S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y . Tian, “Training large language models to reason in a continuous latent space,” 2025. [Online]. Available: https://arxiv.org/abs/2412.06769

  38. [46]

    Plangenllms: A modern survey of llm planning capabilities,

    H. Wei, Z. Zhang, S. He, T. Xia, S. Pan, and F. Liu, “Plangenllms: A modern survey of llm planning capabilities,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 19 497–19 521

  39. [47]

    Task planning and decision-making methods for intelligent agents based on large language models,

    X. Li, “Task planning and decision-making methods for intelligent agents based on large language models,” ser. AIIIP ’25. New York, NY , USA: Association for Computing Machinery, 2026, p. 817–822. [Online]. Available: https://doi.org/10.1145/3778534.3778661

  40. [48]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2201.11903

  41. [49]

    A brain-inspired agentic architecture to improve planning with LLMs,

    T. Webb, S. S. Mondal, and I. Momennejad, “A brain-inspired agentic architecture to improve planning with LLMs,”Nat. Commun., vol. 16, no. 1, p. 8633, Sep. 2025

  42. [50]

    Graph of thoughts: Solving elaborate problems with large language models,

    M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk et al., “Graph of thoughts: Solving elaborate problems with large language models,” inProceedings of the AAAI conference on artificial intelligence,...

  43. [51]

    Hiplan: Hierarchical planning for llm-based agents with adaptive global-local guidance,

    Z. Li, Y . Chang, G. Yu, and X. Le, “Hiplan: Hierarchical planning for llm-based agents with adaptive global-local guidance,”arXiv preprint arXiv:2508.19076, 2025

  44. [52]

    Adaplanner: Adaptive planning from feedback with language models,

    H. Sun, Y . Zhuang, L. Kong, B. Dai, and C. Zhang, “Adaplanner: Adaptive planning from feedback with language models,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, In...

  45. [53]

    Tree of thoughts: Deliberate problem solving with large language models,

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y . Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2305.10601

  46. [54]

    Self-refine: Iterative refinement with self-feedback,

    A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-refine: Iterative refinement with self-feedback,” 2023. [Online]. Available: ht...

  47. [55]

    Coherence-based alignment: A structural architecture for preventing goal drift in agentic ai systems

    A. Abdi, “Coherence-based alignment: A structural architecture for preventing goal drift in agentic ai systems.”

  48. [56]

    Pepa: a persistently autonomous embodied agent with personalities,

    K. Liu, Y . Li, L. Zhu, and W. Zhang, “Pepa: a persistently autonomous embodied agent with personalities,” 2026. [Online]. Available: https://arxiv.org/abs/2603.00117

  49. [57]

    Reflexion: Language agents with verbal reinforcement learning,

    N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023. [Online]. Available: https://arxiv.org/abs/2303.11366

  50. [58]

    Inherited goal drift: Contextual pressure can undermine agentic goals,

    A. Menon, M. Saebo, T. Crosse, S. Gibson, E. Jang, and D. Cruz, “Inherited goal drift: Contextual pressure can undermine agentic goals,” 2026. [Online]. Available: https://arxiv.org/abs/2603.03258

  51. [59]

    The metacognitive demands and opportunities of generative ai,

    L. Tankelevitch, V . Kewenig, A. Simkute, A. E. Scott, A. Sarkar, A. Sellen, and S. Rintel, “The metacognitive demands and opportunities of generative ai,” inProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, ser. CHI ’24. New York, NY , USA: Associa...

  52. [60]

    From human to model overconfidence: Evaluating confidence dynamics in large language models,

    B. Wen, C. Xu, B. HAN, R. Wolfe, L. L. Wang, and B. Howe, “From human to model overconfidence: Evaluating confidence dynamics in large language models,” inNeurIPS 2024 Workshop on Behavioral Machine Learning, 2024

  53. [61]

    Measuring goal-directedness,

    M. MacDermott, J. Fox, F. Belardinelli, and T. Everitt, “Measuring goal-directedness,” inAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 11 4...

  54. [62]

    Do i really know? learning factual self-verification for hallucination reduction,

    E. Altinisik, M. Fatehkia, F. Deniz, N. Durrani, M. Hawasly, M. Raza, and H. T. Sencar, “Do i really know? learning factual self-verification for hallucination reduction,” 2026. [Online]. Available: https://arxiv.org/abs/2602.02018 14

  55. [63]

    ProgCo: Program helps self-correction of large language models,

    X. Song, Y . Wu, W. Wang, J. Liu, W. Su, and B. Zheng, “ProgCo: Program helps self-correction of large language models,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), W. Che, J. Nabende, E. Shutova, and M. T...

  56. [64]

    Uncertainty quantification and confidence calibration in large language models: A survey,

    X. Liu, T. Chen, L. Da, C. Chen, Z. Lin, and H. Wei, “Uncertainty quantification and confidence calibration in large language models: A survey,” ser. KDD ’25. New York, NY , USA: Association for Computing Machinery, 2025, p. 6107–6117. [Online]. Available: https://doi.org/10.1...

  57. [65]

    Large language models can self-correct with minimal effort,

    Z. Wu, Q. Zeng, Z. Zhang, Z. Tan, C. Shen, and M. Jiang, “Large language models can self-correct with minimal effort,” inAI for Math Workshop@ ICML 2024, 2024

  58. [66]

    Out-of-distribution detection with positive and negative prompt super- vision using large language models,

    Z. He, C. Zhao, M. Shao, X. Wu, X. Zhao, D. Li, Q. Tian, and L. Yu, “Out-of-distribution detection with positive and negative prompt super- vision using large language models,”arXiv preprint arXiv:2511.10923, 2025

  59. [67]

    Revisiting epistemic markers in confidence estimation: Can markers accurately reflect large language models’ uncertainty?

    J. Liu, Q. Zong, W. Wang, and Y . Song, “Revisiting epistemic markers in confidence estimation: Can markers accurately reflect large language models’ uncertainty?” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers...

  60. [68]

    Know your limits: A survey of abstention in large language models,

    B. Wen, J. Yao, S. Feng, C. Xu, Y . Tsvetkov, B. Howe, and L. L. Wang, “Know your limits: A survey of abstention in large language models,” Transactions of the Association for Computational Linguistics, vol. 13, pp. 529–556, 2025

  61. [69]

    Envisioning outlier exposure by large language models for out-of-distribution detection,

    C. Cao, Z. Zhong, Z. Zhou, Y . Liu, T. Liu, and B. Han, “Envisioning outlier exposure by large language models for out-of-distribution detection,” 2024. [Online]. Available: https://arxiv.org/abs/2406.00806

  62. [70]

    Risan: Robust instance specific abstention network,

    B. Kalra, K. Shah, and N. Manwani, “Risan: Robust instance specific abstention network,”arXiv preprint arXiv:2107.03090, 2021

  63. [71]

    R-tuning: Instructing large language models to say ‘i don’t know’,

    H. Zhang, S. Diao, Y . Lin, Y . Fung, Q. Lian, X. Wang, Y . Chen, H. Ji, and T. Zhang, “R-tuning: Instructing large language models to say ‘i don’t know’,” inProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Huma...

  64. [72]

    Halluci- nate less by thinking more: Aspect-based causal abstention for large language models,

    V . Nguyen, Z. Xu, J. Chan, E. He, F. Xia, and X. Zhang, “Halluci- nate less by thinking more: Aspect-based causal abstention for large language models,”arXiv preprint arXiv:2511.17170, 2025

  65. [73]

    Selectivenet: A deep neural network with an integrated reject option,

    Y . Geifman and R. El-Yaniv, “Selectivenet: A deep neural network with an integrated reject option,” inInternational conference on machine learning. PMLR, 2019, pp. 2151–2159

  66. [74]

    Language agents meet causality – bridging llms and causal world models,

    J. Gkountouras, M. Lindemann, P. Lippe, E. Gavves, and I. Titov, “Language agents meet causality – bridging llms and causal world models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.19923

  67. [75]

    Language-guided world models: A model-based approach to AI control,

    A. Zhang, K. Nguyen, J. Tuyls, A. Lin, and K. Narasimhan, “Language-guided world models: A model-based approach to AI control,” inProceedings of the 4th Workshop on Spatial Language Understanding and Grounded Communication for Robotics (SpLU-RoboNLP 2024), P. Kordjamshidi, X. ...

  68. [76]

    Causalarc: Abstract reasoning with causal world models,

    J. Maasch, J. Kalantari, and K. Khezeli, “Causalarc: Abstract reasoning with causal world models,” 2026. [Online]. Available: https://arxiv.org/abs/2509.03636

  69. [77]

    Beyond world models: Rethinking understanding in ai models,

    T. Gupta and D. Pruthi, “Beyond world models: Rethinking understanding in ai models,” 2025. [Online]. Available: https://arxiv.org/abs/2511.12239

  70. [78]

    Recthinker: An agentic framework for tool-augmented reasoning in recommendation,

    H. Zhang, Y . Zhu, K. Mao, T. Li, and Z. Dou, “Recthinker: An agentic framework for tool-augmented reasoning in recommendation,”

  71. [79]

    Evaluating and improving tool-augmented computation-intensive math reasoning,

    B. Zhang, K. Zhou, X. Wei, X. Zhao, J. Sha, S. Wang, and J.-R. Wen, “Evaluating and improving tool-augmented computation-intensive math reasoning,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., v...

  72. [80]

    Grounding large language models for robot task planning using closed-loop state feedback,

    V . Bhat, A. U. Kaypak, P. Krishnamurthy, R. Karri, and F. Khorrami, “Grounding large language models for robot task planning using closed-loop state feedback,”Advanced Robotics Research, Nov. 2025. [Online]. Available: http://dx.doi.org/10.1002/adrr.202500072

  73. [81]

    ChatCoT: Tool-augmented chain-of-thought reasoning on chat- based large language models,

    Z. Chen, K. Zhou, B. Zhang, Z. Gong, X. Zhao, and J.-R. Wen, “ChatCoT: Tool-augmented chain-of-thought reasoning on chat- based large language models,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: As...

  74. [82]

    Ai for closed- loop control systems: New opportunities for modeling, designing, and tuning control systems,

    J. Sch ¨oning, A. Riechmann, and H.-J. Pfisterer, “Ai for closed- loop control systems: New opportunities for modeling, designing, and tuning control systems,” inProceedings of the 2022 14th International Conference on Machine Learning and Computing, ser. ICMLC ’22. New York, ...

  75. [83]

    Available: https://arxiv.org/abs/2603.09843

    [Online]. Available: https://arxiv.org/abs/2603.09843

  76. [84]

    Building self-evolving agents via experience- driven lifelong learning: A framework and benchmark,

    Y . Cai, Y . Hao, J. Zhou, H. Yan, Z. Lei, R. Zhen, Z. Han, Y . Yang, J. Li, Q. Panet al., “Building self-evolving agents via experience- driven lifelong learning: A framework and benchmark,”arXiv preprint arXiv:2508.19005, 2025

  77. [85]

    Deep reinforcement learning for robotics: A survey of real- world successes,

    C. Tang, B. Abbatematteo, J. Hu, R. Chandra, R. Mart ´ın-Mart´ın, and P. Stone, “Deep reinforcement learning for robotics: A survey of real- world successes,”Annual Review of Control, Robotics, and Autonomous Systems, vol. 8, no. 1, pp. 153–188, 2025

  78. [86]

    A metacognitive architecture for cor- recting llm errors in ai agents,

    J. Kim, M. Islam, and A. Goel, “A metacognitive architecture for cor- recting llm errors in ai agents,” inProceedings of the AAAI Conference on Artificial Intelligence, 2026

  79. [87]

    Contrastive test-time adaptation,

    D. Chen, D. Wang, T. Darrell, and S. Ebrahimi, “Contrastive test-time adaptation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 295–305

  80. [88]

    How human–ai feedback loops alter human perceptual, emotional and social judgements,

    M. Glickman and T. Sharot, “How human–ai feedback loops alter human perceptual, emotional and social judgements,”Nature Human Behaviour, vol. 9, no. 2, pp. 345–359, 2025

  81. [89]

    Mitigating the alignment tax of rlhf,

    Y . Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y . Yao, and T. Zhang, “Mitigating the alignment tax of rlhf,” 2024. [Online]. Available: https://arxiv.org/abs/2309.06256

  82. [90]

    Analyzing and reducing catastrophic forgetting in parameter efficient tuning,

    W. Ren, X. Li, L. Wang, T. Zhao, and W. Qin, “Analyzing and reducing catastrophic forgetting in parameter efficient tuning,” 2024. [Online]. Available: https://arxiv.org/abs/2402.18865

  83. [91]

    A proximal policy optimization-based reinforcement learning framework for real-time personalized endurance training,

    C. Wu, D. Liang, B. Yang, L. Xu, and Y . Li, “A proximal policy optimization-based reinforcement learning framework for real-time personalized endurance training,”Informatica, vol. 50, no. 8, 2026

  84. [92]

    Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation,

    M. Fawi, “Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation,”arXiv preprint arXiv:2408.14572, 2024

  85. [93]

    Reasoning as meta-learning: An optimization perspective to decipher long cot reasoning in llms

    J. Liu, H. Liu, L. Xiao, S. Liu, T. Zhang, Z. Ma, S. Zhang, and K. Chen, “Reasoning as meta-learning: An optimization perspective to decipher long cot reasoning in llms.”

  86. [94]

    Locating and editing factual associations in gpt,

    K. Meng, D. Bau, A. Andonian, and Y . Belinkov, “Locating and editing factual associations in gpt,”Advances in neural information processing systems, vol. 35, pp. 17 359–17 372, 2022

  87. [95]

    Mass-editing memory in a transformer,

    K. Meng, A. S. Sharma, A. Andonian, Y . Belinkov, and D. Bau, “Mass-editing memory in a transformer,” 2023. [Online]. Available: https://arxiv.org/abs/2210.07229

  88. [96]

    Overcoming catastrophic forgetting in neural networks,

    J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al., “Overcoming catastrophic forgetting in neural networks,”Pro- ceedings of the national academy of sciences, vol. 114, no. 13, pp. 3521–3526, 2017

  89. [97]

    Unveiling the pitfalls of knowledge editing for large language models,

    Z. Li, N. Zhang, Y . Yao, M. Wang, X. Chen, and H. Chen, “Unveiling the pitfalls of knowledge editing for large language models,”arXiv preprint arXiv:2310.02129, 2023

  90. [98]

    Eval- uating the ripple effects of knowledge editing in language models,

    R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva, “Eval- uating the ripple effects of knowledge editing in language models,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 283–298, 2024

  91. [99]

    Think-in-memory: Recalling and post-thinking enable llms with long-term memory,

    L. Liu, X. Yang, Y . Shen, B. Hu, Z. Zhang, J. Gu, and G. Zhang, “Think-in-memory: Recalling and post-thinking enable llms with long-term memory,” 2023. [Online]. Available: https://arxiv.org/abs/2311.08719

  92. [100]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Rayet al., “Training language models to follow instructions with human feedback,”Advances in neural information processing systems, vol. 35, pp. 27 730–27 744, 2022. 15

  93. [101]

    Alphaedit: Null-space constrained knowledge editing for language models,

    J. Fang, H. Jiang, K. Wang, Y . Ma, S. Jie, X. Wang, X. He, and T.-S. Chua, “Alphaedit: Null-space constrained knowledge editing for language models,”arXiv preprint arXiv:2410.02355, 2024

  94. [102]

    Holistic evaluation of language models,

    P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumaret al., “Holistic evaluation of language models,”arXiv preprint arXiv:2211.09110, 2022

  95. [103]

    40 years of cognitive architectures: core cognitive abilities and practical applications,

    I. Kotseruba and J. K. Tsotsos, “40 years of cognitive architectures: core cognitive abilities and practical applications,”Artificial Intelli- gence Review, vol. 53, no. 1, pp. 17–94, 2020

  96. [104]

    Survey of hallucination in natural language generation,

    Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y . Xu, E. Ishii, Y . J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”ACM computing surveys, vol. 55, no. 12, pp. 1–38, 2023

  97. [105]

    Truthfulqa: Measuring how models mimic human falsehoods,

    S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,” inProceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), 2022, pp. 3214–3252

  98. [106]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

    A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonsoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”Transactions on machine learning research, 2023

  99. [108]

    Judging llm-as-a-judge with mt- bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xinget al., “Judging llm-as-a-judge with mt- bench and chatbot arena,”Advances in neural information processing systems, vol. 36, pp. 46 595–46 623, 2023

  100. [111]

    Red teaming language models with language models,

    E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red teaming language models with language models,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3419–3448

  101. [2024]

    Available: https://arxiv.org/abs/2310.01798

    [Online]. Available: https://arxiv.org/abs/2310.01798

  102. [2025]

    Available: http://dx.doi.org/10.1613/jair.1.18675

    [Online]. Available: http://dx.doi.org/10.1613/jair.1.18675

  103. [2026]

    Available: https://arxiv.org/abs/2601.02744

    [Online]. Available: https://arxiv.org/abs/2601.02744

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.