Pith. sign in

REVIEW 3 major objections 8 minor 4 cited by

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

T0 review · 3 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that reaching AGI from large language models will require integrating four principles of human cognition—embodiment, symbol grounding, causality, and memory—and that simply scaling up data and parameters will not suffice.

desk verdict Useful survey of embodiment, grounding, causality, and memory for LLMs, but the necessity claim is asserted, not argued, and the broadened definition of embodiment makes it hard to falsify. read the letter →

arxiv 2501.03151 v1 pith:SQFMIXGH submitted 2025-01-06 cs.AI cs.CVcs.LG

classification cs.AIcs.CVcs.LG
keywords largelanguagemodelsartificialgeneralintelligenceembodimentsymbolgroundingcausalreasoningmemorymechanismsfoundationcognitiveprinciples
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are powerful but brittle: their knowledge comes from statistical patterns in text, so they can mimic reasoning without understanding causes, physical constraints, or meaning. The paper argues that reaching artificial general intelligence from LLMs will require embedding four principles of human cognition into the models themselves—embodiment (having a body that acts in and senses the world), symbol grounding (connecting abstract symbols to real referents), causality (reasoning about cause and effect beyond correlation), and memory (storing and reusing experience). It surveys concrete techniques for each principle and argues that simple scaling of data and parameters will not be enough. A sympathetic reader should take the paper as a road map: AGI will come from a unified architecture that treats these principles as interdependent, not from larger versions of current models.

What carries the argument

The central object is a four-component cognitive architecture: embodiment (a body with sensors and actuators that generates goal-directed experiences), symbol grounding (a mapping from internal symbols such as words to real-world referents), causality (relations organized by the association, intervention, and counterfactual hierarchy), and memory (sensory, working, and long-term stores, the last divided into semantic, episodic, and procedural memory). The machinery does its work through a closed loop: the embodied agent acts and senses, grounding abstract symbols in physical experience and learning causal relationships from feedback; memory then preserves the grounded, causal knowledge as prior knowledge that later perception, reasoning, and planning can draw on. The paper argues that each mechanism addresses a specific failure of current LLMs and that only their integration yields general intelligence.

What would settle it

Take the same pretrained language model and compare two versions, one with access to a physical or simulated body that can act and receive sensorimotor feedback and one that only reads static text and images. If the bodyless version matches or exceeds the embodied version across a broad battery of novel tasks requiring physical common sense, interventions, and counterfactual reasoning, the claim that embodiment, grounding, and causal experience are necessary for AGI is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that the cognitive limitations of current LLMs—superficial context understanding, correlation-based predictions, lack of physical common sense, and inability to accumulate knowledge—trace to the absence of four foundational capabilities. Embodiment supplies agency, goal-directedness, self-awareness, and situatedness; symbol grounding ties internal representations to real-world entities; causality lifts models from association to intervention and counterfactual reasoning; and memory, from sensory buffers to long-term semantic, episodic, and procedural stores, lets knowledge persist and be reused. The paper surveys how each capability is being implemented in LLM-based systems and synthesizes them into a single functional framework in which embodied experiences ground symbols, grounded experiences reveal causal structure, and memory encodes all of it for future perception, reasoning, and action.

Load-bearing premise

The load-bearing premise is that embodiment is necessary for general intelligence, not merely beneficial; if a bodyless, purely text-trained model could match human-level generality, the paper's four-principle framework would not be required.

Editorial extensions

If this is right

  • Scaling data and parameters alone will not produce human-level generality; progress requires coupling LLMs with bodies, grounded representations, causal models, and persistent memory.
  • Embodied training in simulated worlds—game engines, physics simulators, extended reality, and AI-generated environments—becomes a core route because real-world interactive data is too costly and static.
  • Memory must move beyond the context window toward explicit long-term stores, since the context window loses information in its middle and parameter storage suffers from catastrophic forgetting.
  • Causal reasoning must be engineered through causal graphs, structural causal models, or physics-informed world models, because text-trained LLMs predominantly learn correlations.
  • The four principles are mutually reinforcing, so implementing them piecemeal will be less effective than a unified architecture that lets embodied, grounded, causal experience flow into memory and back out into reasoning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the framework predicts a gradient rather than a cliff—models with progressively more embodiment, grounding, causal structure, and memory should generalize progressively better on novel physical and social tasks, making the thesis testable in degrees.
  • Editorial inference: the paper's own caveat that human and machine intelligence are not directly comparable implies that AGI evaluation should be redesigned around generalization across task distributions and transfer efficiency, not benchmark score comparisons.
  • Editorial inference: coupling episodic memory with causal structure—storing records of interventions and their outcomes as counterfactual training signal—is a natural extension the paper leaves implicit, and it could be evaluated on counterfactual reasoning benchmarks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This survey argues that large language models will only achieve artificial general intelligence if they address four foundational cognitive principles: embodiment, symbol grounding, causality, and memory. After reviewing how LLMs can be extended toward generalist behavior, the paper devotes one major section to each principle, surveying real-world and simulated embodiment, knowledge-graph and interaction-based grounding, causal modeling including neuro-symbolic and physics-informed approaches, and multi-level memory systems. It then proposes a conceptual architecture that integrates the four principles, and concludes with a discussion of the prospects and evaluation challenges of LLM-based AGI. The paper is primarily an organizing survey: it synthesizes a large body of recent work, provides summary tables, and offers a unified vocabulary for thinking about these capabilities. Its central normative claim, however, is that the four principles are required for LLM-based AGI, and this necessity claim is asserted rather than derived from evidence.

Significance. If the necessity claim were established, the paper would provide a valuable organizing framework for AGI research and a clear agenda for combining embodiment, grounding, causality, and memory in future LLM architectures. The survey itself is genuinely useful as a reference: it draws on a broad citation base, distinguishes practical implementation families within each principle, and includes helpful comparative tables (Tables 1-3) and a synthesis diagram (Figure 16). It does not contain new experiments, machine-checked proofs, or a formal argument, so its contribution is conceptual and taxonomic rather than empirical. The main risk to significance is the gap between the survey material (which shows that these principles are useful in many systems) and the stronger thesis (that they are necessary for AGI), a gap that the manuscript's own Section 8 discussion partly acknowledges.

major comments (3)
  1. [Abstract, §2.3, §9] The central necessity claim ('required to be addressed' in the abstract; 'essential' in §9) is asserted rather than established. The body of the paper demonstrates that embodiment, grounding, causality, and memory are each beneficial in particular systems and can address specific LLM shortcomings, but benefit does not imply necessity. The authors should either supply a principled argument for why omitting any one of the four principles precludes human-level generality, or explicitly weaken the thesis to 'important design dimensions.' The current wording in §9 is internally unstable, stating that the concepts are 'by no means the only principles necessary' while still calling them 'essential.'
  2. [§3.5.2(b), §8] The broadened definition of embodiment makes the necessity claim trivially satisfiable and unfalsifiable. Section 3.5.2(b) states that autonomous agents operating in virtual mode 'can still be considered as embodied' if they have virtual bodies, sensing, and actuation allowing interaction with the physical environment. Under that definition, an ordinary tool-using LLM agent that receives observations and returns actions already qualifies as embodied, so the claim that embodiment is required for AGI imposes no real constraint. This tension is compounded by §8, which says that current LLM agents are 'not very far from some form of general intelligence' despite lacking the full physical embodiment emphasized in Section 3. The paper should specify a minimal, verifiable notion of embodiment, distinguish degrees of embodiment, and state what observable difference the presence or absence of the four principles would make.
  3. [§7] The proposed holistic framework is purely schematic: it consists of a functional block diagram and a qualitative description of how the four subsystems interact, with no evaluation criteria, no comparison to alternative architectures, and no testable predictions. As a survey, the absence of experiments is acceptable, but if the paper is to support the 'foundational' status of the four principles, the framework should generate at least qualitative predictions that distinguish it from scaling-only approaches, such as differences in data efficiency, robustness to distribution shift, or counterfactual reasoning performance. As written, the framework is compatible with too wide a range of systems to serve as evidence for the paper's central thesis.
minor comments (8)
  1. [§6.2] The bullet list of memory-implementation techniques includes 'Adequate diversity and variability,' which is a requirement for virtual environments from §3.5.2(b) and does not belong in a list of memory mechanisms.
  2. [§3.5.2(b)] The sentence 'instead deploying in cyberphysical systems' is ungrammatical; it should read 'instead of deploying in cyber-physical systems.'
  3. [§5.2.3] The text refers to a 'metal model'; this should be 'mental model.'
  4. [§6.3.3(b)] The phrase 'boots speed' should be 'boosts speed.'
  5. [§8] The sentence 'The power of have LLMs have also be exploited to adapt computer graphics-generated worlds' is garbled and should be rewritten, for example as 'The power of LLMs has also been exploited to adapt computer graphics-generated worlds.'
  6. [Throughout] The abbreviation 'LMM' is used in several places where 'LLM' is intended (e.g., §4.4.4, 'the resulting LMM'), which will confuse readers.
  7. [Figure 13 caption] The caption contains the typo 'and soo forth'; it should read 'and so forth.'
  8. [§6.2.4] The phrase 'as a result of unknown errors, including the presence of unknown errors' repeats 'unknown errors' and should be rephrased.

Circularity Check

1 steps flagged · score 4.0 of 10

The four 'foundational principles' are largely read off the paper's own stipulated AGI definition, making the necessity claim partly definitional rather than derived.

  1. self definitional [Section 1.3 (AGI definition) and Section 2.3/Figure 3 (mapping to foundational concepts); echoed in Abstract and Section 9.]
    "the most important features of a typical AGI system are that ... it retains and accumulates relevant information in memory and reuse the knowledge in future tasks; and it can understand context and perform high-level cognitive tasks such as abstract and commonsense reasoning. ... A number of foundational problems —embodiment, symbol grounding, causality and memory — are required to be addressed for LLMs to attain human-level general intelligence."

    The paper stipulates AGI features in Section 1.3/Figure 1 (memory retention, situatedness, goal-directedness, abstract and commonsense reasoning) and then, in Figure 3, explicitly maps these AGI features to the four 'foundational' concepts: goal-directed behavior/autonomy to embodiment, context understanding to grounding, robust generalization/reasoning to causality, and continual learning to memory. The conclusion that these principles are 'required' for AGI therefore follows by construction from the paper's own definition and mapping, not from independent evidence or a derivation. The survey content about implementation approaches is independent, but the normative necessity claim in the abstract and conclusion reduces to the stipulative AGI features.

full rationale

This is a survey/position paper with no fitted parameters, no predictions, and no equations, so the 'fitted input called prediction' and self-citation-chain patterns do not occur; the authors also do not rely on load-bearing self-citations. The only circularity is conceptual: the necessity of the four principles is built into the paper's definition of AGI and its Figure 3 mapping. That does not undermine the descriptive survey of embodiment, grounding, causality, and memory techniques, which is anchored in external references and is self-contained as a literature review. However, the abstract and conclusion make a stronger claim than the body supports: that these principles are 'required' and 'essential.' That claim is at least partly true by definition rather than demonstrated, so a moderate circularity score is warranted. The Section 8 observation that current LLM agents are 'not very far' from general intelligence despite incomplete embodiment is a tension with the necessity claim, but it is a correctness issue, not a further circular step.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities. The paper's conceptual framework rests on domain assumptions about the necessity of four cognitive principles, which are common in the AGI literature but not formally proven.

assumptions (5)
  • domain assumption Embodiment is a necessary prerequisite for artificial general intelligence.
    Section 3.2 argues that AI can only attain general intelligence if linked with a physical body, citing [109], [129], [131]. This is an unproven philosophical hypothesis.
  • domain assumption Symbol grounding is necessary for robust general intelligence.
    Section 4 claims that grounding bridges the semantic gap between AI and the real world, but this is an assumption about cognitive architecture, not a proven requirement.
  • domain assumption Causal reasoning capacity is necessary for general intelligence.
    Section 5 asserts that causal understanding allows robust generalization, but the necessity is assumed from cognitive science literature.
  • domain assumption Memory mechanisms are necessary for general intelligence.
    Section 6 treats memory as a core building block for continual learning and adaptation, but the necessity is assumed from biological analogy.
  • domain assumption Pearl's three-level hierarchy of causality (association, intervention, counterfactual) is a valid taxonomy.
    Section 5.1.1 relies on Pearl's hierarchy as the organizing framework for causal reasoning levels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches." pith.science (2026). https://pith.science/paper/SQFMIXGH

@misc{pith2026250103151,
  author       = {Pith},
  title        = {Pith review of: Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SQFMIXGH}},
  note         = {Machine review of arXiv:2501.03151}
}
read the original abstract

Generative artificial intelligence (AI) systems based on large-scale pretrained foundation models (PFMs) such as vision-language models, large language models (LLMs), diffusion models and vision-language-action (VLA) models have demonstrated the ability to solve complex and truly non-trivial AI problems in a wide variety of domains and contexts. Multimodal large language models (MLLMs), in particular, learn from vast and diverse data sources, allowing rich and nuanced representations of the world and, thereby, providing extensive capabilities, including the ability to reason, engage in meaningful dialog; collaborate with humans and other agents to jointly solve complex problems; and understand social and emotional aspects of humans. Despite this impressive feat, the cognitive abilities of state-of-the-art LLMs trained on large-scale datasets are still superficial and brittle. Consequently, generic LLMs are severely limited in their generalist capabilities. A number of foundational problems -- embodiment, symbol grounding, causality and memory -- are required to be addressed for LLMs to attain human-level general intelligence. These concepts are more aligned with human cognition and provide LLMs with inherent human-like cognitive properties that support the realization of physically-plausible, semantically meaningful, flexible and more generalizable knowledge and intelligence. In this work, we discuss the aforementioned foundational issues and survey state-of-the art approaches for implementing these concepts in LLMs. Specifically, we discuss how the principles of embodiment, symbol grounding, causality and memory can be leveraged toward the attainment of artificial general intelligence (AGI) in an organic manner.

Figures

Figures reproduced from arXiv: 2501.03151 by the authors.

Figure 1
Figure 1. Some of the most important features of artificial general intelligence (AGI) systems. These features give AGI systems vast cognitive [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. LLM versus human intelligence: Important mechanisms that allow flexible extension of knowledge and cognitive abilities. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. A summary of the essence and role of each of the foundational AGI concepts covered in this work. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: In this scene, two intelligent agents A and B assist during an emergency. When driven by high-level goals that are aligned with human interests and values, such agents can perform good acts spontaneously. Goal-awareness allows them to be proactive, autonomous and capab…
Figure 5
Figure 5. Figure 5: A simplified representation of EmbodiedGPT [172]. The framework utilizes a large-scale egocentric, EgoCOT—curated as part of the [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: A practical case where situatedness (situational-awareness [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: MultiPLY [215], a state-of-the-art embodied LLM trained on simulated worlds, supports a vast array of sensory modalities, including [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: A summary of the important contributions of the main aspects of embodiment to AGI capabilities and the general approaches to [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Human cognition relies on associating abstract mental repre [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: The grounding mechanism allows intelligent systems to rep [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Grounding can be achieved by actively exploring the world [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 13
Figure 13. Figure 13: Interpreting events and observations or performing everyday [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]
Figure 12
Figure 12. Figure 12: Levels of causality per Pearl [402] and the types of problems [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 14
Figure 14. Figure 14: Unlike AI systems, humans naturally have an intuitive under [PITH_FULL_IMAGE:figures/full_fig_p018_14.png]
Figure 15
Figure 15. Figure 15: A simplified representation of the memory system showing information flow as well as the interaction various components and the [PITH_FULL_IMAGE:figures/full_fig_p024_15.png]
Figure 16
Figure 16. Figure 16: Functional block diagram of a generalized AGI system based on the principles covered in this article. The conceptual model consists of [PITH_FULL_IMAGE:figures/full_fig_p025_16.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  2. The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Later OpenAI o-series models substantially outperform GPT-series models on multimodal puzzles, but fine-grained visual perception and algorithmic puzzles remain hard.

  3. Language Games as the Pathway to Artificial Superhuman Intelligence

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.

  4. Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

    cs.AI 2025-07 conditional novelty 2.0 of 10

    A broad survey arguing that AGI requires modular, memory-augmented, embodied architectures rather than scaled-up token prediction, with a brief proposal to decompose intelligence into five components.

Reference graph

Works this paper leans on

300 extracted references · 8 canonical work pages · cited by 4 Pith papers

  1. [1]

    Beyond the octopus: From general intelligence toward a human-like mind,

    S. S. Adams and S. Burbeck, “Beyond the octopus: From general intelligence toward a human-like mind,” in Theoretical foundations of artificial general intelligence. Springer, 2012, pp. 49–65

  2. [2]

    Auto-encoding variational bayes,

    D. P . Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013

  3. [3]

    Image-to-image translation with conditional adversarial networks,

    P . Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1125–1134

  4. [4]

    Deep generative modelling: A comparative review of vaes, gans, nor- malizing flows, energy-based and autoregressive models,

    S. Bond-Taylor, A. Leach, Y. Long, and C. G. Willcocks, “Deep generative modelling: A comparative review of vaes, gans, nor- malizing flows, energy-based and autoregressive models,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 11, pp. 7327–7347, 2021

  5. [5]

    Large language models: A survey,

    S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao, “Large language models: A survey,” arXiv preprint arXiv:2402.06196, 2024

  6. [6]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10 850–10 869, 2023

  7. [7]

    A survey of vision-language pre-trained models,

    Y. Du, Z. Liu, J. Li, and W. X. Zhao, “A survey of vision-language pre-trained models,” arXiv preprint arXiv:2202.10936, 2022

  8. [8]

    A survey on vision-language-action models for embodied ai,

    Y. Ma, Z. Song, Y. Zhuang, J. Hao, and I. King, “A survey on vision-language-action models for embodied ai,” arXiv preprint arXiv:2405.14093, 2024

Show all 300 references
  1. [9]

    Sparks of artificial general intelligence: Early experiments with gpt-4,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P . Lee, Y. T. Lee, Y. Li, S. Lundberg et al. , “Sparks of artificial general intelligence: Early experiments with gpt-4,” arXiv preprint arXiv:2303.12712, 2023

  2. [10]

    Artificial general intelligence is already here,

    B. A. y Arcas, “Artificial general intelligence is already here,” 2023

  3. [11]

    Foundation models in robotics: Applications, challenges, and the future,

    R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y. Zhu, S. Song, A. Kapoor, K. Hausman et al. , “Foundation models in robotics: Applications, challenges, and the future,” The Interna- tional Journal of Robotics Research, p. 02783649241281508, 2023

  4. [12]

    Seqgpt: An out-of-the-box large language model for open domain sequence understanding,

    T. Yu, C. Jiang, C. Lou, S. Huang, X. Wang, W. Liu, J. Cai, Y. Li, Y. Li, K. Tu et al. , “Seqgpt: An out-of-the-box large language model for open domain sequence understanding,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, 2024, pp. 19 458–19 467

  5. [13]

    Transforming human-centered ai collaboration: Redefining embodied agents capabilities through interactive grounded language instructions,

    S. Mohanty, N. Arabzadeh, J. Kiseleva, A. Zholus, M. Teruel, A. Awadallah, Y. Sun, K. Srinet, and A. Szlam, “Transforming human-centered ai collaboration: Redefining embodied agents capabilities through interactive grounded language instructions,” arXiv preprint arXiv:2305.10783, 2023

  6. [14]

    What makes us smart? core knowledge and natural language,

    E. S. Spelke, “What makes us smart? core knowledge and natural language,” 2003

  7. [15]

    Sensory capacities of the chimpanzee: A review

    A. Prestrude, “Sensory capacities of the chimpanzee: A review.” Psychological bulletin, vol. 74, no. 1, p. 47, 1970

  8. [16]

    Animal senses compared to the human sense of time,

    R. Doble, “Animal senses compared to the human sense of time,” 2014

  9. [17]

    Motor system evolution and the emergence of high cognitive functions,

    G. Mendoza and H. Merchant, “Motor system evolution and the emergence of high cognitive functions,” Progress in neurobiology, vol. 122, pp. 73–93, 2014

  10. [18]

    Distributed hierarchical processing in the primate cerebral cortex

    D. J. Felleman and D. C. Van Essen, “Distributed hierarchical processing in the primate cerebral cortex.” Cerebral cortex (New York, NY: 1991), vol. 1, no. 1, pp. 1–47, 1991

  11. [19]

    Learning to walk: Ecological demands and phyloge- netic constraints

    E. Thelen, “Learning to walk: Ecological demands and phyloge- netic constraints.” Advances in infancy research, 1984

  12. [20]

    Brain mechanisms of acoustic communication in humans and nonhuman primates: an evolutionary perspective,

    H. Ackermann, S. R. Hage, and W. Ziegler, “Brain mechanisms of acoustic communication in humans and nonhuman primates: an evolutionary perspective,” Behavioral and Brain Sciences , vol. 37, no. 6, pp. 529–546, 2014

  13. [21]

    Becoming human: human infants link language and cognition, but what about the other great apes?

    M. A. Novack and S. Waxman, “Becoming human: human infants link language and cognition, but what about the other great apes?” Philosophical Transactions of the Royal Society B , vol. 375, no. 1789, p. 20180408, 2020

  14. [22]

    The emergence of language in the human mind and brain- insights from the neurobiology of language, thought and action

    N. J. Bourguignon, “The emergence of language in the human mind and brain- insights from the neurobiology of language, thought and action.” Psychological Review, vol. 130, no. 6, p. 1544, 2023

  15. [23]

    The never-ending innovativeness of homo sapiens,

    E. L. Grigorenko, “The never-ending innovativeness of homo sapiens,” in Creativity, Innovation, and Change Across Cultures . Springer, 2023, pp. 3–27

  16. [24]

    Magic words: How language augments human com- putation,

    A. Clark, “Magic words: How language augments human com- putation,” in Language and meaning in cognitive science . Rout- ledge, 2012, pp. 21–39

  17. [25]

    Language and consumer memory: The impact of linguistic differences between chinese and english,

    B. H. Schmitt, Y. Pan, and N. T. Tavassoli, “Language and consumer memory: The impact of linguistic differences between chinese and english,” Journal of consumer research , vol. 21, no. 3, pp. 419–431, 1994

  18. [26]

    Linguistic versus cultural relativity: On japanese-chinese differences in picture description and re- call,

    Y. Tajima and N. Duffield, “Linguistic versus cultural relativity: On japanese-chinese differences in picture description and re- call,” Cognitive Linguistics, vol. 23, no. 4, pp. 675–709, 2012

  19. [27]

    Perceiving and remembering events cross-linguistically: Evidence from dual-task paradigms,

    J. C. Trueswell and A. Papafragou, “Perceiving and remembering events cross-linguistically: Evidence from dual-task paradigms,” Journal of Memory and Language, vol. 63, no. 1, pp. 64–82, 2010

  20. [28]

    Remembering how: Language, memory, and the salience of manner,

    M. I. Feist and P . C. Férez, “Remembering how: Language, memory, and the salience of manner,” Journal of Cognitive Science, vol. 14, no. 4, pp. 379–398, 2013

  21. [29]

    Can language restructure cognition? the case for space,

    A. Majid, M. Bowerman, S. Kita, D. B. Haun, and S. C. Levinson, “Can language restructure cognition? the case for space,” Trends in cognitive sciences, vol. 8, no. 3, pp. 108–114, 2004

  22. [30]

    Linguistic anthropology: Language as a non-neutral medium,

    A. Duranti, “Linguistic anthropology: Language as a non-neutral medium,” The Cambridge handbook of sociolinguistics , pp. 28–46, 2011

  23. [31]

    The centrality of language in human cognition,

    G. Lupyan, “The centrality of language in human cognition,” Language Learning, vol. 66, no. 3, pp. 516–553, 2016

  24. [32]

    Unconscious effects of language-specific terminology on preattentive color perception,

    G. Thierry, P . Athanasopoulos, A. Wiggett, B. Dering, and J.-R. Kuipers, “Unconscious effects of language-specific terminology on preattentive color perception,” Proceedings of the National Academy of Sciences, vol. 106, no. 11, pp. 4567–4570, 2009

  25. [33]

    Perceptual processing is facilitated by ascribing meaning to novel stimuli,

    G. Lupyan and M. J. Spivey, “Perceptual processing is facilitated by ascribing meaning to novel stimuli,” Current Biology, vol. 18, no. 10, pp. R410–R412, 2008

  26. [34]

    Motion detection and motion verbs: Language affects low-level visual perception,

    L. Meteyard, B. Bahrami, and G. Vigliocco, “Motion detection and motion verbs: Language affects low-level visual perception,” Psychological Science, vol. 18, no. 11, pp. 1007–1013, 2007

  27. [35]

    Evolution of the brain and intelligence in primates,

    G. Roth and U. Dicke, “Evolution of the brain and intelligence in primates,” Progress in brain research, vol. 195, pp. 413–430, 2012

  28. [36]

    Artificial general intelligence: concept, state of the art, and future prospects,

    B. Goertzel, “Artificial general intelligence: concept, state of the art, and future prospects,” Journal of Artificial General Intelligence , vol. 5, no. 1, p. 1, 2014

  29. [37]

    The risks associated with artificial general intelli- gence: A systematic review,

    S. McLean, G. J. Read, J. Thompson, C. Baber, N. A. Stanton, and P . M. Salmon, “The risks associated with artificial general intelli- gence: A systematic review,” Journal of Experimental & Theoretical Artificial Intelligence, vol. 35, no. 5, pp. 649–663, 2023

  30. [38]

    Introduction: Aspects of artificial general intelligence,

    P . Wang and B. Goertzel, “Introduction: Aspects of artificial general intelligence,” in Advances in Artificial General Intelligence: Concepts, Architectures and Algorithms. IOS Press, 2007, pp. 1–16

  31. [39]

    Towards artificial general intelligence via a multimodal foundation model,

    N. Fei, Z. Lu, Y. Gao, G. Yang, Y. Huo, J. Wen, H. Lu, R. Song, X. Gao, T. Xiang et al., “Towards artificial general intelligence via a multimodal foundation model,” Nature Communications, vol. 13, no. 1, p. 3094, 2022

  32. [40]

    Evidence of interrelated cognitive- like capabilities in large language models: Indications of artificial general intelligence or achievement?

    D. Ili´ c and G. E. Gignac, “Evidence of interrelated cognitive- like capabilities in large language models: Indications of artificial general intelligence or achievement?” Intelligence, vol. 106, p. 101858, 2024

  33. [41]

    Hybrid strategies towards safe self- aware superintelligent systems,

    N.-M. Aliman and L. Kester, “Hybrid strategies towards safe self- aware superintelligent systems,” in Artificial General Intelligence: 11th International Conference, AGI 2018, Prague, Czech Republic, August 22-25, 2018, Proceedings 11. Springer, 2018, pp. 1–11

  34. [42]

    Rise of artificial general intelligence: risks and opportunities,

    G. Buttazzo, “Rise of artificial general intelligence: risks and opportunities,” Frontiers in artificial intelligence, vol. 6, p. 1226990, 2023. 28

  35. [43]

    Map- ping the landscape of human-level artificial general intelligence,

    S. Adams, I. Arel, J. Bach, R. Coop, R. Furlan, B. Goertzel, J. S. Hall, A. Samsonovich, M. Scheutz, M. Schlesinger et al. , “Map- ping the landscape of human-level artificial general intelligence,” AI magazine, vol. 33, no. 1, pp. 25–42, 2012

  36. [44]

    Why generality is key to human- level artificial intelligence,

    T. R. Besold and U. Schmid, “Why generality is key to human- level artificial intelligence,” Advances in Cognitive Systems , vol. 4, pp. 13–24, 2016

  37. [45]

    Growth, degrowth, and the challenge of artificial superintelligence,

    S. Pueyo, “Growth, degrowth, and the challenge of artificial superintelligence,” Journal of Cleaner Production , vol. 197, pp. 1731–1736, 2018

  38. [46]

    Ethical issues in advanced artificial intelligence,

    N. Bostrom, “Ethical issues in advanced artificial intelligence,” Machine Ethics and Robot Ethics, pp. 69–75, 2020

  39. [47]

    Contemporary approaches to artificial general intelligence,

    C. Pennachin and B. Goertzel, “Contemporary approaches to artificial general intelligence,” in Artificial general intelligence . Springer, 2007, pp. 1–30

  40. [48]

    Essentials of general intelligence: The direct path to artificial general intelligence,

    P . Voss, “Essentials of general intelligence: The direct path to artificial general intelligence,” Artificial general intelligence , pp. 131–157, 2007

  41. [49]

    The emperor of strong ai has no clothes: limits to artificial intelligence,

    A. Braga and R. K. Logan, “The emperor of strong ai has no clothes: limits to artificial intelligence,” Information, vol. 8, no. 4, p. 156, 2017

  42. [50]

    Strong and weak ai: Deweyan considerations

    J. C. Flowers, “Strong and weak ai: Deweyan considerations.” in AAAI spring symposium: Towards conscious AI systems , vol. 2287, no. 7, 2019

  43. [51]

    Towards strong ai,

    M. V . Butz, “Towards strong ai,” KI-Künstliche Intelligenz, vol. 35, no. 1, pp. 91–101, 2021

  44. [52]

    The cognitive neuroscience of self-awareness: Current framework, clinical im- plications, and future research directions,

    D. C. Mograbi, S. Hall, B. Arantes, and J. Huntley, “The cognitive neuroscience of self-awareness: Current framework, clinical im- plications, and future research directions,” Wiley Interdisciplinary Reviews: Cognitive Science, vol. 15, no. 2, p. e1670, 2024

  45. [53]

    Takeno, Creation of a conscious robot: Mirror image cognition and self-awareness

    J. Takeno, Creation of a conscious robot: Mirror image cognition and self-awareness. CRC Press, 2012

  46. [54]

    Ai is not sentient. why do people say it is?

    C. Metz, “Ai is not sentient. why do people say it is?” International New York Times, pp. NA–NA, 2022

  47. [55]

    Refuting strong ai: Why consciousness cannot be algorithmic,

    A. Knight, “Refuting strong ai: Why consciousness cannot be algorithmic,” arXiv preprint arXiv:1906.10177, 2019

  48. [56]

    In search of the moral status of ai: why sentience is a strong argument,

    M. Gibert and D. Martin, “In search of the moral status of ai: why sentience is a strong argument,”AI & SOCIETY, vol. 37, no. 1, pp. 319–330, 2022

  49. [57]

    Strong artificial intelligence and consciousness,

    G. W. Ng and W. C. Leung, “Strong artificial intelligence and consciousness,” Journal of Artificial Intelligence and Consciousness , vol. 7, no. 01, pp. 63–72, 2020

  50. [58]

    Artificial intelligence and consciousness,

    D. McDermott, “Artificial intelligence and consciousness,” The Cambridge handbook of consciousness, pp. 117–150, 2007

  51. [59]

    Why ai still doesnt have conscious- ness?

    D. Li, W. He, and Y. Guo, “Why ai still doesnt have conscious- ness?” CAAI Transactions on Intelligence Technology , vol. 6, no. 2, pp. 175–179, 2021

  52. [60]

    The problem of moral agency in artificial intelligence,

    R. Manna and R. Nath, “The problem of moral agency in artificial intelligence,” in 2021 IEEE Conference on Norbert Wiener in the 21st Century (21CW). IEEE, 2021, pp. 1–4

  53. [61]

    Autonomous morals: Inferences of mind predict acceptance of ai behavior in sacrificial moral dilemmas,

    A. D. Young and A. E. Monroe, “Autonomous morals: Inferences of mind predict acceptance of ai behavior in sacrificial moral dilemmas,” Journal of Experimental Social Psychology , vol. 85, p. 103870, 2019

  54. [62]

    Gpt-3: Its nature, scope, limits, and consequences,

    L. Floridi and M. Chiriatti, “Gpt-3: Its nature, scope, limits, and consequences,” Minds and Machines, vol. 30, pp. 681–694, 2020

  55. [63]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of naacL-HLT, vol. 1. Minneapolis, Minnesota, 2019, p. 2

  56. [64]

    Palm-e: An embod- ied multimodal language model,

    D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yuet al., “Palm-e: An embod- ied multimodal language model,”arXiv preprint arXiv:2303.03378, 2023

  57. [65]

    Minigpt-4: Enhancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592, 2023

  58. [66]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P . Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems, vol. 35, pp. 23 716–23 736, 2022

  59. [67]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024

  60. [68]

    Making pre-trained language models better few-shot learners,

    T. Gao, A. Fisch, and D. Chen, “Making pre-trained language models better few-shot learners,” arXiv preprint arXiv:2012.15723, 2020

  61. [69]

    Large language models are few-shot health learners,

    X. Liu, D. McDuff, G. Kovacs, I. Galatzer-Levy, J. Sunshine, J. Zhan, M.-Z. Poh, S. Liao, P . Di Achille, and S. Patel, “Large language models are few-shot health learners,” arXiv preprint arXiv:2305.15525, 2023

  62. [70]

    Large language models are zero-shot reasoners,

    T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems, vol. 35, pp. 22 199–22 213, 2022

  63. [71]

    A survey on multimodal large language models,

    S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,” National Science Review , p. nwae403, 2024

  64. [72]

    From data to commonsense rea- soning: The use of large language models for explainable ai,

    S. Krause and F. Stolzenburg, “From data to commonsense rea- soning: The use of large language models for explainable ai,” arXiv preprint arXiv:2407.03778, 2024

  65. [73]

    Large language models for mathematical reasoning: Progresses and challenges,

    J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin, “Large language models for mathematical reasoning: Progresses and challenges,” arXiv preprint arXiv:2402.00157, 2024

  66. [74]

    Mathprompter: Mathe- matical reasoning using large language models,

    S. Imani, L. Du, and H. Shrivastava, “Mathprompter: Mathe- matical reasoning using large language models,” arXiv preprint arXiv:2303.05398, 2023

  67. [75]

    Synergizing spatial optimization with large language models for open-domain urban itinerary planning,

    Y. Tang, Z. Wang, A. Qu, Y. Yan, K. Hou, D. Zhuang, X. Guo, J. Zhao, Z. Zhao, and W. Ma, “Synergizing spatial optimization with large language models for open-domain urban itinerary planning,” arXiv preprint arXiv:2402.07204, 2024

  68. [76]

    Destinai: Your personalized travel itinerary planner and chat catalyst using generative ai,

    A. Kanhed, P . Bhagwat, K. Palande, N. Palav, S. Balpande, K. Deshpande, and U. Patil, “Destinai: Your personalized travel itinerary planner and chat catalyst using generative ai,” in In- ternational Conference on Information and Communication Technology for Intelligent System...

  69. [77]

    Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

    Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems , vol. 36, 2024

  70. [78]

    Ef- ficienteqa: An efficient approach for open vocabulary embodied question answering,

    K. Cheng, Z. Li, X. Sun, B.-C. Min, A. S. Bedi, and A. Bera, “Ef- ficienteqa: An efficient approach for open vocabulary embodied question answering,” arXiv preprint arXiv:2410.20263, 2024

  71. [79]

    Openeqa: Embodied question answering in the era of foun- dation models,

    A. Majumdar, A. Ajay, X. Zhang, P . Putta, S. Yenamandra, M. Henaff, S. Silwal, P . Mcvay, O. Maksymets, S. Arnaud et al., “Openeqa: Embodied question answering in the era of foun- dation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...

  72. [80]

    Wordcraft: story writing with large language models,

    A. Yuan, A. Coenen, E. Reif, and D. Ippolito, “Wordcraft: story writing with large language models,” in Proceedings of the 27th International Conference on Intelligent User Interfaces, 2022, pp. 841– 852

  73. [81]

    On the creativity of large language models,

    G. Franceschelli and M. Musolesi, “On the creativity of large language models,” AI & SOCIETY, pp. 1–11, 2024

  74. [82]

    Llm with tools: A survey,

    Z. Shen, “Llm with tools: A survey,” arXiv preprint arXiv:2409.18807, 2024

  75. [83]

    A review of prominent paradigms for llm-based agents: Tool use (including rag), planning, and feedback learning,

    X. Li, “A review of prominent paradigms for llm-based agents: Tool use (including rag), planning, and feedback learning,” arXiv preprint arXiv:2406.05804, 2024

  76. [84]

    Tool-lmm: A large multi-modal model for tool agent learning,

    C. Wang, W. Luo, Q. Chen, H. Mai, J. Guo, S. Dong, Z. Li, L. Ma, S. Gao et al., “Tool-lmm: A large multi-modal model for tool agent learning,” arXiv preprint arXiv:2401.10727, 2024

  77. [85]

    Retrieval-augmented generation for large language models: A survey,

    Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, and H. Wang, “Retrieval-augmented generation for large language models: A survey,” arXiv preprint arXiv:2312.10997, 2023

  78. [86]

    Kg-rag: Bridging the gap between knowledge and creativity,

    D. Sanmartin, “Kg-rag: Bridging the gap between knowledge and creativity,” arXiv preprint arXiv:2405.12035, 2024

  79. [87]

    Cognitive mirage: A review of hallucinations in large language models,

    H. Ye, T. Liu, A. Zhang, W. Hua, and W. Jia, “Cognitive mirage: A review of hallucinations in large language models,”arXiv preprint arXiv:2309.06794, 2023

  80. [88]

    Funda- mental limitations of alignment in large language models,

    Y. Wolf, N. Wies, O. Avnery, Y. Levine, and A. Shashua, “Funda- mental limitations of alignment in large language models,” arXiv preprint arXiv:2304.11082, 2023

  81. [89]

    Under- standing the capabilities, limitations, and societal impact of large language models,

    A. Tamkin, M. Brundage, J. Clark, and D. Ganguli, “Under- standing the capabilities, limitations, and societal impact of large language models,” arXiv preprint arXiv:2102.02503, 2021

  82. [90]

    A survey on large language models: Applications, challenges, limitations, and practical us- age,

    M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al., “A survey on large language models: Applications, challenges, limitations, and practical us- age,” Authorea Preprints, 2023

  83. [91]

    Causal parrots: Large language models may talk causality but are not causal,

    M. Zeˇ cevi´ c, M. Willig, D. S. Dhami, and K. Kersting, “Causal parrots: Large language models may talk causality but are not causal,” arXiv preprint arXiv:2308.13067, 2023

  84. [92]

    On the dangers of stochastic parrots: Can language models be 29 too big???

    E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be 29 too big???” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 2021, pp. 610–623

  85. [93]

    Evolving embodied intelligence from materials to machines,

    D. Howard, A. E. Eiben, D. F. Kennedy, J.-B. Mouret, P . Valencia, and D. Winkler, “Evolving embodied intelligence from materials to machines,” Nature Machine Intelligence, vol. 1, no. 1, pp. 12–19, 2019

  86. [94]

    Against ai understanding and sentience: large language models, meaning, and the patterns of human language use,

    C. Durt, T. Froese, and T. Fuchs, “Against ai understanding and sentience: large language models, meaning, and the patterns of human language use,” 2023

  87. [95]

    The symbol grounding problem,

    S. Harnad, “The symbol grounding problem,” Physica D: Nonlin- ear Phenomena, vol. 42, no. 1-3, pp. 335–346, 1990

  88. [96]

    The development of human causal learning and reasoning,

    M. K. Goddu and A. Gopnik, “The development of human causal learning and reasoning,” Nature Reviews Psychology , pp. 1–21, 2024

  89. [97]

    Marwala, Causality, correlation and artificial intelligence for ratio- nal decision making

    T. Marwala, Causality, correlation and artificial intelligence for ratio- nal decision making. World Scientific, 2015

  90. [98]

    Attention and working memory as predictors of intelligence,

    K. Schweizer and H. Moosbrugger, “Attention and working memory as predictors of intelligence,” Intelligence, vol. 32, no. 4, pp. 329–347, 2004

  91. [99]

    Working memory ca- pacity and its relation to general intelligence,

    A. R. Conway, M. J. Kane, and R. W. Engle, “Working memory ca- pacity and its relation to general intelligence,” Trends in cognitive sciences, vol. 7, no. 12, pp. 547–552, 2003

  92. [100]

    Intelligence, learning and long-term memory,

    J. Alexander and S. Smales, “Intelligence, learning and long-term memory,” Personality and Individual Differences , vol. 23, no. 5, pp. 815–825, 1997

  93. [101]

    Embodied human language models vs. large language models, or why artificial intelligence cannot explain the modal be able to,

    S. Torres-Martínez, “Embodied human language models vs. large language models, or why artificial intelligence cannot explain the modal be able to,” Biosemiotics, pp. 1–25, 2024

  94. [102]

    Embodied ai with large language models: A survey and new hri framework,

    M.-Y. Lin, O.-W. Lee, and C.-Y. Lu, “Embodied ai with large language models: A survey and new hri framework,” in 2024 International Conference on Advanced Robotics and Mechatronics (ICARM). IEEE, 2024, pp. 978–983

  95. [103]

    Could a large language model be conscious?

    D. J. Chalmers, “Could a large language model be conscious?” arXiv preprint arXiv:2303.07103, 2023

  96. [104]

    The principles of goal-directed decision-making: from neural mechanisms to computation and robotics,

    G. Pezzulo, P . F. Verschure, C. Balkenius, and C. M. Pennartz, “The principles of goal-directed decision-making: from neural mechanisms to computation and robotics,” p. 20130470, 2014

  97. [105]

    Motor cognition: The role of sentience in perception and action,

    E. Morsella, A. G. Velasquez, J. K. Yankulova, Y. Li, C. Y. Wong, and D. Lambert, “Motor cognition: The role of sentience in perception and action,” Kinesiology Review, vol. 9, no. 3, pp. 261– 274, 2020

  98. [106]

    Transforming agency. on the mode of existence of large language models,

    X. E. Barandiaran and L. S. Almendros, “Transforming agency. on the mode of existence of large language models,” arXiv preprint arXiv:2407.10735, 2024

  99. [107]

    Grounding large language models in interactive en- vironments with online reinforcement learning,

    T. Carta, C. Romac, T. Wolf, S. Lamprier, O. Sigaud, and P .-Y. Oudeyer, “Grounding large language models in interactive en- vironments with online reinforcement learning,” in International Conference on Machine Learning. PMLR, 2023, pp. 3676–3713

  100. [108]

    Intuitive physics and cognitive algebra: A review,

    M. Vicovaro, “Intuitive physics and cognitive algebra: A review,” European Review of Applied Psychology , vol. 71, no. 5, p. 100610, 2021

  101. [109]

    A survey of em- bodied ai: From simulators to research tasks,

    J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan, “A survey of em- bodied ai: From simulators to research tasks,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 6, no. 2, pp. 230–244, 2022

  102. [110]

    Are intuitive physics and intuitive psychology indepen- dent? a test with children with asperger syndrome,

    S. Baron-Cohen, S. Wheelwright, A. Spong, V . Scahill, J. Lawson et al. , “Are intuitive physics and intuitive psychology indepen- dent? a test with children with asperger syndrome,” Journal of developmental and learning disorders, vol. 5, no. 1, pp. 47–78, 2001

  103. [111]

    The intuitive psychologist and his shortcomings: Dis- tortions in the attribution process,

    L. Ross, “The intuitive psychologist and his shortcomings: Dis- tortions in the attribution process,” in Advances in experimental social psychology. Academic Press, 1977, vol. 10, pp. 173–220

  104. [112]

    Theory of mind abilities of large language models in human-robot interaction: An illusion?

    M. Verma, S. Bhambri, and S. Kambhampati, “Theory of mind abilities of large language models in human-robot interaction: An illusion?” in Companion of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, 2024, pp. 36–45

  105. [113]

    Theory of mind for multi- agent collaboration via large language models,

    H. Li, Y. Q. Chong, S. Stepputtis, J. Campbell, D. Hughes, M. Lewis, and K. Sycara, “Theory of mind for multi- agent collaboration via large language models,” arXiv preprint arXiv:2310.10701, 2023

  106. [114]

    What you need is what you get: Theory of mind for an llm-based code understanding assistant,

    J. Richards and M. Wessel, “What you need is what you get: Theory of mind for an llm-based code understanding assistant,” arXiv preprint arXiv:2408.04477, 2024

  107. [115]

    Metacognition is all you need? using introspection in generative agents to improve goal-directed behavior,

    J. Toy, J. MacAdam, and P . Tabor, “Metacognition is all you need? using introspection in generative agents to improve goal-directed behavior,” arXiv preprint arXiv:2401.10910, 2024

  108. [116]

    Recursive intro- spection: Teaching llm agents how to self-improve,

    Y. Qu, T. Zhang, N. Garg, and A. Kumar, “Recursive intro- spection: Teaching llm agents how to self-improve,” in ICML 2024 Workshop on Structured Probabilistic Inference {\&} Generative Modeling, 2024

  109. [117]

    Retroformer: Retrospective large language agents with policy gradient opti- mization,

    W. Yao, S. Heinecke, J. C. Niebles, Z. Liu, Y. Feng, L. Xue, R. Murthy, Z. Chen, J. Zhang, D. Arpit et al. , “Retroformer: Retrospective large language agents with policy gradient opti- mization,” arXiv preprint arXiv:2308.02151, 2023

  110. [118]

    Theoretical strategies for an embodied cognitive neuroscience: Mechanistic explanations of brain-body-environment systems,

    D. Mougenot and H. Matheson, “Theoretical strategies for an embodied cognitive neuroscience: Mechanistic explanations of brain-body-environment systems,” Cognitive neuroscience, vol. 15, no. 3-4, pp. 85–97, 2024

  111. [119]

    Embodied cognition is not what you think it is,

    A. D. Wilson and S. Golonka, “Embodied cognition is not what you think it is,” Frontiers in psychology, vol. 4, p. 58, 2013

  112. [120]

    Where does cognition occur: in ones head or in ones embodied/extended environment?

    P . E. Tibbetts, “Where does cognition occur: in ones head or in ones embodied/extended environment?” The Quarterly Review of Biology, vol. 89, no. 4, pp. 359–368, 2014

  113. [121]

    Embodied cognition,

    L. Foglia and R. A. Wilson, “Embodied cognition,” Wiley Inter- disciplinary Reviews: Cognitive Science , vol. 4, no. 3, pp. 319–325, 2013

  114. [122]

    Entangled cognition: Exploring the links between mind, body, and environment in the era of aied,

    L. Hsu, “Entangled cognition: Exploring the links between mind, body, and environment in the era of aied,” Psychology Research and Practice, vol. 3, no. 2, 2024

  115. [123]

    Neurotrophins and neuronal plasticity,

    H. Thoenen, “Neurotrophins and neuronal plasticity,” Science, vol. 270, no. 5236, pp. 593–598, 1995

  116. [124]

    P . R. Huttenlocher, Neural plasticity: The effects of environment on the development of the cerebral cortex . Harvard University Press, 2009

  117. [125]

    Embodied artificial intelligence: Trends and challenges,

    R. Pfeifer and F. Iida, “Embodied artificial intelligence: Trends and challenges,” Lecture notes in computer science, pp. 1–26, 2004

  118. [126]

    Mechanistic versus phenomenal embodiment: Can robot embodiment lead to strong ai?

    N. E. Sharkey and T. Ziemke, “Mechanistic versus phenomenal embodiment: Can robot embodiment lead to strong ai?”Cognitive Systems Research, vol. 2, no. 4, pp. 251–262, 2001

  119. [127]

    Habitat: A platform for embodied ai research,

    M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V . Koltun, J. Maliket al., “Habitat: A platform for embodied ai research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9339–9347

  120. [128]

    A little less conversation, a little more action, please: Investigating the physical common-sense of llms in a 3d embodied environment,

    M. G. Mecattaf, B. Slater, M. Teši´ c, J. Prunty, K. Voudouris, and L. G. Cheke, “A little less conversation, a little more action, please: Investigating the physical common-sense of llms in a 3d embodied environment,” arXiv preprint arXiv:2410.23242, 2024

  121. [129]

    Aligning cyber space with physical world: A comprehensive survey on embodied ai,

    Y. Liu, W. Chen, Y. Bai, X. Liang, G. Li, W. Gao, and L. Lin, “Aligning cyber space with physical world: A comprehensive survey on embodied ai,” arXiv preprint arXiv:2407.06886, 2024

  122. [130]

    Large language models for human-robot interaction: A review,

    C. Zhang, J. Chen, J. Li, Y. Peng, and Z. Mao, “Large language models for human-robot interaction: A review,” Biomimetic Intel- ligence and Robotics, p. 100131, 2023

  123. [131]

    K. M. Lee, Y. Jung, J. Kim, and S. R. Kim, “Are physically em- bodied social agents better than disembodied social agents?: The effects of physical embodiment, tactile interaction, and people’s loneliness in human–robot interaction,” International journal of human-computer stu...

  124. [132]

    Does artificial intelligence have agency?

    D. Swanepoel, “Does artificial intelligence have agency?” The mind-technology problem: Investigating minds, selves and 21st century artefacts, pp. 83–104, 2021

  125. [133]

    Artificial intelligence as digital agency,

    P . J. Ågerfalk, “Artificial intelligence as digital agency,” European Journal of Information Systems, vol. 29, no. 1, pp. 1–8, 2020

  126. [134]

    Belief, values, bias, and agency: Development of and entanglement with artificial intelligence,

    D. P . Williams, “Belief, values, bias, and agency: Development of and entanglement with artificial intelligence,” Ph.D. dissertation, Virginia Polytechnic Institute and State University, 2022

  127. [135]

    Lindblom, Embodied social cognition

    J. Lindblom, Embodied social cognition. Springer, 2015, vol. 26

  128. [136]

    Voestermans and T

    P . Voestermans and T. Verheggen,Culture as embodiment: The social tuning of behavior. John Wiley & Sons, 2013

  129. [137]

    Coca: Regaining safety-awareness of multimodal large language models with constitutional calibration,

    J. Gao, R. Pi, T. Han, H. Wu, L. Hong, L. Kong, X. Jiang, and Z. Li, “Coca: Regaining safety-awareness of multimodal large language models with constitutional calibration,” arXiv preprint arXiv:2409.11365, 2024

  130. [138]

    Ethical reasoning and moral value alignment of llms de- pend on the language we prompt them in,

    U. Agarwal, K. Tanmay, A. Khandelwal, and M. Choudhury, “Ethical reasoning and moral value alignment of llms de- pend on the language we prompt them in,” arXiv preprint arXiv:2404.18460, 2024

  131. [139]

    Prediction of goal-directed behavior: Attitudes, intentions, and perceived behavioral control,

    I. Ajzen and T. J. Madden, “Prediction of goal-directed behavior: Attitudes, intentions, and perceived behavioral control,” Journal of experimental social psychology, vol. 22, no. 5, pp. 453–474, 1986

  132. [140]

    Goal directed behavior: The concept of action in psychology,

    M. Frese and J. Sabini, “Goal directed behavior: The concept of action in psychology,” 2021. 30

  133. [141]

    Habits as knowledge structures: automaticity in goal-directed behavior

    H. Aarts and A. Dijksterhuis, “Habits as knowledge structures: automaticity in goal-directed behavior.” Journal of personality and social psychology, vol. 78, no. 1, p. 53, 2000

  134. [142]

    Intelligence and the frontal lobe: The organization of goal- directed behavior,

    J. Duncan, H. Emslie, P . Williams, R. Johnson, and C. Freer, “Intelligence and the frontal lobe: The organization of goal- directed behavior,” Cognitive psychology , vol. 30, no. 3, pp. 257– 303, 1996

  135. [143]

    Position: Levels of agi for operationalizing progress on the path to agi,

    M. R. Morris, J. Sohl-Dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, C. Farabet, and S. Legg, “Position: Levels of agi for operationalizing progress on the path to agi,” in Forty-first Inter- national Conference on Machine Learning

  136. [144]

    Xcs for self-awareness in autonomous computing systems

    T. Hansmeier, “Xcs for self-awareness in autonomous computing systems.” Ph.D. dissertation, University of Paderborn, Germany, 2023

  137. [145]

    Golem: towards an agi meta-architecture enabling both goal preservation and radical self-improvement,

    B. Goertzel, “Golem: towards an agi meta-architecture enabling both goal preservation and radical self-improvement,” Journal of Experimental & Theoretical Artificial Intelligence , vol. 26, no. 3, pp. 391–403, 2014

  138. [146]

    A model of prefrontal cortical mechanisms for goal-directed behavior,

    M. E. Hasselmo, “A model of prefrontal cortical mechanisms for goal-directed behavior,” Journal of cognitive neuroscience , vol. 17, no. 7, pp. 1115–1129, 2005

  139. [147]

    The motivational factors underlying delay discounting

    C. E. Köpetz, J. L. Briskin, R. Sultana, and S. C. Stanciu, “The motivational factors underlying delay discounting.” Motivation Science, vol. 7, no. 3, p. 264, 2021

  140. [148]

    On the complexity of exploration in goal-driven navigation,

    M. Al-Shedivat, L. Lee, R. Salakhutdinov, and E. Xing, “On the complexity of exploration in goal-driven navigation,” arXiv preprint arXiv:1811.06889, 2018

  141. [149]

    A diffusion model for the congruency sequence effect,

    C. Luo and R. W. Proctor, “A diffusion model for the congruency sequence effect,” Psychonomic Bulletin & Review, vol. 29, no. 6, pp. 2034–2051, 2022

  142. [150]

    Goaliath: A theory of goal-directed behavior,

    B. Hommel, “Goaliath: A theory of goal-directed behavior,” Psychological Research, vol. 86, no. 4, pp. 1054–1077, 2022

  143. [151]

    Selfgoal: Your language agents already know how to achieve high-level goals,

    R. Yang, J. Chen, Y. Zhang, S. Yuan, A. Chen, K. Richard- son, Y. Xiao, and D. Yang, “Selfgoal: Your language agents already know how to achieve high-level goals,” arXiv preprint arXiv:2406.04784, 2024

  144. [152]

    Can large language models reason about goal-oriented tasks?

    F. Bellos, Y. Li, W. Liu, and J. Corso, “Can large language models reason about goal-oriented tasks?” in Proceedings of the First edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024), 2024, pp. 24–34

  145. [153]

    Do llms selectively encode the goal of an agent’s reach?

    L. Ruis, A. Findeis, H. Bradley, H. A. Rahmani, K. W. Choe, E. Grefenstette, and T. Rocktäschel, “Do llms selectively encode the goal of an agent’s reach?” in First Workshop on Theory of Mind in Communicating Agents, 2023

  146. [154]

    I think, therefore i am: Awareness in large language models,

    Y. Li, Y. Huang, Y. Lin, S. Wu, Y. Wan, and L. Sun, “I think, therefore i am: Awareness in large language models,” arXiv preprint arXiv:2401.17882, 2024

  147. [155]

    Incorporating exter- nal knowledge and goal guidance for llm-based conversational recommender systems,

    C. Li, Y. Deng, H. Hu, M.-Y. Kan, and H. Li, “Incorporating exter- nal knowledge and goal guidance for llm-based conversational recommender systems,” arXiv preprint arXiv:2405.01868, 2024

  148. [156]

    Goal awareness for conversational ai: proactivity, non-collaborativity, and beyond,

    Y. Deng, W. Lei, M. Huang, and T.-S. Chua, “Goal awareness for conversational ai: proactivity, non-collaborativity, and beyond,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 6: Tutorial Abstracts) , 2023, pp. 1–10

  149. [157]

    Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond,

    ——, “Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond,” in Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region, 2023, pp. 298–301

  150. [158]

    Hitkg: Towards goal-oriented conversations via multi-hierarchy learn- ing,

    J. Ni, V . Pandelea, T. Young, H. Zhou, and E. Cambria, “Hitkg: Towards goal-oriented conversations via multi-hierarchy learn- ing,” in Proceedings of the AAAI conference on artificial intelligence , vol. 36, no. 10, 2022, pp. 11 112–11 120

  151. [159]

    Towards goal-oriented large language model prompting: A survey,

    H. Li, J. Leung, and Z. Shen, “Towards goal-oriented large language model prompting: A survey,” arXiv preprint arXiv:2401.14043, 2024

  152. [160]

    Large language models still can’t plan (a benchmark for llms on planning and reasoning about change),

    K. Valmeekam, A. Olmo, S. Sreedharan, and S. Kambhampati, “Large language models still can’t plan (a benchmark for llms on planning and reasoning about change),” in NeurIPS 2022 Foundation Models for Decision Making Workshop , 2022

  153. [161]

    Value fulcra: Mapping large language models to the multidimensional spectrum of basic human values,

    J. Yao, X. Yi, X. Wang, Y. Gong, and X. Xie, “Value fulcra: Mapping large language models to the multidimensional spectrum of basic human values,” arXiv preprint arXiv:2311.10766, 2023

  154. [162]

    Keyword-guided neural conversational model,

    P . Zhong, Y. Liu, H. Wang, and C. Miao, “Keyword-guided neural conversational model,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 16, 2021, pp. 14 568–14 576

  155. [163]

    A unified multi-task learning framework for multi-goal conversa- tional recommender systems,

    Y. Deng, W. Zhang, W. Xu, W. Lei, T.-S. Chua, and W. Lam, “A unified multi-task learning framework for multi-goal conversa- tional recommender systems,” ACM Transactions on Information Systems, vol. 41, no. 3, pp. 1–25, 2023

  156. [164]

    Graph-grounded goal planning for con- versational recommendation,

    Z. Liu, D. Zhou, H. Liu, H. Wang, Z.-Y. Niu, H. Wu, W. Che, T. Liu, and H. Xiong, “Graph-grounded goal planning for con- versational recommendation,” IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 5, pp. 4923–4939, 2022

  157. [165]

    Discover- ing language model behaviors with model-written evaluations,

    E. Perez, S. Ringer, K. Lukoši ¯ut˙e, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath et al. , “Discover- ing language model behaviors with model-written evaluations,” arXiv preprint arXiv:2212.09251, 2022

  158. [166]

    Towards automatic learning of procedures from web instructional videos,

    L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018

  159. [167]

    Gomaa-geo: Goal modality agnostic active geo- localization,

    A. Sarkar, S. Sastry, A. Pirinen, C. Zhang, N. Jacobs, and Y. Vorobeychik, “Gomaa-geo: Goal modality agnostic active geo- localization,” arXiv preprint arXiv:2406.01917, 2024

  160. [168]

    Context-aware language modeling for goal-oriented dialogue systems,

    C. Snell, M. Yang, J. Fu, Y. Su, and S. Levine, “Context-aware language modeling for goal-oriented dialogue systems,” arXiv preprint arXiv:2204.10198, 2022

  161. [169]

    Instruc- tion following with goal-conditioned reinforcement learning in virtual environments,

    Z. Volovikova, A. Skrynnik, P . Kuderov, and A. I. Panov, “Instruc- tion following with goal-conditioned reinforcement learning in virtual environments,” in ECAI 2024. IOS Press, 2024, pp. 650– 657

  162. [170]

    Mp5: A multi-modal open-ended embodied system in minecraft via active perception,

    Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, Y. Qiao, and J. Shao, “Mp5: A multi-modal open-ended embodied system in minecraft via active perception,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2024, pp. 16 307–16 316

  163. [171]

    Reinforcement learning from llm feedback to counteract goal misgeneralization,

    H. N. E. Barj and T. Sautory, “Reinforcement learning from llm feedback to counteract goal misgeneralization,” arXiv preprint arXiv:2401.07181, 2024

  164. [172]

    Embodiedgpt: Vision-language pre-training via embodied chain of thought,

    Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P . Luo, “Embodiedgpt: Vision-language pre-training via embodied chain of thought,” Advances in Neural Information Processing Systems, vol. 36, 2024

  165. [173]

    Generate subgoal images before act: Unlocking the chain-of-thought reasoning in diffusion model for robot manipu- lation with multimodal prompts,

    F. Ni, J. Hao, S. Wu, L. Kou, J. Liu, Y. Zheng, B. Wang, and Y. Zhuang, “Generate subgoal images before act: Unlocking the chain-of-thought reasoning in diffusion model for robot manipu- lation with multimodal prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vis...

  166. [174]

    Pre-training goal-based mod- els for sample-efficient reinforcement learning,

    H. Yuan, Z. Mu, F. Xie, and Z. Lu, “Pre-training goal-based mod- els for sample-efficient reinforcement learning,” in The Twelfth International Conference on Learning Representations, 2024

  167. [175]

    Guest editorial- special issue on situation, activity, and goal awareness in cyber- physical human–machine systems,

    L. Chen, D. J. Cook, B. Guo, and W. Leister, “Guest editorial- special issue on situation, activity, and goal awareness in cyber- physical human–machine systems,” IEEE Transactions on Human- Machine Systems, vol. 47, no. 3, pp. 305–309, 2017

  168. [176]

    Situation, activity and goal awareness in ubiquitous computing,

    L. Chen and P . Rashidi, “Situation, activity and goal awareness in ubiquitous computing,” International Journal of Pervasive Com- puting and Communications, vol. 8, no. 3, pp. 216–224, 2012

  169. [177]

    Improving conversational recommender system via contextual and time- aware modeling with less domain-specific knowledge,

    L. Wang, S. Joty, W. Gao, X. Zeng, and K.-F. Wong, “Improving conversational recommender system via contextual and time- aware modeling with less domain-specific knowledge,” IEEE Transactions on Knowledge and Data Engineering, 2024

  170. [178]

    Large language models encode clinical knowledge,

    K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl et al. , “Large language models encode clinical knowledge,” Nature, vol. 620, no. 7972, pp. 172–180, 2023

  171. [179]

    Gpt understands, too,

    X. Liu, Y. Zheng, Z. Du, M. Ding, Y. Qian, Z. Yang, and J. Tang, “Gpt understands, too,” AI Open, vol. 5, pp. 208–215, 2024

  172. [180]

    A survey on evaluation of large language models,

    Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al., “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, no. 3, pp. 1–45, 2024

  173. [181]

    Towards generalist biomedical ai,

    T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P .-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena et al. , “Towards generalist biomedical ai,” NEJM AI, vol. 1, no. 3, p. AIoa2300138, 2024

  174. [182]

    Hypo- thetical minds: Scaffolding theory of mind for multi-agent tasks with large language models,

    L. Cross, V . Xiang, A. Bhatia, D. L. Yamins, and N. Haber, “Hypo- thetical minds: Scaffolding theory of mind for multi-agent tasks with large language models,” arXiv preprint arXiv:2407.07086 , 2024

  175. [183]

    Applying oppo- nent and environment modelling in decentralised multi-agent 31 reinforcement learning,

    A. Chernyavskiy, A. Skrynnik, and A. Panov, “Applying oppo- nent and environment modelling in decentralised multi-agent 31 reinforcement learning,” Cognitive Systems Research , p. 101306, 2024

  176. [184]

    Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,

    Y. Huang, G. Chen, J. Xu, M. Zhang, L. Yang, B. Pei, H. Zhang, L. Dong, Y. Wang, L. Wang et al. , “Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,” in Proceedings of the IEEE/CVF Conference on Computer Vision an...

  177. [185]

    Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world,

    X. Wang, T. Kwon, M. Rad, B. Pan, I. Chakraborty, S. Andrist, D. Bohus, A. Feniello, B. Tekin, F. V . Frujeriet al., “Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world,” in Proceedings of the IEEE/CVF International Conference o...

  178. [186]

    Ego- tracks: A long-term egocentric visual object tracking dataset,

    H. Tang, K. J. Liang, K. Grauman, M. Feiszli, and W. Wang, “Ego- tracks: A long-term egocentric visual object tracking dataset,” Advances in Neural Information Processing Systems , vol. 36, 2024

  179. [187]

    Egochoir: Capturing 3d human-object interaction regions from egocentric views,

    Y. Yang, W. Zhai, C. Wang, C. Yu, Y. Cao, and Z.-J. Zha, “Egochoir: Capturing 3d human-object interaction regions from egocentric views,” arXiv preprint arXiv:2405.13659, 2024

  180. [188]

    Alanavlm: A multimodal embodied ai foundation model for egocentric video understanding,

    A. Suglia, C. Greco, K. Baker, J. L. Part, I. Papaioannou, A. Eshghi, I. Konstas, and O. Lemon, “Alanavlm: A multimodal embodied ai foundation model for egocentric video understanding,” arXiv preprint arXiv:2406.13807, 2024

  181. [189]

    Towards continual egocentric activity recognition: A multi-modal egocentric activity dataset for continual learning,

    L. Xu, Q. Wu, L. Pan, F. Meng, H. Li, C. He, H. Wang, S. Cheng, and Y. Dai, “Towards continual egocentric activity recognition: A multi-modal egocentric activity dataset for continual learning,” IEEE Transactions on Multimedia, 2023

  182. [190]

    Emhi: A multimodal egocentric human mo- tion dataset with hmd and body-worn imus,

    Z. Fan, P . Dai, Z. Su, X. Gao, Z. Lv, J. Zhang, T. Du, G. Wang, and Y. Zhang, “Emhi: A multimodal egocentric human mo- tion dataset with hmd and body-worn imus,” arXiv preprint arXiv:2408.17168, 2024

  183. [191]

    Egocentric+: A multipurpose data set for head-mounted wearable computing devices,

    Y. Ashok and M. K. Rohil, “Egocentric+: A multipurpose data set for head-mounted wearable computing devices,” in Proceedings of the 2024 International Conference on Advanced Visual Interfaces , 2024, pp. 1–5

  184. [192]

    Towards versatile embodied navigation,

    H. Wang, W. Liang, L. V . Gool, and W. Wang, “Towards versatile embodied navigation,” Advances in neural information processing systems, vol. 35, pp. 36 858–36 874, 2022

  185. [193]

    Out of the box: embodied navigation in the real world,

    R. Bigazzi, F. Landi, M. Cornia, S. Cascianelli, L. Baraldi, and R. Cucchiara, “Out of the box: embodied navigation in the real world,” in Computer Analysis of Images and Patterns: 19th International Conference, CAIP 2021, Virtual Event, September 28– 30, 2021, Proceedings, Pa...

  186. [194]

    Vision-language navigation with embodied intelligence: A survey,

    P . Gao, P . Wang, F. Gao, F. Wang, and R. Yuan, “Vision-language navigation with embodied intelligence: A survey,” arXiv preprint arXiv:2402.14304, 2024

  187. [195]

    Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,

    X. Li, M. Zhang, Y. Geng, H. Geng, Y. Long, Y. Shen, R. Zhang, J. Liu, and H. Dong, “Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1...

  188. [196]

    Empowering embodied manipulation: A bimanual-mobile robot manipulation dataset for household tasks,

    T. Zhang, D. Li, Y. Li, Z. Zeng, L. Zhao, L. Sun, Y. Chen, X. Wei, Y. Zhan, L. Li et al. , “Empowering embodied manipulation: A bimanual-mobile robot manipulation dataset for household tasks,” arXiv preprint arXiv:2405.18860, 2024

  189. [197]

    Embodied communication: How robots and people communicate through physical interaction,

    A. Kalinowska, P . M. Pilarski, and T. D. Murphey, “Embodied communication: How robots and people communicate through physical interaction,” Annual review of control, robotics, and au- tonomous systems, vol. 6, no. 1, pp. 205–232, 2023

  190. [198]

    Adapting the number of questions based on detected psychological distress for cognitive behavioral therapy with an embodied conversational agent: Comparative study,

    K. Shidara, H. Tanaka, H. Adachi, D. Kanayama, T. Kudo, and S. Nakamura, “Adapting the number of questions based on detected psychological distress for cognitive behavioral therapy with an embodied conversational agent: Comparative study,” JMIR Formative Research, vol. 8, p. e...

  191. [199]

    Benchmarking 2d egocentric hand pose datasets,

    O. Taran, D. M. Manzone, and J. Zariffa, “Benchmarking 2d egocentric hand pose datasets,” arXiv preprint arXiv:2409.07337 , 2024

  192. [200]

    Ego- centric human-object interaction detection exploiting synthetic data,

    R. Leonardi, F. Ragusa, A. Furnari, and G. M. Farinella, “Ego- centric human-object interaction detection exploiting synthetic data,” in International Conference on Image Analysis and Processing . Springer, 2022, pp. 237–248

  193. [201]

    Hup-3d: A 3d multi-view synthetic dataset for assisted-egocentric hand-ultrasound pose estimation,

    M. Birlo, R. Caramalau, P . J. Edwards, B. Dromey, M. J. Clarkson, D. Stoyanov et al. , “Hup-3d: A 3d multi-view synthetic dataset for assisted-egocentric hand-ultrasound pose estimation,” arXiv preprint arXiv:2407.09215, 2024

  194. [202]

    Leap: Llm-generation of egocentric action programs,

    E. Dessalene, M. Maynord, C. Fermüller, and Y. Aloimonos, “Leap: Llm-generation of egocentric action programs,” arXiv preprint arXiv:2312.00055, 2023

  195. [203]

    Egogen: An egocentric synthetic data generator,

    G. Li, K. Zhao, S. Zhang, X. Lyu, M. Dusmanu, Y. Zhang, M. Pollefeys, and S. Tang, “Egogen: An egocentric synthetic data generator,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 497–14 509

  196. [204]

    Parse-ego4d: Personal action recom- mendation suggestions for egocentric videos,

    S. Abreu, T. D. Do, K. Ahuja, E. J. Gonzalez, L. Payne, D. McDuff, and M. Gonzalez-Franco, “Parse-ego4d: Personal action recom- mendation suggestions for egocentric videos,” arXiv preprint arXiv:2407.09503, 2024

  197. [205]

    A survey on large language model based autonomous agents,

    L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al., “A survey on large language model based autonomous agents,” Frontiers of Computer Science, vol. 18, no. 6, p. 186345, 2024

  198. [206]

    Proagent: building proactive cooperative agents with large language models,

    C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu et al., “Proagent: building proactive cooperative agents with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 591–17 599

  199. [207]

    Large language model based multi- agents: A survey of progress and challenges,

    T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V . Chawla, O. Wiest, and X. Zhang, “Large language model based multi- agents: A survey of progress and challenges,” arXiv preprint arXiv:2402.01680, 2024

  200. [208]

    Building cooperative embodied agents modularly with large language models,

    H. Zhang, W. Du, J. Shan, Q. Zhou, Y. Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building cooperative embodied agents modularly with large language models,” arXiv preprint arXiv:2307.02485 , 2023

  201. [209]

    Llm-agent-umf: Llm-based agent unified modeling framework for seamless in- tegration of multi active/passive core-agents,

    A. B. Hassouna, H. Chaari, and I. Belhaj, “Llm-agent-umf: Llm-based agent unified modeling framework for seamless in- tegration of multi active/passive core-agents,” arXiv preprint arXiv:2409.11393, 2024

  202. [210]

    A human-computer collaborative tool for training a single large language model agent into a network through few examples,

    L. Pan, Y. Li, C. Yu, and Y. Shi, “A human-computer collaborative tool for training a single large language model agent into a network through few examples,” arXiv preprint arXiv:2404.15974, 2024

  203. [211]

    Voicepilot: Harnessing llms as speech interfaces for physically assistive robots,

    A. Padmanabha, J. Yuan, J. Gupta, Z. Karachiwalla, C. Majidi, H. Admoni, and Z. Erickson, “Voicepilot: Harnessing llms as speech interfaces for physically assistive robots,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 2024, pp. 1–18

  204. [212]

    Llm-based multi-agent rein- forcement learning: Current and future directions,

    C. Sun, S. Huang, and D. Pompili, “Llm-based multi-agent rein- forcement learning: Current and future directions,” arXiv preprint arXiv:2405.11106, 2024

  205. [213]

    En- abling intelligent interactions between an agent and an llm: A re- inforcement learning approach,

    B. Hu, C. Zhao, P . Zhang, Z. Zhou, Y. Yang, Z. Xu, and B. Liu, “En- abling intelligent interactions between an agent and an llm: A re- inforcement learning approach,” arXiv preprint arXiv:2306.03604 , 2023

  206. [214]

    Dynamic llm-agent network: An llm-agent collaboration framework with agent team optimization,

    Z. Liu, Y. Zhang, P . Li, Y. Liu, and D. Yang, “Dynamic llm-agent network: An llm-agent collaboration framework with agent team optimization,” arXiv preprint arXiv:2310.02170, 2023

  207. [215]

    Multiply: A multisensory object-centric embodied large language model in 3d world,

    Y. Hong, Z. Zheng, P . Chen, Y. Wang, J. Li, and C. Gan, “Multiply: A multisensory object-centric embodied large language model in 3d world,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 406–26 416

  208. [216]

    See and think: Embodied agent in virtual environ- ment,

    Z. Zhao, W. Chai, X. Wang, B. Li, S. Hao, S. Cao, T. Ye, and G. Wang, “See and think: Embodied agent in virtual environ- ment,” in European Conference on Computer Vision. Springer, 2025, pp. 187–204

  209. [217]

    Mindagent: Emergent gaming interaction,

    R. Gong, Q. Huang, X. Ma, H. Vo, Z. Durante, Y. Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Feiet al., “Mindagent: Emergent gaming interaction,” arXiv preprint arXiv:2309.09971, 2023

  210. [218]

    Pybullet, a python module for physics simulation for games, robotics and machine learning,

    E. Coumans and Y. Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” 2016

  211. [219]

    Mhrc: Closed- loop decentralized multi-heterogeneous robot collaboration with large language models,

    W. Yu, J. Peng, Y. Ying, S. Li, J. Ji, and Y. Zhang, “Mhrc: Closed- loop decentralized multi-heterogeneous robot collaboration with large language models,” arXiv preprint arXiv:2409.16030, 2024

  212. [220]

    Grounding llms for robot task planning using closed-loop state feedback,

    V . Bhat, A. U. Kaypak, P . Krishnamurthy, R. Karri, and F. Khor- rami, “Grounding llms for robot task planning using closed-loop state feedback,” arXiv preprint arXiv:2402.08546, 2024

  213. [221]

    Cuify the xr: An open- source package to embed llm-powered conversational agents in xr,

    K. B. Buldu, S. Özdel, K. H. C. Lau, M. Wang, D. Saad, S. Schön- born, A. Boch, E. Kasneci, and E. Bozkir, “Cuify the xr: An open- source package to embed llm-powered conversational agents in xr,” arXiv preprint arXiv:2411.04671, 2024

  214. [222]

    Llmr: Real-time prompting of interactive worlds using large language models,

    F. De La Torre, C. M. Fang, H. Huang, A. Banburski-Fahey, J. Amores Fernandez, and J. Lanier, “Llmr: Real-time prompting of interactive worlds using large language models,” in Proceed- 32 ings of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–22

  215. [223]

    Envgen: Gener- ating and adapting environments via llms for training embodied agents,

    A. Zala, J. Cho, H. Lin, J. Yoon, and M. Bansal, “Envgen: Gener- ating and adapting environments via llms for training embodied agents,” arXiv preprint arXiv:2403.12014, 2024

  216. [224]

    Can language models serve as text-based world simulators?

    R. Wang, G. Todd, Z. Xiao, X. Yuan, M.-A. Côté, P . Clark, and P . Jansen, “Can language models serve as text-based world simulators?” arXiv preprint arXiv:2406.06485, 2024

  217. [225]

    Airsim: High-fidelity visual and physical simulation for autonomous vehicles,

    S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and Service Robotics: Results of the 11th International Conference . Springer, 2018, pp. 621–635

  218. [226]

    Ai2- thor: An interactive 3d environment for visual ai,

    E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Her- rasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu et al. , “Ai2- thor: An interactive 3d environment for visual ai,” arXiv preprint arXiv:1712.05474, 2017

  219. [227]

    Carla: An open urban driving simulator,

    A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun, “Carla: An open urban driving simulator,” in Conference on robot learning. PMLR, 2017, pp. 1–16

  220. [228]

    Eai-sim: An open- source embodied ai simulation framework with large language models,

    G. Liu, T. Sun, W. Li, X. Li, X. Liu, and J. Cui, “Eai-sim: An open- source embodied ai simulation framework with large language models,” in 2024 IEEE 18th International Conference on Control & Automation (ICCA). IEEE, 2024, pp. 994–999

  221. [229]

    Aeroverse: Uav-agent benchmark suite for simulating, pre-training, finetuning, and evaluating aerospace embodied world models,

    F. Yao, Y. Yue, Y. Liu, X. Sun, and K. Fu, “Aeroverse: Uav-agent benchmark suite for simulating, pre-training, finetuning, and evaluating aerospace embodied world models,” arXiv preprint arXiv:2408.15511, 2024

  222. [230]

    Objaverse: A universe of annotated 3d objects,

    M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. Van- derBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2023, pp. 13 1...

  223. [231]

    Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations,

    R. Gao, Y.-Y. Chang, S. Mall, L. Fei-Fei, and J. Wu, “Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations,” arXiv preprint arXiv:2109.07991, 2021

  224. [232]

    Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai,

    S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang et al., “Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai,” arXiv preprint arXiv:2109.08238, 2021

  225. [233]

    Shriek: a role playing game using unreal engine 4 and behaviour trees,

    S. Rodrigues, H. K. Rayat, R. M. Kurichithanam, and S. Rukhande, “Shriek: a role playing game using unreal engine 4 and behaviour trees,” in 2021 4th biennial international conference on nascent technologies in engineering (ICNTE) . IEEE, 2021, pp. 1–6

  226. [234]

    Exploring the viability of conversational ai for non-playable characters: A comprehensive survey,

    A. Mehta, Y. Kunjadiya, A. Kulkarni, and M. Nagar, “Exploring the viability of conversational ai for non-playable characters: A comprehensive survey,” in 2021 4th International Conference on Recent Trends in Computer Science and Technology (ICRTCST) . IEEE, 2022, pp. 96–102

  227. [235]

    Exploring presence in inter- actions with llm-driven npcs: A comparative study of speech recognition and dialogue options,

    F. R. Christiansen, L. N. Hollensberg, N. B. Jensen, K. Julsgaard, K. N. Jespersen, and I. Nikolov, “Exploring presence in inter- actions with llm-driven npcs: A comparative study of speech recognition and dialogue options,” in Proceedings of the 30th ACM Symposium on Virtual ...

  228. [236]

    A review of nine physics engines for reinforcement learning research,

    M. Kaup, C. Wolff, H. Hwang, J. Mayer, and E. Bruni, “A review of nine physics engines for reinforcement learning research,” arXiv preprint arXiv:2407.08590, 2024

  229. [237]

    Isaac gym: High performance gpu-based physics simulation for robot learning,

    V . Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa et al. , “Isaac gym: High performance gpu-based physics simulation for robot learning,” arXiv preprint arXiv:2108.10470, 2021

  230. [238]

    Difftactile: A physics-based differentiable tactile sim- ulator for contact-rich robotic manipulation,

    Z. Si, G. Zhang, Q. Ben, B. Romero, Z. Xian, C. Liu, and C. Gan, “Difftactile: A physics-based differentiable tactile sim- ulator for contact-rich robotic manipulation,” arXiv preprint arXiv:2403.08716, 2024

  231. [239]

    Ros- llm: A ros framework for embodied ai with task feedback and structured reasoning,

    C. E. Mower, Y. Wan, H. Yu, A. Grosnit, J. Gonzalez-Billandon, M. Zimmer, J. Wang, X. Zhang, Y. Zhao, A. Zhai et al. , “Ros- llm: A ros framework for embodied ai with task feedback and structured reasoning,” arXiv preprint arXiv:2406.19741, 2024

  232. [240]

    Lan- car: Leveraging language for context-aware robot locomotion in unstructured environments,

    C. L. Shek, X. Wu, D. Manocha, P . Tokekar, and A. S. Bedi, “Lan- car: Leveraging language for context-aware robot locomotion in unstructured environments,” arXiv preprint arXiv:2310.00481 , 2023

  233. [241]

    Embardiment: an embodied ai agent for productivity in xr,

    R. Bovo, S. Abreu, K. Ahuja, E. J. Gonzalez, L.-T. Cheng, and M. Gonzalez-Franco, “Embardiment: an embodied ai agent for productivity in xr,” arXiv preprint arXiv:2408.08158, 2024

  234. [242]

    Extended reality in iot scenarios: Concepts, applications and future trends,

    T. Andrade and D. Bastos, “Extended reality in iot scenarios: Concepts, applications and future trends,” in 2019 5th Experiment International Conference (exp. at’19). IEEE, 2019, pp. 107–112

  235. [243]

    Building llm-based ai agents in social virtual reality,

    H. Wan, J. Zhang, A. A. Suria, B. Yao, D. Wang, Y. Coady, and M. Prpa, “Building llm-based ai agents in social virtual reality,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2024, pp. 1–7

  236. [244]

    Y social: an llm-powered social media digital twin,

    G. Rossetti, M. Stella, R. Cazabet, K. Abramski, E. Cau, S. Cit- raro, A. Failla, R. Improta, V . Morini, and V . Pansanella, “Y social: an llm-powered social media digital twin,” arXiv preprint arXiv:2408.00818, 2024

  237. [245]

    Magicitem: Dynamic behav- ior design of virtual objects with large language models in a consumer metaverse platform,

    R. Kurai, T. Hiraki, Y. Hiroi, Y. Hirao, M. Perusquia-Hernandez, H. Uchiyama, and K. Kiyokawa, “Magicitem: Dynamic behav- ior design of virtual objects with large language models in a consumer metaverse platform,” arXiv preprint arXiv:2406.13242 , 2024

  238. [246]

    Pan- otree: Autonomous photo-spot explorer in virtual reality scenes,

    T. Hayase, S. Braun, H. Yanagawa, I. Orito, and Y. Hiroi, “Pan- otree: Autonomous photo-spot explorer in virtual reality scenes,” arXiv preprint arXiv:2405.17136, 2024

  239. [247]

    Dreamcodevr: Towards democratizing behavior design in virtual reality with speech-driven programming,

    D. Giunchi, N. Numan, E. Gatti, and A. Steed, “Dreamcodevr: Towards democratizing behavior design in virtual reality with speech-driven programming,” in 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR) . IEEE, 2024, pp. 579–589

  240. [248]

    Virtual reality and language models, a new frontier in learning,

    J. Izquierdo-Domenech, J. Linares-Pellicer, and I. Ferri-Molla, “Virtual reality and language models, a new frontier in learning,” 2024

  241. [249]

    Situationadapt: Contextual ui optimization in mixed reality with situation awareness via llm reasoning,

    Z. Li, C. Gebhardt, Y. Inglin, N. Steck, P . Streli, and C. Holz, “Situationadapt: Contextual ui optimization in mixed reality with situation awareness via llm reasoning,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 2024, pp. 1–13

  242. [250]

    Gui-world: A dataset for gui-oriented multimodal llm-based agents,

    D. Chen, Y. Huang, S. Wu, J. Tang, L. Chen, Y. Bai, Z. He, C. Wang, H. Zhou, Y. Li et al. , “Gui-world: A dataset for gui-oriented multimodal llm-based agents,” arXiv preprint arXiv:2406.10819 , 2024

  243. [251]

    Voxposer: Composable 3d value maps for robotic manipulation with language models,

    W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” arXiv preprint arXiv:2307.05973, 2023

  244. [252]

    Bytesized32: A corpus and challenge task for generating task- specific world models expressed as text games,

    R. Wang, G. Todd, E. Yuan, Z. Xiao, M.-A. Côté, and P . Jansen, “Bytesized32: A corpus and challenge task for generating task- specific world models expressed as text games,” arXiv preprint arXiv:2305.14879, 2023

  245. [253]

    Llm+ p: Empowering large language models with optimal planning proficiency,

    B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P . Stone, “Llm+ p: Empowering large language models with optimal planning proficiency,” arXiv preprint arXiv:2304.11477 , 2023

  246. [254]

    Embodied multi-modal agent trained by an llm from a parallel textworld,

    Y. Yang, T. Zhou, K. Li, D. Tao, L. Li, L. Shen, X. He, J. Jiang, and Y. Shi, “Embodied multi-modal agent trained by an llm from a parallel textworld,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 275–26 285

  247. [255]

    Scenecraft: An llm agent for synthesizing 3d scenes as blender code,

    Z. Hu, A. Iscen, A. Jain, T. Kipf, Y. Yue, D. A. Ross, C. Schmid, and A. Fathi, “Scenecraft: An llm agent for synthesizing 3d scenes as blender code,” in Forty-first International Conference on Machine Learning, 2024

  248. [256]

    Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment,

    H. Tang, D. Key, and K. Ellis, “Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment,” arXiv preprint arXiv:2402.12275, 2024

  249. [257]

    When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models,

    X. Ma, Y. Bhalgat, B. Smart, S. Chen, X. Li, J. Ding, J. Gu, D. Z. Chen, S. Peng, J.-W. Bianet al., “When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models,” arXiv preprint arXiv:2405.10255, 2024

  250. [258]

    In- struct2act: Mapping multi-modality instructions to robotic ac- tions with large language model,

    S. Huang, Z. Jiang, H. Dong, Y. Qiao, P . Gao, and H. Li, “In- struct2act: Mapping multi-modality instructions to robotic ac- tions with large language model,”arXiv preprint arXiv:2305.11176, 2023

  251. [259]

    Scenemotifcoder: Example-driven visual program learning for generating 3d object arrangements,

    H. I. I. Tam, H. I. D. Pun, A. T. Wang, A. X. Chang, and M. Savva, “Scenemotifcoder: Example-driven visual program learning for generating 3d object arrangements,” arXiv preprint arXiv:2408.02211, 2024

  252. [260]

    Craft an iron sword: Dynamically generating interactive game characters by prompting large language models tuned on code,

    R. Volum, S. Rao, M. Xu, G. DesGarennes, C. Brockett, B. Van Durme, O. Deng, A. Malhotra, and W. B. Dolan, “Craft an iron sword: Dynamically generating interactive game characters by prompting large language models tuned on code,” in Proceed- ings of the 3rd Wordplay: When Lan...

  253. [261]

    Spaceblender: Creating context-rich collaborative spaces through generative 3d scene blending,

    N. Numan, S. Rajaram, B. T. Kumaravel, N. Marquardt, and A. D. Wilson, “Spaceblender: Creating context-rich collaborative spaces through generative 3d scene blending,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 2024, pp. 1–25

  254. [262]

    Shape- it: Exploring text-to-shape-display for generative shape-changing behaviors with llms,

    W. Qian, C. Gao, A. Sathya, R. Suzuki, and K. Nakagaki, “Shape- it: Exploring text-to-shape-display for generative shape-changing behaviors with llms,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology , 2024, pp. 1– 29

  255. [263]

    Active lower limb prosthetics: a systematic review of design issues and solutions,

    M. Windrich, M. Grimmer, O. Christ, S. Rinderknecht, and P . Beckerle, “Active lower limb prosthetics: a systematic review of design issues and solutions,” Biomedical engineering online, vol. 15, pp. 5–19, 2016

  256. [264]

    Toward higher-performance bionic limbs for wider clinical use,

    D. Farina, I. Vujaklija, R. Brånemark, A. M. Bull, H. Dietl, B. Graimann, L. J. Hargrove, K.-P . Hoffmann, H. Huang, T. In- gvarsson et al. , “Toward higher-performance bionic limbs for wider clinical use,” Nature biomedical engineering , vol. 7, no. 4, pp. 473–485, 2023

  257. [265]

    Benchmarking retrieval-augmented large language models in biomedical nlp: Application, robustness, and self-awareness,

    M. Li, Z. Zhan, H. Yang, Y. Xiao, J. Huang, and R. Zhang, “Benchmarking retrieval-augmented large language models in biomedical nlp: Application, robustness, and self-awareness,” arXiv preprint arXiv:2405.08151, 2024

  258. [266]

    From persona to personalization: A survey on role-playing language agents,

    J. Chen, X. Wang, R. Xu, S. Yuan, Y. Zhang, W. Shi, J. Xie, S. Li, R. Yang, T. Zhuet al., “From persona to personalization: A survey on role-playing language agents,”arXiv preprint arXiv:2404.18231, 2024

  259. [267]

    Mm- sap: A comprehensive benchmark for assessing self-awareness of multimodal large language models in perception,

    Y. Wang, Y. Liao, H. Liu, H. Liu, Y. Wang, and Y. Wang, “Mm- sap: A comprehensive benchmark for assessing self-awareness of multimodal large language models in perception,” arXiv preprint arXiv:2401.07529, 2024

  260. [268]

    Do large language models know what they don’t know?

    Z. Yin, Q. Sun, Q. Guo, J. Wu, X. Qiu, and X. Huang, “Do large language models know what they don’t know?” arXiv preprint arXiv:2305.18153, 2023

  261. [269]

    Survey on factuality in large language models: Knowledge, retrieval and domain-specificity,

    C. Wang, X. Liu, Y. Yue, X. Tang, T. Zhang, C. Jiayang, Y. Yao, W. Gao, X. Hu, Z. Qi et al. , “Survey on factuality in large language models: Knowledge, retrieval and domain-specificity,” arXiv preprint arXiv:2310.07521, 2023

  262. [270]

    Do large language models know how much they know?

    G. Prato, J. Huang, P . Parthasarathi, S. Sodhani, and S. Chandar, “Do large language models know how much they know?” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 6054–6070

  263. [271]

    Introspective planning: Guid- ing language-enabled agents to refine their own uncertainty,

    K. Liang, Z. Zhang, and J. F. Fisac, “Introspective planning: Guid- ing language-enabled agents to refine their own uncertainty,” arXiv preprint arXiv:2402.06529, 2024

  264. [272]

    Exploreself: Fostering user-driven exploration and reflection on personal challenges with adaptive guidance by large language models,

    I. Song, S. Park, S. R. Pendse, J. L. Schleider, M. De Choudhury, and Y.-H. Kim, “Exploreself: Fostering user-driven exploration and reflection on personal challenges with adaptive guidance by large language models,” arXiv preprint arXiv:2409.09662, 2024

  265. [273]

    Self-knowledge guided retrieval augmentation for large language models,

    Y. Wang, P . Li, M. Sun, and Y. Liu, “Self-knowledge guided retrieval augmentation for large language models,” arXiv preprint arXiv:2310.05002, 2023

  266. [274]

    Me, myself, and ai: The situational awareness dataset (sad) for llms,

    R. Laine, B. Chughtai, J. Betley, K. Hariharan, J. Scheurer, M. Balesni, M. Hobbhahn, A. Meinke, and O. Evans, “Me, myself, and ai: The situational awareness dataset (sad) for llms,” arXiv preprint arXiv:2407.04694, 2024

  267. [275]

    Self-recognition in language models,

    T. R. Davidson, V . Surkov, V . Veselovsky, G. Russo, R. West, and C. Gulcehre, “Self-recognition in language models,” arXiv preprint arXiv:2407.06946, 2024

  268. [276]

    Perception of knowledge boundary for large language models through semi-open-ended question answering,

    Z. Wen, Z. Tian, Z. Jian, Z. Huang, P . Ke, Y. Gao, M. Huang, and D. Li, “Perception of knowledge boundary for large language models through semi-open-ended question answering,” arXiv preprint arXiv:2405.14383, 2024

  269. [277]

    Popular large language model chatbots accuracy, compre- hensiveness, and self-awareness in answering ocular symptom queries,

    K. Pushpanathan, Z. W. Lim, S. M. E. Yew, D. Z. Chen, H. A. H. Lin, J. H. L. Goh, W. M. Wong, X. Wang, M. C. J. Tan, V . T. C. Koh et al., “Popular large language model chatbots accuracy, compre- hensiveness, and self-awareness in answering ocular symptom queries,” Iscience, v...

  270. [278]

    Train- ing language models to follow instructions with human feed- back,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Train- ing language models to follow instructions with human feed- back,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022

  271. [279]

    Can ai assistants know what they don’t know?

    Q. Cheng, T. Sun, X. Liu, W. Zhang, Z. Yin, S. Li, L. Li, Z. He, K. Chen, and X. Qiu, “Can ai assistants know what they don’t know?” arXiv preprint arXiv:2401.13275, 2024

  272. [280]

    Bootstrapping cognitive agents with a large language model,

    F. Zhu and R. Simmons, “Bootstrapping cognitive agents with a large language model,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 1, 2024, pp. 655–663

  273. [281]

    Knowl- edge of knowledge: Exploring known-unknowns uncertainty with large language models,

    A. Amayuelas, K. Wong, L. Pan, W. Chen, and W. Wang, “Knowl- edge of knowledge: Exploring known-unknowns uncertainty with large language models,” arXiv preprint arXiv:2305.13712 , 2023

  274. [282]

    Don’t hallucinate, abstain: Identifying llm knowledge gaps via multi-llm collaboration,

    S. Feng, W. Shi, Y. Wang, W. Ding, V . Balachandran, and Y. Tsvetkov, “Don’t hallucinate, abstain: Identifying llm knowledge gaps via multi-llm collaboration,” arXiv preprint arXiv:2402.00367, 2024

  275. [283]

    Seakr: Self-aware knowledge retrieval for adaptive retrieval augmented generation,

    Z. Yao, W. Qi, L. Pan, S. Cao, L. Hu, W. Liu, L. Hou, and J. Li, “Seakr: Self-aware knowledge retrieval for adaptive retrieval augmented generation,” arXiv preprint arXiv:2406.19215, 2024

  276. [284]

    Internal consistency and self- feedback in large language models: A survey,

    X. Liang, S. Song, Z. Zheng, H. Wang, Q. Yu, X. Li, R.-H. Li, Y. Wang, Z. Wang, F. Xiong et al., “Internal consistency and self- feedback in large language models: A survey,” arXiv preprint arXiv:2407.14507, 2024

  277. [285]

    Self-controller: Controlling llms with multi-round step-by-step self-awareness,

    X. Peng and X. Geng, “Self-controller: Controlling llms with multi-round step-by-step self-awareness,” arXiv preprint arXiv:2410.00359, 2024

  278. [286]

    Retrieval- augmented generation for knowledge-intensive nlp tasks,

    P . Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al., “Retrieval- augmented generation for knowledge-intensive nlp tasks,” Ad- vances in Neural Information Processing Systems , vol. 33, pp. 9459– 9474, 2020

  279. [287]

    Active retrieval augmented genera- tion,

    Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig, “Active retrieval augmented genera- tion,” arXiv preprint arXiv:2305.06983, 2023

  280. [288]

    Safe task planning for language-instructed multi-robot systems using conformal predic- tion,

    J. Wang, G. He, and Y. Kantaros, “Safe task planning for language-instructed multi-robot systems using conformal predic- tion,” arXiv preprint arXiv:2402.15368, 2024

  281. [289]

    Autogen: Enabling next-gen llm applications via multi-agent conversation framework,

    Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang, “Autogen: Enabling next-gen llm applications via multi-agent conversation framework,” arXiv preprint arXiv:2308.08155, 2023

  282. [290]

    Instruction- following agents with multimodal transformer,

    H. Liu, L. Lee, K. Lee, and P . Abbeel, “Instruction- following agents with multimodal transformer,” arXiv preprint arXiv:2210.13431, 2022

  283. [291]

    Translating natural language to planning goals with large-language models,

    Y. Xie, C. Yu, T. Zhu, J. Bai, Z. Gong, and H. Soh, “Translating natural language to planning goals with large-language models,” arXiv preprint arXiv:2302.05128, 2023

  284. [292]

    On the planning abilities of large language models-a critical investigation,

    K. Valmeekam, M. Marquez, S. Sreedharan, and S. Kambhampati, “On the planning abilities of large language models-a critical investigation,” Advances in Neural Information Processing Systems , vol. 36, pp. 75 993–76 005, 2023

  285. [293]

    Gg-llm: Geometrically grounding large language models for zero-shot human activity forecasting in human-aware task planning,

    M. A. Graule and V . Isler, “Gg-llm: Geometrically grounding large language models for zero-shot human activity forecasting in human-aware task planning,” in2024 IEEE International Confer- ence on Robotics and Automation (ICRA). IEEE, 2024, pp. 568–574

  286. [294]

    Generalized planning in pddl domains with pretrained large language models,

    T. Silver, S. Dan, K. Srinivas, J. B. Tenenbaum, L. Kaelbling, and M. Katz, “Generalized planning in pddl domains with pretrained large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 18, 2024, pp. 20 256–20 264

  287. [295]

    Pddl planning with pretrained large language models,

    T. Silver, V . Hariprasad, R. S. Shuttleworth, N. Kumar, T. Lozano- Pérez, and L. P . Kaelbling, “Pddl planning with pretrained large language models,” in NeurIPS 2022 foundation models for decision making workshop, 2022

  288. [296]

    Fine-tuning large vision-language models as decision-making agents via reinforcement learning,

    Y. Zhai, H. Bai, Z. Lin, J. Pan, S. Tong, Y. Zhou, A. Suhr, S. Xie, Y. LeCun, Y. Maet al., “Fine-tuning large vision-language models as decision-making agents via reinforcement learning,” arXiv preprint arXiv:2405.10292, 2024

  289. [297]

    Planning with large language models via corrective re-prompting,

    S. S. Raman, V . Cohen, E. Rosen, I. Idrees, D. Paulius, and S. Tellex, “Planning with large language models via corrective re-prompting,” in NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022

  290. [298]

    Large language models are learnable planners for long- term recommendation,

    W. Shi, X. He, Y. Zhang, C. Gao, X. Li, J. Zhang, Q. Wang, and F. Feng, “Large language models are learnable planners for long- term recommendation,” in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024, pp. 1893–1903

  291. [299]

    Leveraging pre-trained large language models to construct and utilize world models for model-based task planning,

    L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Leveraging pre-trained large language models to construct and utilize world models for model-based task planning,” Advances in Neural Information Processing Systems, vol. 36, pp. 79 081–79 094, 2023. 34

  292. [300]

    Leveraging environ- ment interaction for automated pddl translation and planning with large language models,

    S. Mahdavi, R. Aoki, K. Tang, and Y. Cao, “Leveraging environ- ment interaction for automated pddl translation and planning with large language models,” arXiv preprint arXiv:2407.12979 , 2024

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.