Pith. sign in

REVIEW 3 major objections 4 minor 78 references

To explain, validate, or govern an AI agent's behavior, first find out which layer produced it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 18:50 UTC pith:F2TATNUN

load-bearing objection A clear conceptual synthesis of layer attribution for AI agents, but the stability-based diagnostic rule is underdetermined and can misclassify constant modulation factors as foundational. the 3 major comments →

arxiv 2607.17149 v1 pith:F2TATNUN submitted 2026-07-19 cs.AI cs.CY

A Diagnostic Framework for AI Agent Behavior

classification cs.AI cs.CY
keywords AI agent behaviorlayer attributionbehavioral sciencesurrogate validityhuman-AI divergenceAI governance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that a behavioral pattern shown by an AI agent can originate in either of two layers: the foundational computational layer (architecture, memory, perception, representation) that defines what behavior is possible, or the behavioral modulation layer (roles, objectives, interaction structures, governance) that shapes how that capacity is expressed. The proposed diagnostic move, layer attribution, asks which layer generated a pattern before deciding how to explain, validate, or govern it. The paper's operational rule is that patterns persisting across changes in prompts, roles, metrics, and deployment constraints point to foundational sources, while patterns that shift point to modulation. This matters because the same observable behavior can demand very different interventions depending on its source: output guardrails cannot fix a distortion embedded in the substrate. If the framework is right, surrogate validity becomes a relation between model, task, and layer; human-AI divergence becomes diagnostic evidence; and governance must attribute source before intervening.

Core claim

The central claim is that AI agent behavior should be diagnosed by source before it is explained or acted upon. The paper defines AI agent behavior as context-sensitive, goal-directed action that unfolds over time and affects human systems, and it identifies two generating layers. The foundational computational layer — architecture, memory, perception and attention, and representational structure — sets what the agent can do; the behavioral modulation layer — identity, resources, objectives, social interaction, institutional constraints, and governance — shapes how those capacities appear in context. The diagnostic rule is simple: patterns that stay stable as prompts, roles, metrics, and dep

What carries the argument

The central object is the two-layer diagnostic framework the paper calls layer attribution. It works by treating behavioral science as a selective diagnostic resource: cognitive science distinguishes representation, memory, perception, and attention; social psychology and sociology distinguish role, norm, interaction, and institutional effects; economics and governance scholarship specify how objectives and constraints shape action. The framework then maps these onto AI agents and prescribes a perturbation method — compare behavior across changes in prompts, roles, objectives, interaction structures, and governance constraints — and infer the source from stability or shift. The definition of

Load-bearing premise

The paper's diagnostic rule assumes that behavioral stability across changes in roles, prompts, metrics, and deployment constraints cleanly separates foundational from modulation sources; but a foundational limitation may show up only in specific contexts, and a modulation effect can be very stable, so the mapping is a heuristic rather than a proven diagnostic.

What would settle it

A controlled experiment where a behavioral bias (for example, the compression of minority opinions in a social simulation) persists across all tested prompt, role, and metric variations, yet is then eliminated by a purely modulation-layer change such as rewriting the reward function or adding a governance constraint, would falsify the claim that persistence indicates a foundational source.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Surrogate validity is not a fixed property of a model; the same agent can be a valid stand-in for one behavioral question and an invalid one for another, depending on which layer the question requires to align.
  • Human-AI divergence should be read differentially: a gap that persists across all tested variations points to a foundational difference, while one that shifts with roles or objectives points to modulation, so the same gap can serve different scientific purposes.
  • Governance tools that operate on outputs, incentives, and deployment constraints (such as RLHF, constitutional AI, and guardrails) may be ineffective when the risk originates in the foundational layer; source attribution should precede intervention choice.
  • The framework prescribes a concrete evaluation practice: stress tests that vary roles, goals, memory, interaction partners, and governance constraints, then compare the resulting behavioral patterns to attribute their source.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the stability heuristic holds, evaluation of AI agents should routinely include perturbation batteries; but the boundary between layers is likely fuzzy, since changing a role may also change available context or objectives, so attribution may require disentangling correlated modulation factors.
  • A testable extension: measure whether a persistent behavioral bias (e.g., compressed minority views) can be removed by a purely modulation-layer intervention; if it can, the stability rule would need refinement.
  • The framework implies that studies using LLMs as human surrogates should report which layer they intend to align; otherwise their conclusions are underspecified about the scope of validity.
  • It also suggests that alignment success at the output surface can mask foundational risk, so governance evaluations should consider probing internal representations, not just final responses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a conceptual framework, 'layer attribution', for diagnosing whether an AI agent's observed behavior originates in a foundational computational layer (architecture, memory, perception/attention, representation) or in a behavioral modulation layer (identity/role, resources, objectives, interaction structure, governance). The central diagnostic rule (Section 4, first paragraph) is that behavioral patterns persisting across changes in roles, prompts, metrics, and deployment constraints indicate foundational sources, while patterns shifting with objectives, roles, interaction structures, or governance rules indicate modulation-layer sources. The framework is then applied to three areas: surrogate validity is characterized as a model-task-layer relation, human-AI divergence is reinterpreted as diagnostic evidence, and governance is said to require source attribution before intervention. The paper is explicitly a Perspective and does not present empirical validation; its conclusion states that making the framework operational is the 'next agenda'.

Significance. If the layer-attribution distinction can be made operational, the framework would give behavioral scientists and AI governance researchers a useful common vocabulary for asking where an agent behavior originates, and it connects machine-behavior research to existing behavioral science constructs. The paper's three consequences are plausible, clearly stated, and broadly grounded in the literature; they are not derived from fitted parameters, so there is no circularity of the prediction-from-fit type. The manuscript is honest that operationalization is future work, which is appropriate for a Perspective. Its main contribution is synthetic and conceptual rather than empirical; the value depends on whether the diagnostic rule can be sharpened into a reliable procedure.

major comments (3)
  1. [Section 4, first paragraph; Table 2] The stability rule is underdetermined as stated. It classifies patterns that persist across roles, prompts, metrics, and deployment constraints as foundational, but it does not require that modulation-layer variables be varied. A fixed modulation factor, such as a constant RLHF/constitutional-AI safety objective — which Section 4.3 itself places in the modulation layer — will produce stable behavior across the listed perturbations and would therefore be misdiagnosed as foundational. Conversely, foundational limitations such as weak grounding or compressed minority representation may appear only in novel or demanding settings, so a context-sensitive shift is not uniquely diagnostic of modulation. The paper should either restrict the rule to contrasts that vary both layer families, or explicitly present stability as a conditional evidence with stated confounds and a calibration requirement
  2. [Section 3.1, Table 1 vs Section 4.3] There is an internal tension in layer assignment. Table 1 lists 'training regularities' as part of the foundational substrate under 'Weighted architectures; training regularities', while Section 4.3 assigns RLHF and constitutional AI to the behavioral modulation layer. RLHF and constitutional AI alter weights and training regularities, so the same intervention can be attributed to both layers depending on whether one emphasizes mechanism or policy intent. This ambiguity matters for the governance consequence: the framework's recommendation to choose intervention layer requires a clear criterion for when a weight-level change counts as foundational repair rather than modulation. The manuscript should disambiguate source of behavior from locus of intervention.
  3. [Section 5; Section 4, Table 2] The manuscript explicitly states in Section 5 that making the framework operational is the 'next agenda', which is an honest limitation but also a load-bearing one for the title's claim to provide a 'Diagnostic Framework'. Table 2 presents three consequences as established, but they rest on the unvalidated stability heuristic. For a Perspective, a full empirical demonstration is not required, but the paper should either provide a worked example showing how the diagnostic rule distinguishes a foundational from a modulation source in a concrete setting, or state clearly, in the abstract and introduction, that the framework is a hypothesis-generating heuristic rather than a validated diagnostic. This would align the presentation with the admitted scope and avoid overclaiming.
minor comments (4)
  1. [References] Reference [22] contains typos in the title: 'nNnetworks and lLearning' should be 'Neural Networks and Learning Systems'.
  2. [References] Reference [73] uses 'mises-timation' in the title; this appears to be a misspelling of 'misestimation' and should be corrected.
  3. [Section 2] The definition of 'attributable' as 'traceable, in principle, to the underlying mechanisms' is not operationalized; it would help to state what kind of evidence would count as tracing a behavior to a layer.
  4. [Throughout] The terms 'model behavior' and 'agent behavior' are used side by side (e.g., Abstract, Section 2); the relationship between a single-query response and an agent behavior is defined, but the manuscript should ensure that 'model behavior' is either defined or replaced consistently.

Circularity Check

0 steps flagged

No significant circularity: the framework is a definitional proposal with no fitted parameters, equations, or load-bearing self-citations.

full rationale

The paper is a Perspective that proposes a conceptual two-layer framework (foundational computational layer vs. behavioral modulation layer) and draws three consequences (surrogate validity, human-AI divergence, governance). There are no equations, no fitted parameters, and no quantitative predictions derived from data. The diagnostic rule—stable patterns suggest foundational sources, shifting patterns suggest modulation—is presented as a heuristic, not as a theorem derived from the definitions, so it is not circular in the sense of reducing to its inputs by construction. The one self-citation ([41], Zhang & Cavusoglu) appears alongside multiple independent references and supports an ancillary empirical claim about LLM reasoning; it is not load-bearing and does not introduce circularity. The paper's central claims follow from the proposed definitions, but that is the normal structure of a conceptual framework, not a circular reduction. The main potential weakness is the empirical validity of the stability heuristic under conditions where modulation factors are held constant, but that is a correctness/robustness concern, not circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The framework rests on a few key assumptions: that AI behavior is attributable, that a two-layer partition captures the relevant sources, and that stability across perturbations indicates the foundational layer. No free parameters are fitted; no new entities are postulated.

axioms (4)
  • domain assumption AI agent behavior is attributable in principle to underlying mechanisms
    Section 2 lists 'attributable' as a defining property of AI agent behavior; the entire diagnostic approach presumes this traceability.
  • domain assumption The two-layer partition is sufficient to classify sources of behavior
    Section 3 decomposes all sources into foundational vs behavioral modulation; interactions are mentioned but not structurally incorporated.
  • ad hoc to paper Stability across perturbations diagnoses foundational source
    Section 4 states 'patterns that persist across roles, prompts, metrics, or deployment constraints suggest foundational sources'—this is the key operational inference, asserted without proof.
  • domain assumption Functional analogy between human cognitive/social constructs and AI components
    Section 3 maps memory, attention, role, governance to AI counterparts; the whole comparison assumes this translation is valid.

pith-pipeline@v1.3.0-alltime-deepseek · 9838 in / 8603 out tokens · 75505 ms · 2026-08-01T18:50:16.727160+00:00 · methodology

0 comments
read the original abstract

AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 7 linked inside Pith

  1. [1]

    Nature, 1–8 (2025)

    Lin, H., Czarnek, G., Lewis, B., White, J.P., Berinsky, A.J., Costello, T., Pen- nycook, G., Rand, D.G.: Persuading voters using human–artificial intelligence dialogues. Nature, 1–8 (2025)

  2. [2]

    Science385, 1814 (2024)

    Costello, T.H., Pennycook, G., Rand, D.G.: Durably reducing conspiracy beliefs through dialogues with AI. Science385, 1814 (2024)

  3. [3]

    npj Digital Medicine8(1), 640 (2025)

    Croxford, E., Gao, Y., First, E., Pellegrino, N., Schnier, M., Caskey, J., Oguss, M., Wills, G., Chen, G., Dligach, D.,et al.: Evaluating clinical AI summaries with large language models as judges. npj Digital Medicine8(1), 640 (2025)

  4. [4]

    Nature Communications 16(1), 9076 (2025)

    Wang, S., Zhao, F., Bu, D., Lu, Y., Gong, M., Liu, H., Yang, Z., Zeng, X., Yuan, Z., Wan, B.,et al.: Lins: A general medical q&a framework for enhancing the quality and credibility of LLM-generated responses. Nature Communications 16(1), 9076 (2025)

  5. [5]

    npj Digital Medicine8(1), 119 (2025)

    Shusterman, R., Waters, A.C., O’Neill, S., Bangs, M., Luu, P., Tucker, D.M.: An active inference strategy for prompting reliable responses from large language models in medical practice. npj Digital Medicine8(1), 119 (2025)

  6. [6]

    Proceedings of the National Academy of Sciences122(24), 2501660122 (2025)

    Gao, Y., Lee, D., Burtch, G., Fazelpour, S.: Take caution in using LLMs as human surrogates. Proceedings of the National Academy of Sciences122(24), 2501660122 (2025)

  7. [7]

    In: Findings of the Association for Computational Linguistics: NAACL 2024, pp

    Chuang, Y.-S., Goyal, A., Harlalka, N., Suresh, S., Hawkins, R., Yang, S., Shah, D., Hu, J., Rogers, T.: Simulating opinion dynamics with networks of LLM-based agents. In: Findings of the Association for Computational Linguistics: NAACL 2024, pp. 3326–3346 (2024)

  8. [8]

    Advances in Neural Information Processing Systems36, 51991–52008 (2023)

    Li, G., Hammoud, H., Itani, H., Khizbullin, D., Ghanem, B.: Camel: Commu- nicative agents for “mind” exploration of large language model society. Advances in Neural Information Processing Systems36, 51991–52008 (2023)

  9. [9]

    In: The 12th International Conference on Learning Representations, pp

    Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S.K.S., Lin, Z.,et al.: MetaGPT: Meta programming for a multi- agent collaborative framework. In: The 12th International Conference on Learning Representations, pp. 23247–23275 (2023)

  10. [10]

    In: First Conference on Language Modeling (2024)

    Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J.,et al.: Autogen: Enabling next-gen LLM applications via multi-agent conversations. In: First Conference on Language Modeling (2024)

  11. [11]

    Nature620(7972), 47–60 (2023) 11

    Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Van Katwyk, P., Deac, A.,et al.: Scientific discovery in the age of artificial intelligence. Nature620(7972), 47–60 (2023) 11

  12. [12]

    Science Advances11(20), 9368 (2025)

    Ashery, A.F., Aiello, L.M., Baronchelli, A.: Emergent social conventions and collective bias in LLM populations. Science Advances11(20), 9368 (2025)

  13. [13]

    Murthy, S.K., Ullman, T., Hu, J.: One fish, two fish, but not the whole sea: Align- ment reduces language models’ conceptual diversity. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Com- putational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 11241–11258 (2025)

  14. [14]

    Nature568(7753), 477–486 (2019)

    Rahwan, I., Cebrian, M., Obradovich, N., Bongard, J., Bonnefon, J.-F., Breazeal, C., Crandall, J.W., Christakis, N.A., Couzin, I.D., Jackson, M.O.,et al.: Machine behaviour. Nature568(7753), 477–486 (2019)

  15. [15]

    Nature, 1–9 (2025)

    Tu, T., Schaekermann, M., Palepu, A., Saab, K., Freyberg, J., Tanno, R., Wang, A., Li, B., Amin, M., Cheng, Y., et al.: Towards conversational diagnostic artificial intelligence. Nature, 1–9 (2025)

  16. [16]

    npj Heritage Science13(1), 608 (2025)

    Hao, X., Xu, J., Wang, Y.: How generative AI shapes user perceived value and adoption intention in digital museum experiences. npj Heritage Science13(1), 608 (2025)

  17. [17]

    N´ u˜ nez, R., Allen, M., Gao, R., Miller Rigoli, C., Relaford-Doyle, J., Semenuks, A.: What happened to cognitive science? Nature Human Behaviour3(8), 782–791 (2019)

  18. [18]

    Nature Human Behaviour8(10), 1864–1876 (2024)

    Tsvetkova, M., Yasseri, T., Pescetelli, N., Werner, T.: A new sociology of humans and machines. Nature Human Behaviour8(10), 1864–1876 (2024)

  19. [19]

    Nature521(7553), 436–444 (2015)

    LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature521(7553), 436–444 (2015)

  20. [20]

    Nature601(7894), 549–555 (2022)

    Wright, L.G., Onodera, T., Stein, M.M., Wang, T., Schachter, D.T., Hu, Z., McMahon, P.L.: Deep physical neural networks trained with backpropagation. Nature601(7894), 549–555 (2022)

  21. [21]

    arXiv:2512.13564 (2025)

    Hu, Y., Liu, S., Yue, Y., Zhang, G., Liu, B., Zhu, F., Lin, J., Guo, H., Dou, S., Xi, Z., et al.: Memory in the age of AI agents. arXiv:2512.13564 (2025)

  22. [22]

    IEEE Transactions on Neural nNetworks and lLearning Systems30(11), 3212–3232 (2019)

    Zhao, Z.-Q., Zheng, P., Xu, S.-t., Wu, X.: Object detection with deep learning: A review. IEEE Transactions on Neural nNetworks and lLearning Systems30(11), 3212–3232 (2019)

  23. [23]

    Neurocomputing452, 48–62 (2021)

    Niu, Z., Zhong, G., Yu, H.: A review on the attention mechanism of deep learning. Neurocomputing452, 48–62 (2021)

  24. [24]

    Advances in Neural 12 Information Processing Systems30(2017)

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. Advances in Neural 12 Information Processing Systems30(2017)

  25. [25]

    Proceedings of the National Academy of Sciences122(44), 2512514122 (2025)

    Xu, N., Zhang, Q., Du, C., Luo, Q., Qiu, X., Huang, X., Zhang, M.: Reveal- ing emergent human-like conceptual representations from language prediction. Proceedings of the National Academy of Sciences122(44), 2512514122 (2025)

  26. [26]

    Nature638(8051), 769–778 (2025)

    Xiang, J., Wang, X., Zhang, X., Xi, Y., Eweje, F., Chen, Y., Li, Y., Bergstrom, C., Gopaulchan, M., Kim, T.,et al.: A vision–language foundation model for precision oncology. Nature638(8051), 769–778 (2025)

  27. [27]

    Nature Machine Intelligence7(1), 96–106 (2025)

    Schulze Buschoff, L.M., Akata, E., Bethge, M., Schulz, E.: Visual cognition in multimodal large language models. Nature Machine Intelligence7(1), 96–106 (2025)

  28. [28]

    MIT press, Cambridge (1986)

    Rumelhart, D.E., McClelland, J.L., Group, P.R.,et al.: Parallel Distributed Pro- cessing, Volume 1: Explorations in the Microstructure of Cognition: Foundations. MIT press, Cambridge (1986)

  29. [29]

    Psychology Press, London (2005)

    Hebb, D.O.: The Organization of Behavior: A Neuropsychological Theory. Psychology Press, London (2005)

  30. [30]

    Psychology of Learning and Motivation2, 89–195 (1968)

    Atkinson, R.C., Shiffrin, R.M.: Human memory: A proposed system and its control processes. Psychology of Learning and Motivation2, 89–195 (1968)

  31. [31]

    Psychological Review63(2), 81 (1956)

    Miller, G.A.: The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review63(2), 81 (1956)

  32. [32]

    Elsevier, New York (2013)

    Broadbent, D.E.: Perception and Communication. Elsevier, New York (2013)

  33. [33]

    Cognitive Psychology12(1), 97–136 (1980)

    Treisman, A.M., Gelade, G.: A feature-integration theory of attention. Cognitive Psychology12(1), 97–136 (1980)

  34. [34]

    Psychology Press, London (2013)

    Paivio, A.: Imagery and Verbal Processes. Psychology Press, London (2013)

  35. [35]

    Behavioral and Brain Sciences22(4), 577–660 (1999)

    Barsalou, L.W.: Perceptual symbol systems. Behavioral and Brain Sciences22(4), 577–660 (1999)

  36. [36]

    Artificial Intelligence Review 59(1), 11 (2025)

    Abou Ali, M., Dornaika, F., Charafeddine, J.: Agentic AI: a comprehensive survey of architectures, applications, and future directions. Artificial Intelligence Review 59(1), 11 (2025)

  37. [37]

    arXiv:2602.12285 (2026)

    Cao, L., Sun, L., Yue, Y.: From biased chatbots to biased agents: Examining role assignment effects on LLM agent robustness. arXiv:2602.12285 (2026)

  38. [38]

    In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

    Li, J., Liu, X., Feng, Y.: From single to societal: Analyzing persona-induced bias in multi-agent interactions. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, pp. 31609–31617 (2026) 13

  39. [39]

    Humanities and Social Sciences Communications11(1), 1–24 (2024)

    Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., Li, Y.: Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanities and Social Sciences Communications11(1), 1–24 (2024)

  40. [40]

    Transactions of the Association for Computational Linguistics12, 157–173 (2024)

    Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics12, 157–173 (2024)

  41. [41]

    Available at SSRN 5337019 (2025)

    Zhang, X., Cavusoglu, H.: Unpacking the decision logic of LLM-based agents: Evidence from the newsvendor problem. Available at SSRN 5337019 (2025)

  42. [42]

    Data Science and Engineering, 1–31 (2025)

    Xu, W., Huang, C., Gao, S., Shang, S.: LLM-based agents for tool learning: A survey. Data Science and Engineering, 1–31 (2025)

  43. [43]

    ACM Computing Surveys (2025)

    Du, S., Zhao, J., Shi, J., Xie, Z., Jiang, X., Bai, Y., He, L.: A survey on the optimization of large language model-based agents. ACM Computing Surveys (2025)

  44. [44]

    Advances in Neural Information Processing Systems37, 5244–5284 (2024)

    Feng, P., He, Y., Huang, G., Lin, Y., Zhang, H., Zhang, Y., Li, H.: Agile: A novel reinforcement learning framework of LLM agents. Advances in Neural Information Processing Systems37, 5244–5284 (2024)

  45. [45]

    arXiv:2502.10325 (2025)

    Choudhury, S.: Process reward models for LLM agents: Practical framework and directions. arXiv:2502.10325 (2025)

  46. [46]

    In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

    Mordatch, I., Abbeel, P.: Emergence of grounded compositional language in multi-agent populations. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018)

  47. [47]

    Advances in Neural Information Processing Systems37, 83548–83599 (2024)

    Abdelnabi, S., Gomaa, A., Sivaprasad, S., Sch¨ onherr, L., Fritz, M.: Coopera- tion, competition, and maliciousness: LLM-stakeholders interactive negotiation. Advances in Neural Information Processing Systems37, 83548–83599 (2024)

  48. [48]

    Advances in Neural Information Processing Systems36, 8634–8652 (2023)

    Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., Yao, S.: Reflexion: Lan- guage agents with verbal reinforcement learning. Advances in Neural Information Processing Systems36, 8634–8652 (2023)

  49. [49]

    arXiv:2405.06682 (2024)

    Renze, M., Guven, E.: Self-reflection in LLM agents: Effects on problem-solving performance. arXiv:2405.06682 (2024)

  50. [50]

    In: Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, pp

    Liu, Z., Bai, X., Chen, K., Chen, X., Li, X., Xiang, Y., Liu, J., Li, H.-D., Wang, Y., Nie, L.,et al.: A survey on the feedback mechanism of LLM-based AI agents. In: Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, pp. 10582–10592 (2025)

  51. [51]

    Advances in Neural Information Processing Systems37, 15497–15525 (2024)

    Ma, H., Hu, T., Pu, Z., Boyin, L., Ai, X., Liang, Y., Chen, M.: Coevolving with the other you: Fine-tuning LLM with sequential cooperative multi-agent 14 reinforcement learning. Advances in Neural Information Processing Systems37, 15497–15525 (2024)

  52. [52]

    In: ICML 2025 Workshop on Computer Use Agents (2025)

    Xiang, Z., Zheng, L., Li, Y., Hong, J., Li, Q., Xie, H., Zhang, J., Xiong, Z., Xie, C., Bastian, N.D.,et al.: GuardAgent: Safeguard LLM agents via knowledge-enabled reasoning. In: ICML 2025 Workshop on Computer Use Agents (2025)

  53. [53]

    Advances in Neural Information Processing Systems37, 52040–52094 (2024)

    Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T.J., Cheng, Z., Shin, D., Lei, F.,et al.: Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems37, 52040–52094 (2024)

  54. [54]

    arXiv:2601.18491 (2026)

    Liu, D., Ren, Q., Qian, C., Shao, S., Xie, Y., Li, Y., Yang, Z., Luo, H., Wang, P., Liu, Q., et al.: AgentDog: A diagnostic guardrail framework for AI agent safety and security. arXiv:2601.18491 (2026)

  55. [55]

    IEEE Transactions on Network Science and Engineering (2026)

    Qi, M., Zhu, T., Zhang, L., Li, N., Tan, Y.-a., Zhou, W.: Towards transparent and incentive-compatible collaboration in decentralized LLM multi-agent sys- tems: A blockchain-driven approach. IEEE Transactions on Network Science and Engineering (2026)

  56. [56]

    IEEE Access13, 18912–18936 (2025)

    Acharya, D.B., Kuppan, K., Divya, B.: Agentic AI: Autonomous intelligence for complex goals: A comprehensive survey. IEEE Access13, 18912–18936 (2025)

  57. [57]

    arXiv:2506.04133 (2025)

    Raza, S., Sapkota, R., Karkee, M., Emmanouilidis, C.: Trism for agentic AI: A review of trust, risk, and security management in LLM-based agentic multi-agent systems. arXiv:2506.04133 (2025)

  58. [58]

    Research Paper, OpenAI (2023)

    Shavit, Y., Agarwal, S., Brundage, M., Adler, S., O’Keefe, C., Campbell, R., Lee, T., Mishkin, P., Eloundou, T., Hickey, A., et al.: Practices for governing agentic AI systems. Research Paper, OpenAI (2023)

  59. [59]

    IEEE Open Journal of Intelligent Transportation Systems7, 615–657 (2026)

    Ferrag, M.A., Lakas, A., Tihanyi, N., Debbah, M.: LLM and AI agents for autonomous systems: A survey of applications, datasets, and security challenges. IEEE Open Journal of Intelligent Transportation Systems7, 615–657 (2026)

  60. [60]

    IEEE Intelligent Systems40(2), 8–14 (2025)

    Murugesan, S.: The rise of agentic AI: implications, concerns, and the path forward. IEEE Intelligent Systems40(2), 8–14 (2025)

  61. [61]

    Nature, 1–8 (2025)

    Binz, M., Akata, E., Bethge, M., Br¨ andle, F., Callaway, F., Coda-Forno, J., Dayan, P., Demircan, C., Eckstein, M.K., ´Eltet˝ o, N., et al.: A foundation model to predict and capture human cognition. Nature, 1–8 (2025)

  62. [62]

    Proceed- ings of the National Academy of Sciences120(6), 2218523120 (2023)

    Binz, M., Schulz, E.: Using cognitive psychology to understand GPT-3. Proceed- ings of the National Academy of Sciences120(6), 2218523120 (2023)

  63. [63]

    Nature Communications16(1), 8633 (2025)

    Webb, T., Mondal, S.S., Momennejad, I.: A brain-inspired agentic architecture 15 to improve planning with LLMs. Nature Communications16(1), 8633 (2025)

  64. [64]

    Chen, Y., Kirshner, S.N., Ovchinnikov, A., Andiappan, M., Jenkin, T.: A manager and an AI walk into a bar: does ChatGPT make biased decisions like we do? Manufacturing & Service Operations Management27(2), 354–368 (2025)

  65. [65]

    In: Findings of the Association for Computational Linguistics: EMNLP 2024, pp

    Echterhoff, J.M., Liu, Y., Alessa, A., McAuley, J., He, Z.: Cognitive bias in decision-making with LLMs. In: Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 12640–12653 (2024)

  66. [66]

    In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

    Zhuang, N., Cao, B., Yang, Y., Xu, J., Xu, M., Wang, Y., Liu, Q.: LLM agents can be choice-supportive biased evaluators: An empirical study. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 26436–26444 (2025)

  67. [67]

    Proceedings of the National Academy of Sciences120(49), 2309350120 (2023)

    Le Mens, G., Kov´ acs, B., Hannan, M.T., Pros, G.: Uncovering the semantics of concepts using GPT-4. Proceedings of the National Academy of Sciences120(49), 2309350120 (2023)

  68. [68]

    arXiv:2210.03629 (2022)

    Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., Cao, Y.: ReAct: Synergizing reasoning and acting in language models. arXiv:2210.03629 (2022)

  69. [69]

    14–23 (2025)

    Ng, L.H.X., Carley, K.M.: Are LLM-powered social media bots realistic? In: Inter- national Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, pp. 14–23 (2025). Springer

  70. [70]

    Nature646(8085), 716–723 (2025)

    Swanson, K., Wu, W., Bulaong, N.L., Pak, J.E., Zou, J.: The virtual lab of AI agents designs new SARS-CoV-2 nanobodies. Nature646(8085), 716–723 (2025)

  71. [71]

    arXiv:2410.10934 (2024)

    Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., et al.: Agent-as-a-Judge: Evaluate agents with agents. arXiv:2410.10934 (2024)

  72. [72]

    Nature Communications16(1), 9104 (2025)

    Mandal, I., Soni, J., Zaki, M., Smedskjaer, M.M., Wondraczek, K., Wondraczek, L., Gosvami, N.N., Krishnan, N.A.: Evaluating large language model agents for automation of atomic force microscopy. Nature Communications16(1), 9104 (2025)

  73. [73]

    Proceedings of the National Academy of Sciences122(48), 2519394122 (2025)

    Pataranutaporn, P., Powdthavee, N., Archiwaranguprok, C., Maes, P.: Simulating human well-being with large language models: Systematic validation and mises- timation across 64,000 individuals from 64 countries. Proceedings of the National Academy of Sciences122(48), 2519394122 (2025)

  74. [74]

    Nature Machine Intelligence5(10), 1076–1086 (2023) 16

    Pataranutaporn, P., Liu, R., Finn, E., Maes, P.: Influencing human-AI interaction by priming beliefs about AI can increase perceived trustworthiness, empathy and effectiveness. Nature Machine Intelligence5(10), 1076–1086 (2023) 16

  75. [75]

    In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol

    Batzner, J., Stocker, V., Tang, B., Natarajan, A., Chen, Q., Schmid, S., Kas- neci, G.: Whose personae? Synthetic persona experiments in LLM research and pathways to transparency. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 8, pp. 343–354 (2025)

  76. [76]

    Advances in Neural Information Processing Systems30(2017)

    Christiano, P.F., Leike, J., Brown, T., Martic, M., Legg, S., Amodei, D.: Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems30(2017)

  77. [77]

    arXiv:2212.08073 (2022)

    Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al.: Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073 (2022)

  78. [78]

    In: Proceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp

    Luo, W., Dai, S., Liu, X., Banerjee, S., Sun, H., Chen, M., Xiao, C.: Agrail: A lifelong agent guardrail with effective and adaptive safety detection. In: Proceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8104–8139 (2025) 17