Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

AI Agent Behavioral Science

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read AI agent behavior is not determined by the model alone but emerges from situated interaction, making it a scientific subject in its own right.

desk verdict A useful, occasionally sloppy consolidation of the agent-behavior literature, whose central promise—measuring agents as behaviors rather than mechanisms—is real but not yet earned. read the letter →

arxiv 2506.06366 v3 pith:I35BS4PP submitted 2025-06-04 q-bio.NC cs.CYcs.MA

classification q-bio.NCcs.CYcs.MA
keywords AIagentsbehavioralsciencelargelanguagemodelsmulti-agentsystemshuman-agentinteractionresponsibleFoggBehaviorModelmachine
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the behavior of an AI agent is not fixed by the model inside it; it emerges from the agent's situated interaction with environments, other agents, and humans. It proposes AI Agent Behavioral Science as a new paradigm that studies what agents actually do, adapt, and become over time, treating the model as a substrate that enables but does not determine behavior. The paper systematizes findings across individual, multi-agent, and human-agent settings, and shows how responsible AI concerns such as fairness, safety, and privacy can be reframed as measurable behavioral properties. A sympathetic reader would care because, if the paradigm holds, evaluating and governing AI shifts from inspecting internal weights to observing and shaping behavioral trajectories.

What carries the argument

The paradigm itself is the central object: AI Agent Behavioral Science, defined as 'the study of how AI agents act, adapt, and interact in situated contexts.' Carrying the argument are three organizing devices: the brain-to-action analogy, which licenses transferring behavioral-science methods to agents; the social cognitive theory triad of intrinsic attributes, environmental constraints, and behavioral feedback for individual behavior; and the Fogg Behavior Model (ability, motivation, trigger) for classifying adaptation techniques. These devices convert scattered empirical results into a structured scientific field with its own measurement, intervention, and theory-guided interpretation.

What would settle it

A concrete observation that would settle the claim: if a fixed model and fixed task, with only incidental prompt phrasing or model version varied, produces behavior distributions that vary as much as, or more than, changes in the environmental and social variables the paradigm treats as formative, the situated-interaction foundation is undermined. Alternatively, demonstrating that agent behaviors observed in sandbox simulations do not predict behaviors of the same agents in deployment settings would falsify the paradigm's predictive promise.

Watch

Extended reading notes

Core claim

The central claim is that AI agent behavior is a legitimate empirical subject in its own right. The paper states this as an ontological analogy: 'the model is to behavior what the brain is to action: a substrate that enables but does not determine.' From this, it follows that complex behaviors such as negotiation, deception, cooperation, and institutional formation are not properties of the LLM alone but products of the agentic system—memory, planning, tools, roles, feedback—embedded in context. The paper organizes the emerging literature into individual, multi-agent, and human-agent interaction layers, and interprets adaptation methods through the Fogg Behavior Model by mapping ability to pretraining, motivation to reward signals, and trigger to prompting. It then argues that responsible AI principles should be treated as dynamic, context-dependent behavioral attributes rather than static model properties, opening a research agenda on behavioral entropy, macro-level adaptation, artificial societies, and behavioral warning signs.

Load-bearing premise

The paradigm assumes that human behavioral-science methods transfer to LLM-based agents, meaning that agent behavior has stable, context-dependent, causally structured regularities that hold beyond the specific sandbox where they were observed.

Editorial extensions

If this is right

  • If the paradigm is right, evaluating an AI system means running behavioral experiments—observing trajectories over time and across contexts—rather than only auditing weights or static outputs.
  • Responsible AI metrics (fairness, safety, interpretability, accountability, privacy) would each need trajectory-level operationalizations, since one-shot assessments miss drift, deception, and feedback effects.
  • Adaptation research would be organized around behavioral levers—ability, motivation, trigger—so that prompting, fine-tuning, and reinforcement learning are seen as complementary ways to shape behavior, not competing paradigms.
  • Artificial societies made of agents become legitimate instruments for behavioral theory, allowing controlled, replicable, counterfactual experiments that are impossible with human subjects.
  • Hybrid human-agent teams and machine culture become objects of scientific study, with the field predicting that behavior, not architecture, determines collective outcomes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The logic of the paper implies that behavioral benchmarks should eventually replace or complement static model benchmarks for deployment decisions; a model card would be paired with a behavioral profile measured across contexts. This is our inference, not a claim the paper explicitly makes.
  • If behavioral science transfers to agents, then behavioral entropy and other summary statistics could function like temperament inventories for AI, enabling comparison across model versions and providers—an extension the paper proposes as a research direction, not an established result.
  • A testable consequence the authors leave implicit: two agents with the same underlying model but different memory, role, or feedback structures should diverge behaviorally over interaction rounds, and this divergence should be measurable and stable across repeated runs. A failure to observe such divergence would challenge the substrate-enables-behavior claim.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes a new research paradigm called "AI Agent Behavioral Science," arguing that LLM-based agents should be studied not only through internal model mechanisms but as behavioral entities whose actions emerge from situated interaction with environments, other agents, and humans. It synthesizes recent empirical work across three settings—individual agents, multi-agent systems, and human-agent interaction—organizes behavioral adaptation methods using a reinterpretation of the Fogg Behavior Model (ability, motivation, trigger), and reframes responsible AI principles (fairness, safety, interpretability, accountability, privacy) as behavioral properties. The paper closes with six proposed research directions. It is a synthesis and position paper rather than a new experimental study.

Significance. If its central premise is accepted, the paper makes a timely contribution by connecting machine behavior, behavioral economics, social simulation, and AI governance under one framework. Its strengths include broad literature coverage, a useful taxonomy of emergent agent behaviors, and a concrete agenda for studying behavioral reliability and adaptation. However, the key conceptual claim that this paradigm is a "necessary complement" to model-centric approaches is asserted rather than demonstrated, and the paper's own cited evidence raises substantial concerns about the stability and reproducibility of agent behavior that the paradigm presupposes. The internal contradictions in summarizing key references further weaken the synthesis. With revisions that address these issues, the paper could be a valuable roadmap for an emerging field.

major comments (4)
  1. [Section 1, 8] The central claim that AI Agent Behavioral Science is a "necessary complement" to model-centric approaches is asserted rather than demonstrated. The paper surveys examples of emergent behavior, but it does not identify a case where a behavioral-level account yields predictions, explanations, or governance insights that are unavailable in principle from model-centric analysis, nor does it state what evidence would count against the necessity claim. The conclusion should be either softened to a "valuable complement" or supported by an explicit argument and testable criteria.
  2. [Sections 2.1, 6.2, 6.3, 2.4] The paradigm presupposes that agent behavior has sufficient stability and reproducibility to be measured with behavioral-science methods, but the paper's own cited evidence repeatedly shows acute sensitivity to prompts, context, and model version. Section 2.1 reports "LLM sensitivity to prompts" [95, 171]; Section 6.2 notes that a single added emoji can significantly alter outputs [174]; Section 6.3 reports order effects in similarity judgments [156]; and Section 2.4 concedes that evaluations are limited in scale and scenario diversity. The paper never provides test-retest reliability data, variance decompositions, or a comparison of within-condition variability across paraphrases and model versions versus between-condition effects. Without such evidence, the "emergent behavior" the paradigm studies could be largely an artifact of prompt or version variance, undermining the brain-to-behavior analogy. This needs to be addressed explicitly, for example by adding a section on behavioral reliability, measurement invariance, and variance decomposition.
  3. [Sections 2.1 vs 2.2; Table 2] The characterization of Mozikov et al. [107] is internally contradictory. Section 2.1 states that "emotions can influence the strategic decision-making of LLMs in a manner similar to how they affect humans," while Section 2.2 states the same work shows "many LLMs have emotional tendencies distinct from those of humans, making them potentially more rational," and Table 2 lists [107] under "Decision making is not affected by emotions like humans." The same citation is used to support opposite conclusions in adjacent subsections, which undermines the reliability of the synthesis and must be corrected.
  4. [Sections 2 and 5] The paper transfers human behavioral theories—social cognitive theory in Section 2 and the Fogg Behavior Model in Section 5—to LLM-based agents without validating the transfer. The Fogg mapping (ability=pretraining, motivation=RL/fine-tuning, trigger=prompting) is presented as a post-hoc classification; the paper itself acknowledges in Section 7 that most methods "were not originally developed with behavioral theory in mind" and that the framework is retrospective. As presented, the framework cannot be falsified and does not generate novel predictions. The paper should clarify whether these are intended as testable theories or as organizing heuristics, and if the former, specify what empirical observations would disconfirm them.
minor comments (5)
  1. [Section 7] The first research direction heading contains a typo: "ehavior" should be "behavior."
  2. [Section 5.2] The citation [79] for the "dual-reward reinforcement learning architecture" appears mismatched: the listed reference is "Socially situated artificial intelligence enables learning from human interaction," which does not obviously describe the dual-reward RL method summarized in the text. Please verify and correct the citation.
  3. [References] Reference [85] is missing a year and publication status; complete the bibliographic details.
  4. [Figure 4] Figure 4 has cramped and overlapping labels, making the Fogg-model mapping difficult to read; a cleaner layout would improve clarity.
  5. [Table 1] The row describing "ontological assumption" under the behavioral perspective reads more like a normative stance than a descriptive contrast; consider rephrasing to maintain the table's analytic tone.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a position/survey whose claims are conceptual and whose evidence is independent; the one retrospective framework is explicitly acknowledged as retrospective.

full rationale

This paper does not derive quantitative predictions from fitted parameters. Its central proposal—that AI agent behavior should be studied as situated, emergent, and context-dependent—is advanced as a conceptual paradigm and supported by a broad literature, most of it external to the author group. The only framework that maps external theory onto AI methods is the Fogg Behavior Model reinterpretation in Section 5; the paper explicitly states that existing methods 'were not originally developed with behavioral theory in mind' and 'emerged through empirical iteration,' and that the triadic mapping is used to 'retrospectively organize and interpret' them. That is an organizing lens, not a circular derivation. Self-cited systems (S3 [61], EconAgent [86], AgentSquare [133], OpenCity [176]) appear as illustrative examples in taxonomies of emergent behavior; they are not used as the premise for the paradigm, and no uniqueness claim or theorem from prior work is imported to force the paper's framing. Consequently, no equation or argument reduces the paper's conclusions to its inputs, and no load-bearing self-citation chain exists.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

This ledger captures the paper's unstated premises. There are no fitted numerical parameters because the paper is a narrative synthesis. It depends on three domain assumptions: that AI agents can be studied like human subjects, that human behavioral theories transfer to agents, and that sandbox environments represent deployment. The only invented entity is the paradigm label itself.

assumptions (3)
  • domain assumption LLM-based agents can be meaningfully studied as behavioral entities with stable regularities, analogous to human subjects in behavioral experiments.
    Invoked in Section 1 with the brain-to-action analogy and used throughout the paper; the whole paradigm depends on this transfer of behavioral-science methods to AI agents.
  • domain assumption Social cognitive theory and the Fogg Behavior Model apply to AI agents with the same causal roles as in humans.
    Social cognitive theory frames Section 2 and Fogg's model frames Section 5. The mappings of ability to pretraining, motivation to reinforcement learning, and trigger to prompting are asserted, not derived.
  • domain assumption Sandbox environments and games used in the cited studies are representative of real-world deployment contexts.
    Sections 3 and 4 generalize from simulated towns, social deduction games, negotiation experiments, and macroeconomic simulations to claims about AI agent behavior in practice.
invented entities (1)
  • AI Agent Behavioral Science paradigm
    purpose: A newly proposed discipline label and research agenda that treats behavior in context as the unit of analysis for studying and governing AI agents.
    No falsifiable claim or measurable output is attached to the paradigm itself; its utility is asserted through selected examples and analogies. This is a conceptual construct, not an independently observable entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Agent Behavioral Science." pith.science (2026). https://pith.science/paper/I35BS4PP

@misc{pith2026250606366,
  author       = {Pith},
  title        = {Pith review of: AI Agent Behavioral Science},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I35BS4PP}},
  note         = {Machine review of arXiv:2506.06366}
}
read the original abstract

Recent advances in large language models (LLMs) have enabled the development of AI agents that exhibit increasingly human-like behaviors, including planning, adaptation, and social dynamics across diverse, interactive, and open-ended scenarios. These behaviors are not solely the product of the internal architectures of the underlying models, but emerge from their integration into agentic systems operating within specific contexts, where environmental factors, social cues, and interaction feedbacks shape behavior over time. This evolution necessitates a new scientific perspective: AI Agent Behavioral Science. Rather than focusing only on internal mechanisms, this perspective emphasizes the systematic observation of behavior, design of interventions to test hypotheses, and theory-guided interpretation of how AI agents act, adapt, and interact over time. We systematize a growing body of research across individual agent, multi-agent, and human-agent interaction settings, and further demonstrate how this perspective informs responsible AI by treating fairness, safety, interpretability, accountability, and privacy as behavioral properties. By unifying recent findings and laying out future directions, we position AI Agent Behavioral Science as a necessary complement to traditional model-centric approaches, providing essential tools for understanding, evaluating, and governing the real-world behavior of increasingly autonomous AI systems.

Figures

Figures reproduced from arXiv: 2506.06366 by the authors.

Figure 1
Figure 1. Development of AI technologies and understanding of AI agent behavior. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Determinants of individual AI agent behavior: a social cognitive perspective. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Three types of multi-agent interaction dynamics. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Fogg Behavior Model-informed framework for AI agent behavior adaptation. [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Exemplifying four types of prompting on a shared task. [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]
Figure 6
Figure 6. Figure 6: Examples of how AI Agent Behavioral Science informs the measurement and optimization [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.

Reference graph

Works this paper leans on

196 extracted references · 14 canonical work pages · cited by 1 Pith paper

  1. [107]

    Eai: Emotional decision-making of llms in strategic games and ethical dilemmas

    Mikhail Mozikov, Nikita Severin, Valeria Bodishtianu, Maria Glushanina, Ivan Nasonov, Daniil Orekhov, Pekhotin Vladislav, Ivan Makovetskiy, Mikhail Baklashkin, Vasily Lavren- tyev, et al. Eai: Emotional decision-making of llms in strategic games and ethical dilemmas. Advances in Neural Information Processing Systems, 37:53969–54002, 2024

  2. [174]

    An llm can fool itself: A prompt-based adversarial attack.arXiv preprint arXiv:2310.13345, 2023

    Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan Kankanhalli. An llm can fool itself: A prompt-based adversarial attack.arXiv preprint arXiv:2310.13345, 2023

  3. [156]

    Investigating context effects in similarity judgements in large language models, 2024

    Sagar Uprety, Amit Kumar Jaiswal, Haiming Liu, and Dawei Song. Investigating context effects in similarity judgements in large language models, 2024

  4. [1]

    Coopera- tion, competition, and maliciousness: Llm-stakeholders interactive negotiation.Advances in Neural Information Processing Systems, 37:83548–83599, 2024

    Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Sch ¨onherr, and Mario Fritz. Coopera- tion, competition, and maliciousness: Llm-stakeholders interactive negotiation.Advances in Neural Information Processing Systems, 37:83548–83599, 2024

  5. [2]

    Large language models show human-like content biases in transmission chain experiments.Proceedings of the National Academy of Sciences, 120(44):e2313790120, 2023

    Alberto Acerbi and Joseph M Stubbersfield. Large language models show human-like content biases in transmission chain experiments.Proceedings of the National Academy of Sciences, 120(44):e2313790120, 2023

  6. [3]

    Playing repeated games with large language models.arXiv preprint arXiv:2305.16867, 2023

    Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz. Playing repeated games with large language models.arXiv preprint arXiv:2305.16867, 2023

  7. [4]

    Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

  8. [5]

    Characterizing the role of bots’ in polarized stance on social media.Social Network Analysis and Mining, 12(1):30, 2022

    Abeer Aldayel and Walid Magdy. Characterizing the role of bots’ in polarized stance on social media.Social Network Analysis and Mining, 12(1):30, 2022

Show all 196 references
  1. [6]

    Investigating Cultural Alignment of Large Language Models

    Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. Investigating Cultural Alignment of Large Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computa- tional Linguistics (...

  2. [7]

    Direct preference optimization with an offset

    Afra Amini, Tim Vieira, and Ryan Cotterell. Direct preference optimization with an offset. arXiv preprint arXiv:2402.10571, 2024

  3. [8]

    Embodied cognition: A field guide.Artificial intelligence, 149(1):91– 130, 2003

    Michael L Anderson. Embodied cognition: A field guide.Artificial intelligence, 149(1):91– 130, 2003. 30

  4. [9]

    The moral machine experiment.Nature, 563(7729):59–64, 2018

    Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shar- iff, Jean-Franc ¸ois Bonnefon, and Iyad Rahwan. The moral machine experiment.Nature, 563(7729):59–64, 2018

  5. [10]

    Explicitly un- biased large language models still form biased associations.Proceedings of the National Academy of Sciences, 122(8):e2416228122, 2025

    Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths. Explicitly un- biased large language models still form biased associations.Proceedings of the National Academy of Sciences, 122(8):e2416228122, 2025

  6. [11]

    Emergent tool use from multi-agent autocurricula.ArXiv, abs/1909.07528, 2019

    Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch. Emergent tool use from multi-agent autocurricula.ArXiv, abs/1909.07528, 2019

  7. [12]

    Miller, Sandra Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David J

    Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sandra Mitts, Adithya Renduchintala, Step...

  8. [13]

    Social cognitive theory: An agentic perspective.Annual review of psychol- ogy, 52(1):1–26, 2001

    Albert Bandura. Social cognitive theory: An agentic perspective.Annual review of psychol- ogy, 52(1):1–26, 2001

  9. [14]

    Preventing rogue agents improves multi-agent collab- oration.arXiv preprint arXiv:2502.05986, 2025

    Ohav Barbi, Ori Yoran, and Mor Geva. Preventing rogue agents improves multi-agent collab- oration.arXiv preprint arXiv:2502.05986, 2025

  10. [15]

    Multi-agent large language models for conversational task-solving.arXiv preprint arXiv:2410.22932, 2024

    Jonas Becker. Multi-agent large language models for conversational task-solving.arXiv preprint arXiv:2410.22932, 2024

  11. [16]

    Reflective multi-agent collaboration based on large language models.Advances in Neural Information Processing Systems, 37:138595–138631, 2024

    Xiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng, Lei Wang, Rui Li, Xu Chen, and Ji- Rong Wen. Reflective multi-agent collaboration based on large language models.Advances in Neural Information Processing Systems, 37:138595–138631, 2024

  12. [17]

    Machine culture.Nature Human Behaviour, 7(11):1855–1868, 2023

    Levin Brinkmann, Fabian Baumann, Jean-Franc ¸ois Bonnefon, Maxime Derex, Thomas F M¨uller, Anne-Marie Nussberger, Agnieszka Czaplicka, Alberto Acerbi, Thomas L Griffiths, Joseph Henrich, et al. Machine culture.Nature Human Behaviour, 7(11):1855–1868, 2023

  13. [18]

    Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2022

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2022

  14. [19]

    Scalable ai safety via doubly- efficient debate.arXiv preprint arXiv:2311.14125, 2023

    Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras. Scalable ai safety via doubly- efficient debate.arXiv preprint arXiv:2311.14125, 2023

  15. [20]

    Behavioural science is unlikely to change the world without a heterogeneity revolution.Nature human behaviour, 5(8):980– 989, 2021

    Christopher J Bryan, Elizabeth Tipton, and David S Yeager. Behavioural science is unlikely to change the world without a heterogeneity revolution.Nature human behaviour, 5(8):980– 989, 2021

  16. [21]

    Motivat- ing voter turnout by invoking the self.Proceedings of the National Academy of Sciences, 108(31):12653–12656, 2011

    Christopher J Bryan, Gregory M Walton, Todd Rogers, and Carol S Dweck. Motivat- ing voter turnout by invoking the self.Proceedings of the National Academy of Sciences, 108(31):12653–12656, 2011

  17. [22]

    Discovering latent knowledge in language models without supervision.arXiv preprint arXiv:2212.03827, 2022

    Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision.arXiv preprint arXiv:2212.03827, 2022

  18. [23]

    ´Angel Alexander Cabrera, Adam Perer, and Jason I. Hong. Improving human-ai collaboration with descriptions of ai behavior.Proc. ACM Hum.-Comput. Interact., 7(CSCW1), April 2023

  19. [24]

    Chateval: Towards better llm-based evaluators through multi-agent debate

    Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201, 2023

  20. [25]

    Autoagents: A framework for automatic agent generation.arXiv preprint arXiv:2309.17288, 2023

    Guangyao Chen, Siwei Dong, Yu Shu, Ge Zhang, Jaward Sesay, B ¨orje F Karlsson, Jie Fu, and Yemin Shi. Autoagents: A framework for automatic agent generation.arXiv preprint arXiv:2309.17288, 2023. 31

  21. [26]

    Step-level value preference optimiza- tion for mathematical reasoning.arXiv preprint arXiv:2406.10858, 2024

    Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan. Step-level value preference optimiza- tion for mathematical reasoning.arXiv preprint arXiv:2406.10858, 2024

  22. [27]

    Multi-agent consensus seeking via large language models.arXiv preprint arXiv:2310.20151, 2023

    Huaben Chen, Wenkang Ji, Lufeng Xu, and Shiyu Zhao. Multi-agent consensus seeking via large language models.arXiv preprint arXiv:2310.20151, 2023

  23. [28]

    Put your money where your mouth is: Evaluating strategic planning and execution of llm agents in an auction arena.arXiv preprint arXiv:2310.05746, 2023

    Jiangjie Chen, Siyu Yuan, Rong Ye, Bodhisattwa Prasad Majumder, and Kyle Richardson. Put your money where your mouth is: Evaluating strategic planning and execution of llm agents in an auction arena.arXiv preprint arXiv:2310.05746, 2023

  24. [29]

    S-agents: Self-organizing agents in open-ended environments.arXiv preprint arXiv:2402.04578, 2024

    Jiaqi Chen, Yuxian Jiang, Jiachen Lu, and Li Zhang. S-agents: Self-organizing agents in open-ended environments.arXiv preprint arXiv:2402.04578, 2024

  25. [30]

    Decision transformer: Reinforcement learn- ing via sequence modeling.Advances in neural information processing systems, 34:15084– 15097, 2021

    Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learn- ing via sequence modeling.Advances in neural information processing systems, 34:15084– 15097, 2021

  26. [31]

    Grath: Gradual self-truthifying for large language models.arXiv preprint arXiv:2401.12292, 2024

    Weixin Chen, Dawn Song, and Bo Li. Grath: Gradual self-truthifying for large language models.arXiv preprint arXiv:2401.12292, 2024

  27. [32]

    Agentverse: Facilitating multi-agent collabo- ration and exploring emergent behaviors in agents.arXiv preprint arXiv:2308.10848, 2(4):6, 2023

    Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al. Agentverse: Facilitating multi-agent collabo- ration and exploring emergent behaviors in agents.arXiv preprint arXiv:2308.10848, 2(4):6, 2023

  28. [33]

    Emotionqueen: A benchmark for evaluating empathy of large language models.arXiv preprint arXiv:2409.13359, 2024

    Yuyan Chen, Hao Wang, Songzhou Yan, Sijia Liu, Yueze Li, Yi Zhao, and Yanghua Xiao. Emotionqueen: A benchmark for evaluating empathy of large language models.arXiv preprint arXiv:2409.13359, 2024

  29. [34]

    Deep reinforcement learning from human preferences.Advances in neural information pro- cessing systems, 30, 2017

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences.Advances in neural information pro- cessing systems, 30, 2017

  30. [35]

    Simulating opinion dynamics with networks of llm-based agents.arXiv preprint arXiv:2311.09618, 2023

    Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. Simulating opinion dynamics with networks of llm-based agents.arXiv preprint arXiv:2311.09618, 2023

  31. [36]

    The wisdom of partisan crowds: Comparing collective intelligence in humans and llm-based agents.arXiv preprint arXiv:2311.09665, 2023

    Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. The wisdom of partisan crowds: Comparing collective intelligence in humans and llm-based agents.arXiv preprint arXiv:2311.09665, 2023

  32. [37]

    MIT press, 1998

    Andy Clark.Being there: Putting brain, body, and world together again. MIT press, 1998

  33. [38]

    What does bert look at? an analysis of bert’s attention.arXiv preprint arXiv:1906.04341, 2019

    Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. What does bert look at? an analysis of bert’s attention.arXiv preprint arXiv:1906.04341, 2019

  34. [39]

    Cogbench: a large language model walks into a psychology lab.arXiv preprint arXiv:2402.18225, 2024

    Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab.arXiv preprint arXiv:2402.18225, 2024

  35. [40]

    Durably reducing conspiracy beliefs through dialogues with ai.Science, 385(6714):eadq1814, 2024

    Thomas H Costello, Gordon Pennycook, and David G Rand. Durably reducing conspiracy beliefs through dialogues with ai.Science, 385(6714):eadq1814, 2024

  36. [41]

    Ultrafeedback: Boosting language models with high-quality feedback

    Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. Ultrafeedback: Boosting language models with high-quality feedback. 2023

  37. [42]

    Artificial leviathan: Exploring social evolution of llm agents through the lens of hobbe- sian social contract theory.arXiv preprint arXiv:2406.14373, 2024

    Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Srihas Rao, Arthur Caetano, Misha Sra, et al. Artificial leviathan: Exploring social evolution of llm agents through the lens of hobbe- sian social contract theory.arXiv preprint arXiv:2406.14373, 2024

  38. [43]

    Behavioural nudges increase covid-19 vaccinations.Nature, 597(7876):404–409, 2021

    Hengchen Dai, Silvia Saccardo, Maria A Han, Lily Roh, Naveen Raja, Sitaram Vangala, Hardikkumar Modi, Shital Pandya, Michael Sloyan, and Daniel M Croymans. Behavioural nudges increase covid-19 vaccinations.Nature, 597(7876):404–409, 2021. 32

  39. [44]

    Mmrole: A com- prehensive framework for developing and evaluating multimodal role-playing agents.arXiv preprint arXiv:2408.04203, 2024

    Yanqi Dai, Huanran Hu, Lei Wang, Shengjie Jin, Xu Chen, and Zhiwu Lu. Mmrole: A com- prehensive framework for developing and evaluating multimodal role-playing agents.arXiv preprint arXiv:2408.04203, 2024

  40. [45]

    Springer Science & Business Media, 2013

    Edward L Deci and Richard M Ryan.Intrinsic motivation and self-determination in human behavior. Springer Science & Business Media, 2013

  41. [46]

    Shared mental models: Ideologies and institutions

    Arthur T Denzau, Douglass C North, et al. Shared mental models: Ideologies and institutions. KYKLOS-BERNE-, 47(1):3–31, 1994

  42. [47]

    Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

  43. [48]

    Springer, 2019

    Virginia Dignum.Responsible artificial intelligence: how to develop and use AI in a respon- sible way, volume 2156. Springer, 2019

  44. [49]

    Humaine: human multi-agent immersive negotiation competition

    Rahul R Divekar, Hui Su, Jeffrey O Kephart, Maira Gratti DeBayser, Melina Guerra, Xi- angyang Mou, Matthew Peveler, and Lisha Chen. Humaine: human multi-agent immersive negotiation competition. InExtended abstracts of the 2020 CHI conference on human factors in computing syste...

  45. [50]

    Towards a rigorous science of interpretable machine learning, 2017

    Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning, 2017

  46. [51]

    Algorithmic agents in the hybrid media system: Social bots, selective amplifica- tion, and partisan news about covid-19.Human Communication Research, 48(3):516–542, 2022

    Zening Duan, Jianing Li, Josephine Lukito, Kai-Cheng Yang, Fan Chen, Dhavan V Shah, and Sijia Yang. Algorithmic agents in the hybrid media system: Social bots, selective amplifica- tion, and partisan news about covid-19.Human Communication Research, 48(3):516–542, 2022

  47. [52]

    EtiCor: Corpus for Analyzing LLMs for Etiquettes

    Ashutosh Dwivedi, Pradhyumna Lavania, and Ashutosh Modi. EtiCor: Corpus for Analyzing LLMs for Etiquettes. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6921– 6931, Singapore,...

  48. [53]

    Chatgpt outperforms humans in emotional awareness evaluations.Frontiers in psychology, 14:1199058, 2023

    Zohar Elyoseph, Dorit Hadar-Shoval, Kfir Asraf, and Maya Lvovsky. Chatgpt outperforms humans in emotional awareness evaluations.Frontiers in psychology, 14:1199058, 2023

  49. [54]

    Brookings Institution Press, 1996

    Joshua M Epstein and Robert Axtell.Growing artificial societies: social science from the bottom up. Brookings Institution Press, 1996

  50. [55]

    Why does unsuper- vised pre-training help deep learning? InProceedings of the thirteenth international confer- ence on artificial intelligence and statistics, pages 201–208

    Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent. Why does unsuper- vised pre-training help deep learning? InProceedings of the thirteenth international confer- ence on artificial intelligence and statistics, pages 201–208. JMLR Workshop and Conference Proceed...

  51. [56]

    Can large language models serve as ra- tional players in game theory? a systematic analysis

    Caoyun Fan, Jindou Chen, Yaohui Jin, and Hao He. Can large language models serve as ra- tional players in game theory? a systematic analysis. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17960–17967, 2024

  52. [57]

    A behavior model for persuasive design

    Brian J Fogg. A behavior model for persuasive design. InProceedings of the 4th international Conference on Persuasive Technology, pages 1–7, 2009

  53. [58]

    Nicer than humans: How do large language models behave in the prisoner’s dilemma?arXiv preprint arXiv:2406.13605, 2024

    Nicol ´o Fontana, Francesco Pierri, and Luca Maria Aiello. Nicer than humans: How do large language models behave in the prisoner’s dilemma?arXiv preprint arXiv:2406.13605, 2024

  54. [59]

    Bias and fairness in large language models: A survey.Computational Linguistics, 50(3):1097–1179, 2024

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. Bias and fairness in large language models: A survey.Computational Linguistics, 50(3):1097–1179, 2024

  55. [60]

    Chang Gao, Haiyun Jiang, Deng Cai, Shuming Shi, and Wai Lam. Strategyllm: Large lan- guage models as strategy generators, executors, optimizers, and evaluators for problem solv- ing.Advances in Neural Information Processing Systems, 37:96797–96846, 2024. 33

  56. [61]

    S3: Social-network simulation system with large language model- empowered agents.arXiv preprint arXiv:2307.14984, 2023

    Chen Gao, Xiaochong Lan, Zhihong Lu, Jinzhu Mao, Jinghua Piao, Huandong Wang, De- peng Jin, and Yong Li. S3: Social-network simulation system with large language model- empowered agents.arXiv preprint arXiv:2307.14984, 2023

  57. [62]

    How human–ai feedback loops alter human perceptual, emotional and social judgements.Nature Human Behaviour, pages 1–15, 2024

    Moshe Glickman and Tali Sharot. How human–ai feedback loops alter human perceptual, emotional and social judgements.Nature Human Behaviour, pages 1–15, 2024

  58. [63]

    Deception abilities emerged in large language models.Proceedings of the National Academy of Sciences, 121(24):e2317967121, 2024

    Thilo Hagendorff. Deception abilities emerged in large language models.Proceedings of the National Academy of Sciences, 121(24):e2317967121, 2024

  59. [64]

    Human-like intuitive behavior and rea- soning biases emerged in large language models but disappeared in chatgpt.Nature Compu- tational Science, 3(10):833–838, 2023

    Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. Human-like intuitive behavior and rea- soning biases emerged in large language models but disappeared in chatgpt.Nature Compu- tational Science, 3(10):833–838, 2023

  60. [65]

    A survey on vision transformer.IEEE transactions on pattern analysis and machine intelligence, 45(1):87–110, 2022

    Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. A survey on vision transformer.IEEE transactions on pattern analysis and machine intelligence, 45(1):87–110, 2022

  61. [66]

    Large language models are bad game theoretic reasoners: Evaluating performance and bias in two- player non-zero-sum games

    Nathan Herr, Fernando Acero, Roberta Raileanu, Maria Perez-Ortiz, and Zhibin Li. Large language models are bad game theoretic reasoners: Evaluating performance and bias in two- player non-zero-sum games. InICML 2024 Workshop on LLMs and Cognition

  62. [67]

    Ai generates covertly racist decisions about people based on their dialect.Nature, 633(8028):147–154, 2024

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. Ai generates covertly racist decisions about people based on their dialect.Nature, 633(8028):147–154, 2024

  63. [68]

    Wang, Chenhui Zhang, Zhangheng LI, Bo Li, and Zhangyang Wang

    Junyuan Hong, Jiachen T. Wang, Chenhui Zhang, Zhangheng LI, Bo Li, and Zhangyang Wang. DP-OPT: Make large language model your privacy-preserving prompt engineer. In The Twelfth International Conference on Learning Representations, 2024

  64. [69]

    Generative Language Models Exhibit Social Identity Biases, June 2024

    Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier, Sander van der Linden, and Jon Roozenbeek. Generative Language Models Exhibit Social Identity Biases, June 2024

  65. [70]

    War and peace (waragent): Large language model-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227, 2023

    Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang. War and peace (waragent): Large language model-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227, 2023

  66. [71]

    Competing large language models in multi-agent gaming environments

    Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael Lyu. Competing large language models in multi-agent gaming environments. InThe Thirteenth International Conference on Learning Representations, 2025

  67. [72]

    Assessing ai utility: The random guesser test for sequential decision-making systems.arXiv preprint arXiv:2407.20276, 2024

    Shun Ide, Allison Blunt, and Djallel Bouneffouf. Assessing ai utility: The random guesser test for sequential decision-making systems.arXiv preprint arXiv:2407.20276, 2024

  68. [73]

    Personalized soups: Per- sonalized large language model alignment via post-hoc parameter merging.arXiv preprint arXiv:2310.11564, 2023

    Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. Personalized soups: Per- sonalized large language model alignment via post-hoc parameter merging.arXiv preprint arXiv:2310.11564, 2023

  69. [74]

    Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny T

    Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny T. Liang, Sydney Levine, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jack Hessel, Jon Borchardt, Tay- lor Sorensen, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi....

  70. [75]

    The global landscape of ai ethics guidelines

    Anna Jobin, Marcello Ienca, and Effy Vayena. The global landscape of ai ethics guidelines. Nature machine intelligence, 1(9):389–399, 2019

  71. [76]

    Hachette UK, 2021

    Daniel Kahneman, Olivier Sibony, and Cass R Sunstein.Noise: A flaw in human judgment. Hachette UK, 2021

  72. [77]

    Porover: Improving safety and reducing overrefusal in large language models with overgeneration and preference optimization.arXiv preprint arXiv:2410.12999, 2024

    Batuhan K Karaman, Ishmam Zabir, Alon Benhaim, Vishrav Chaudhary, Mert R Sabuncu, and Xia Song. Porover: Improving safety and reducing overrefusal in large language models with overgeneration and preference optimization.arXiv preprint arXiv:2410.12999, 2024. 34

  73. [78]

    Hauser, Duncan Williams, Lucy Campbell-Gillingham, Phoebe Thacker, Matthew M

    Raphael Koster, Jan Balaguer, Andrea Tacchetti, Ari Weinstein, Tina Zhu, Oliver P. Hauser, Duncan Williams, Lucy Campbell-Gillingham, Phoebe Thacker, Matthew M. Botvinick, and Christopher Summerfield. Human-centered mechanism design with democratic ai.ArXiv, abs/2201.11441, 2022

  74. [79]

    Socially situated artificial intelligence enables learning from human interaction.Proceedings of the National Academy of Sciences, 119(39):e2115730119, 2022

    Ranjay Krishna, Donsuk Lee, Li Fei-Fei, and Michael S Bernstein. Socially situated artificial intelligence enables learning from human interaction.Proceedings of the National Academy of Sciences, 119(39):e2115730119, 2022

  75. [80]

    Understanding the effects of iterative prompting on truthfulness.arXiv preprint arXiv:2402.06625, 2024

    Satyapriya Krishna, Chirag Agarwal, and Himabindu Lakkaraju. Understanding the effects of iterative prompting on truthfulness.arXiv preprint arXiv:2402.06625, 2024

  76. [81]

    The first workshop on ai behavioral science

    Himabindu Lakkaraju, Qiaozhu Mei, Chenhao Tan, Jie Tang, and Yutong Xie. The first workshop on ai behavioral science. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 6724–6725, 2024

  77. [82]

    Llm-based agent society investigation: Collaboration and confronta- tion in avalon gameplay.arXiv preprint arXiv:2310.14985, 2023

    Yihuai Lan, Zhiqiang Hu, Lei Wang, Yang Wang, Deheng Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong, and Hao Wang. Llm-based agent society investigation: Collaboration and confronta- tion in avalon gameplay.arXiv preprint arXiv:2310.14985, 2023

  78. [83]

    Uncovering the semantics of concepts using gpt-4.Proceedings of the National Academy of Sciences, 120(49):e2309350120, 2023

    Ga ¨el Le Mens, Bal ´azs Kov ´acs, Michael T Hannan, and Guillem Pros. Uncovering the semantics of concepts using gpt-4.Proceedings of the National Academy of Sciences, 120(49):e2309350120, 2023

  79. [84]

    Oxford University Press Oxford, 164–182, 2022

    Theodore M Lechterman.The concept of accountability in AI ethics and governance. Oxford University Press Oxford, 164–182, 2022

  80. [85]

    Prompting fairness: Integrating causality to debias large language models

    Jingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes, Kun Zhang, Liu Leqi, and Yang Liu. Prompting fairness: Integrating causality to debias large language models. InThe Thirteenth International Conference on Learning Representations

  81. [86]

    Econagent: large lan- guage model-empowered agents for simulating macroeconomic activities.arXiv preprint arXiv:2310.10436, 2023

    Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qingmin Liao. Econagent: large lan- guage model-empowered agents for simulating macroeconomic activities.arXiv preprint arXiv:2310.10436, 2023

  82. [87]

    Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents.arXiv preprint arXiv:2310.06500, 2023

    Yuan Li, Yixuan Zhang, and Lichao Sun. Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents.arXiv preprint arXiv:2310.06500, 2023

  83. [88]

    Zhuoyan Li, Zhuoran Lu, and Ming Yin. Decoding ai’s nudge: A unified framework to predict human behavior in ai-assisted decision making.Proceedings of the AAAI Conference on Artificial Intelligence, 38(9):10083–10091, Mar. 2024

  84. [89]

    Strategy selection as rational metareasoning.Psycho- logical review, 124(6):762, 2017

    Falk Lieder and Thomas L Griffiths. Strategy selection as rational metareasoning.Psycho- logical review, 124(6):762, 2017

  85. [90]

    Toward a better understanding of the emo- tional dynamics of negotiation with large language models

    Eleanor Lin, James Hale, and Jonathan Gratch. Toward a better understanding of the emo- tional dynamics of negotiation with large language models. InProceedings of the Twenty- fourth International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Net...

  86. [91]

    Truthfulqa: Measuring how models mimic human falsehoods.arXiv preprint arXiv:2109.07958, 2021

    Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods.arXiv preprint arXiv:2109.07958, 2021

  87. [92]

    Explainable ai: A review of machine learning interpretability methods.Entropy, 23(1):18, 2020

    Pantelis Linardatos, Vasilis Papastefanopoulos, and Sotiris Kotsiantis. Explainable ai: A review of machine learning interpretability methods.Entropy, 23(1):18, 2020

  88. [93]

    Lidao: towards limited interventions for debiasing (large) language models.arXiv preprint arXiv:2406.00548, 2024

    Tianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng, and Jing Gao. Lidao: towards limited interventions for debiasing (large) language models.arXiv preprint arXiv:2406.00548, 2024

  89. [94]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InPro- ceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 35

  90. [95]

    Strategic behavior of large language models and the role of game structure versus contextual framing.Scientific Reports, 14(1):18490, 2024

    Nunzio Lor `e and Babak Heydari. Strategic behavior of large language models and the role of game structure versus contextual framing.Scientific Reports, 14(1):18490, 2024

  91. [96]

    Multi-agent actor-critic for mixed cooperative-competitive environments

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. InProceedings of the 31st Inter- national Conference on Neural Information Processing Systems, NIPS’17, page 6382–6393, Red Hook,...

  92. [97]

    Llm discussion: Enhancing the creativity of large language models via discussion frame- work and role-play.arXiv preprint arXiv:2405.06373, 2024

    Li-Chun Lu, Shou-Jen Chen, Tsung-Min Pai, Chan-Hung Yu, Hung-yi Lee, and Shao-Hua Sun. Llm discussion: Enhancing the creativity of large language models via discussion frame- work and role-play.arXiv preprint arXiv:2405.06373, 2024

  93. [98]

    Down the bot hole: Actionable insights from a one-year analysis of bot activity on twitter.First Monday, 2021

    Luca Luceri, Felipe Cardoso, and Silvia Giordano. Down the bot hole: Actionable insights from a one-year analysis of bot activity on twitter.First Monday, 2021

  94. [99]

    Red bots do it better: Compar- ative analysis of social bot partisan behavior

    Luca Luceri, Ashok Deb, Adam Badawy, and Emilio Ferrara. Red bots do it better: Compar- ative analysis of social bot partisan behavior. InCompanion proceedings of the 2019 world wide web conference, pages 1007–1012, 2019

  95. [100]

    Reft: Reasoning with reinforced fine-tuning.arXiv preprint arXiv:2401.08967, 3, 2024

    Trung Quoc Luong, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. Reft: Reasoning with reinforced fine-tuning.arXiv preprint arXiv:2401.08967, 3, 2024

  96. [101]

    Eureka: Human-level reward design via coding large language models.arXiv preprint arXiv:2310.12931, 2023

    Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Eureka: Human-level reward design via coding large language models.arXiv preprint arXiv:2310.12931, 2023

  97. [102]

    O’Reilly Media, Incorpo- rated, 2020

    Trisha Mahoney, Kush Varshney, and Michael Hind.AI fairness. O’Reilly Media, Incorpo- rated, 2020

  98. [103]

    A turing test of whether ai chatbots are behaviorally similar to humans.Proceedings of the National Academy of Sciences, 121(9):e2313925121, 2024

    Qiaozhu Mei, Yutong Xie, Walter Yuan, and Matthew O Jackson. A turing test of whether ai chatbots are behaviorally similar to humans.Proceedings of the National Academy of Sciences, 121(9):e2313925121, 2024

  99. [104]

    Results of the first annual human-agent league of the automated negotiating agents competi- tion

    Johnathan Mell, Jonathan Gratch, Tim Baarslag, Reyhan Aydo ˘gan, and Catholijn M Jonker. Results of the first annual human-agent league of the automated negotiating agents competi- tion. InProceedings of the 18th International Conference on Intelligent Virtual Agents, pages 23...

  100. [105]

    Ai emerges as the frontier in behavioral science.Proceedings of the National Academy of Sciences, 121(10):e2401336121, 2024

    Juanjuan Meng. Ai emerges as the frontier in behavioral science.Proceedings of the National Academy of Sciences, 121(10):e2401336121, 2024

  101. [106]

    Secret collusion among ai agents: Multi- agent deception via steganography.Advances in Neural Information Processing Systems, 37:73439–73486, 2024

    Sumeet Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip Torr, Lewis Hammond, and Christian Schroeder de Witt. Secret collusion among ai agents: Multi- agent deception via steganography.Advances in Neural Information Processing Systems, 37:73439–73486, 2024

  102. [108]

    Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew A

    Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki A. Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew A. Ayele, V´ıctor Guti´errez-Basulto, Yazm´ın Ib´a˜nez-Garc´ıa, Hwaran Lee, Shamsuddeen H. Muhammad, Kiwoong Park, Anar S. Rzayev, Nina W...

  103. [109]

    Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth

    Thao Nguyen, Maithra Raghu, and Simon Kornblith. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. arXiv preprint arXiv:2010.15327, 2020. 36

  104. [110]

    Extracting Cultural Commonsense Knowledge at Scale

    Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. Extracting Cultural Commonsense Knowledge at Scale. InProceedings of the ACM Web Conference 2023, pages 1907–1917, Austin TX USA, April 2023. ACM

  105. [111]

    Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models

    Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877, 2021

  106. [112]

    Accountability in artificial intel- ligence: what it is and how it works.AI & SOCIETY, 39:1–12, 02 2023

    Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. Accountability in artificial intel- ligence: what it is and how it works.AI & SOCIETY, 39:1–12, 02 2023

  107. [113]

    Hoodwinked: Deception and cooperation in a text-based game for language models.arXiv preprint arXiv:2308.01404, 2023

    Aidan O’Gara. Hoodwinked: Deception and cooperation in a text-based game for language models.arXiv preprint arXiv:2308.01404, 2023

  108. [114]

    Agentcoord: Visually exploring coordination strategy for llm-based multi-agent collaboration.arXiv preprint arXiv:2404.11943, 2024

    Bo Pan, Jiaying Lu, Ke Wang, Li Zheng, Zhen Wen, Yingchaojie Feng, Minfeng Zhu, and Wei Chen. Agentcoord: Visually exploring coordination strategy for llm-based multi-agent collaboration.arXiv preprint arXiv:2404.11943, 2024

  109. [115]

    Generative agents: Interactive simulacra of human behavior

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. InPro- ceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22, 2023

  110. [116]

    Cooperate or collapse: Emergence of sustainable cooperation in a soci- ety of llm agents.Advances in Neural Information Processing Systems, 37:111715–111759, 2024

    Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Sch ¨olkopf, Mrinmaya Sachan, and Rada Mihalcea. Cooperate or collapse: Emergence of sustainable cooperation in a soci- ety of llm agents.Advances in Neural Information Processing Systems, 37:111715–111759, 2024

  111. [117]

    Harnessing the power of large language mod- els for empathetic response generation: Empirical investigations and improvements.arXiv preprint arXiv:2310.05140, 2023

    Yushan Qian, Wei-Nan Zhang, and Ting Liu. Harnessing the power of large language mod- els for empathetic response generation: Empirical investigations and improvements.arXiv preprint arXiv:2310.05140, 2023

  112. [118]

    Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

  113. [119]

    Llms among us: Generative ai par- ticipating in digital discourse

    Kristina Radivojevic, Nicholas Clark, and Paul Brenner. Llms among us: Generative ai par- ticipating in digital discourse. InProceedings of the AAAI Symposium Series, volume 3, pages 209–218, 2024

  114. [120]

    Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

  115. [121]

    Machine behaviour.Nature, 568(7753):477–486, 2019

    Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-Franc ¸ois Bonnefon, Cynthia Breazeal, Jacob W Crandall, Nicholas A Christakis, Iain D Couzin, Matthew O Jack- son, et al. Machine behaviour.Nature, 568(7753):477–486, 2019

  116. [122]

    Steer: Assessing the economic rationality of large language models

    Narun Raman, Taylor Lundy, Samuel Amouyal, Yoav Levine, Kevin Leyton-Brown, and Moshe Tennenholtz. Steer: Assessing the economic rationality of large language models. arXiv preprint arXiv:2402.09552, 2024

  117. [123]

    Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data.arXiv preprint arXiv:2406.18921, 2024

    Yiting Ran, Xintao Wang, Rui Xu, Xinfeng Yuan, Jiaqing Liang, Deqing Yang, and Yanghua Xiao. Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data.arXiv preprint arXiv:2406.18921, 2024

  118. [124]

    Mitigating bias in conversations: a hate speech classifier and debiaser with prompts.arXiv preprint arXiv:2307.10213, 2023

    Shaina Raza, Chen Ding, and Deval Pandya. Mitigating bias in conversations: a hate speech classifier and debiaser with prompts.arXiv preprint arXiv:2307.10213, 2023

  119. [125]

    A generalist agent.arXiv preprint arXiv:2205.06175, 2022

    Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent.arXiv preprint arXiv:2205.06175, 2022

  120. [126]

    Lamp: When large language models meet personalization.arXiv preprint arXiv:2304.11406, 2023

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. Lamp: When large language models meet personalization.arXiv preprint arXiv:2304.11406, 2023. 37

  121. [127]

    Whose Opinions Do Language Models Reflect? InProceedings of the 40th International Conference on Machine Learning, pages 29971–30004

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose Opinions Do Language Models Reflect? InProceedings of the 40th International Conference on Machine Learning, pages 29971–30004. PMLR, July 2023

  122. [128]

    Training language models for social deduction with multi-agent reinforcement learning.arXiv preprint arXiv:2502.06060, 2025

    Bidipta Sarkar, Warren Xia, C Karen Liu, and Dorsa Sadigh. Training language models for social deduction with multi-agent reinforcement learning.arXiv preprint arXiv:2502.06060, 2025

  123. [129]

    Large language models can strate- gically deceive their users when put under pressure

    J ´er´emy Scheurer, Mikita Balesni, and Marius Hobbhahn. Large language models can strate- gically deceive their users when put under pressure. InICLR 2024 Workshop on Large Lan- guage Model (LLM) Agents, 2024

  124. [130]

    Team reflexivity and in- novation: The moderating role of team context.Journal of Management, 41(3):769–788, 2015

    Micha ´ela C Schippers, Michael A West, and Jeremy F Dawson. Team reflexivity and in- novation: The moderating role of team context.Journal of Management, 41(3):769–788, 2015

  125. [131]

    Negotiating with llms: Prompt hacks, skill gaps, and reasoning deficits

    Johannes Schneider, Steffi Haag, and Leona Chandra Kruse. Negotiating with llms: Prompt hacks, skill gaps, and reasoning deficits. InInternational Conference on Computer-Human Interaction Research and Applications, pages 238–259. Springer, 2024

  126. [132]

    Rewarding progress: Scaling auto- mated process verifiers for llm reasoning.arXiv preprint arXiv:2410.08146, 2024

    Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agar- wal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar. Rewarding progress: Scaling auto- mated process verifiers for llm reasoning.arXiv preprint arXiv:2410.08146, 2024

  127. [133]

    Agentsquare: Automatic llm agent search in modular design space.arXiv preprint arXiv:2410.06153, 2024

    Yu Shang, Yu Li, Keyu Zhao, Likai Ma, Jiahe Liu, Fengli Xu, and Yong Li. Agentsquare: Automatic llm agent search in modular design space.arXiv preprint arXiv:2410.06153, 2024

  128. [134]

    The spread of low-credibility content by social bots.Nature communications, 9(1):4787, 2018

    Chengcheng Shao, Giovanni Luca Ciampaglia, Onur Varol, Kai-Cheng Yang, Alessandro Flammini, and Filippo Menczer. The spread of low-credibility content by social bots.Nature communications, 9(1):4787, 2018

  129. [135]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

  130. [136]

    The dynamics of collec- tive creativity in human-ai social networks.arXiv preprint arXiv:2502.17962, 2025

    Shota Shiiku, Raja Marjieh, Manuel Anglada-Tort, and Nori Jacoby. The dynamics of collec- tive creativity in human-ai social networks.arXiv preprint arXiv:2502.17962, 2025

  131. [137]

    Superhuman artificial intelligence can improve human decision-making by increasing novelty.Proceedings of the National Academy of Sciences, 120(12):e2214840120, 2023

    Minkyu Shin, Jin Kim, Bas Van Opheusden, and Thomas L Griffiths. Superhuman artificial intelligence can improve human decision-making by increasing novelty.Proceedings of the National Academy of Sciences, 120(12):e2214840120, 2023

  132. [138]

    Locally noisy autonomous agents improve global human coordination in network experiments.Nature, 545(7654):370–374, 2017

    Hirokazu Shirado and Nicholas A Christakis. Locally noisy autonomous agents improve global human coordination in network experiments.Nature, 545(7654):370–374, 2017

  133. [139]

    Lillicrap, Fan Hui, L

    David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, L. Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering ...

  134. [140]

    Explaining ma- chine learning models with interactive natural language conversations using TalkToModel

    Dylan Slack, Satyapriya Krishna, Himabindu Lakkaraju, and Sameer Singh. Explaining ma- chine learning models with interactive natural language conversations using TalkToModel. Nature Machine Intelligence, 5(8):873–883, 2023

  135. [141]

    On early detection of hallucina- tions in factual question answering

    Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. On early detection of hallucina- tions in factual question answering. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2721–2732, 2024

  136. [142]

    Beyond memorization: Vi- olating privacy via inference with large language models

    Robin Staab, Mark Vero, Mislav Balunovic, and Martin Vechev. Beyond memorization: Vi- olating privacy via inference with large language models. InThe Twelfth International Con- ference on Learning Representations, 2024

  137. [143]

    Bots increase exposure to nega- tive and inflammatory content in online social systems.Proceedings of the National Academy of Sciences, 115(49):12435–12440, 2018

    Massimo Stella, Emilio Ferrara, and Manlio De Domenico. Bots increase exposure to nega- tive and inflammatory content in online social systems.Proceedings of the National Academy of Sciences, 115(49):12435–12440, 2018. 38

  138. [144]

    Information gerrymandering and undemocratic decisions.Nature, 573(7772):117–121, 2019

    Alexander J Stewart, Mohsen Mosleh, Marina Diakonova, Antonio A Arechar, David G Rand, and Joshua B Plotkin. Information gerrymandering and undemocratic decisions.Nature, 573(7772):117–121, 2019

  139. [145]

    What large language models know and what people think they know.Nature Machine Intelligence, pages 1–11, 2025

    Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W Mayer, and Padhraic Smyth. What large language models know and what people think they know.Nature Machine Intelligence, pages 1–11, 2025

  140. [146]

    Testing theory of mind in large language models and humans.Nature Human Behaviour, 8(7):1285– 1295, 2024

    James W A Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, et al. Testing theory of mind in large language models and humans.Nature Human Behaviour, 8(7):1285– 1295, 2024

  141. [147]

    Identity-driven hierarchical role-playing agents.arXiv preprint arXiv:2407.19412, 2024

    Libo Sun, Siyuan Wang, Xuanjing Huang, and Zhongyu Wei. Identity-driven hierarchical role-playing agents.arXiv preprint arXiv:2407.19412, 2024

  142. [148]

    Vintage, 2005

    James Surowiecki.The wisdom of crowds. Vintage, 2005

  143. [149]

    Medagents: Large language models as collaborators for zero-shot medical reasoning.arXiv preprint arXiv:2311.10537, 2023

    Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein. Medagents: Large language models as collaborators for zero-shot medical reasoning.arXiv preprint arXiv:2311.10537, 2023

  144. [150]

    Enhancing role- playing systems through aggressive queries: Evaluation and improvement.arXiv preprint arXiv:2402.10618, 2024

    Yihong Tang, Jiao Ou, Che Liu, Fuzheng Zhang, Di Zhang, and Kun Gai. Enhancing role- playing systems through aggressive queries: Evaluation and improvement.arXiv preprint arXiv:2402.10618, 2024

  145. [151]

    Cultural bias and cultural align- ment of large language models.PNAS nexus, 3(9):pgae346, 2024

    Yan Tao, Olga Viberg, Ryan S Baker, and Ren ´e F Kizilcec. Cultural bias and cultural align- ment of large language models.PNAS nexus, 3(9):pgae346, 2024

  146. [152]

    People judge others more harshly after talking to bots.PNAS nexus, 3(9):pgae397, 2024

    Kian Siong Tey, Asaf Mazar, Geoff Tomaino, Angela L Duckworth, and Lyle H Ungar. People judge others more harshly after talking to bots.PNAS nexus, 3(9):pgae397, 2024

  147. [153]

    Penguin, 2009

    Richard H Thaler and Cass R Sunstein.Nudge: Improving decisions about health, wealth, and happiness. Penguin, 2009

  148. [154]

    Vulnerable robots positively shape human conversational dynamics in a human– robot team.Proceedings of the National Academy of Sciences, 117(12):6370–6375, 2020

    Margaret L Traeger, Sarah Strohkorb Sebo, Malte Jung, Brian Scassellati, and Nicholas A Christakis. Vulnerable robots positively shape human conversational dynamics in a human– robot team.Proceedings of the National Academy of Sciences, 117(12):6370–6375, 2020

  149. [155]

    A new sociology of humans and machines.Nature Human Behaviour, 8(10):1864–1876, 2024

    Milena Tsvetkova, Taha Yasseri, Niccolo Pescetelli, and Tobias Werner. A new sociology of humans and machines.Nature Human Behaviour, 8(10):1864–1876, 2024

  150. [157]

    Large language models that replace human participants can harmfully misportray and flatten identity groups.Nature Machine Intelligence, pages 1–12, 2025

    Angelina Wang, Jamie Morgenstern, and John P Dickerson. Large language models that replace human participants can harmfully misportray and flatten identity groups.Nature Machine Intelligence, pages 1–12, 2025

  151. [158]

    Investigating and extending homans’ social ex- change theory with large language model based agents.arXiv preprint arXiv:2502.12450, 2025

    Lei Wang, Zheqing Zhang, and Xu Chen. Investigating and extending homans’ social ex- change theory with large language model based agents.arXiv preprint arXiv:2502.12450, 2025

  152. [159]

    Transactive memory: A contemporary analysis of the group mind

    Daniel M Wegner. Transactive memory: A contemporary analysis of the group mind. In Theories of group behavior, pages 185–208. Springer, 1987

  153. [160]

    Llm-powered autonomous agents

    Lilian Weng. Llm-powered autonomous agents. lilianweng. github. io, jun 2023.URL https://lilianweng. github. io/posts/2023-06-23-agent, 2023

  154. [161]

    Deciphering digital detectives: Un- derstanding llm behaviors and capabilities in multi-agent mystery games.arXiv preprint arXiv:2312.00746, 2023

    Dekun Wu, Haochen Shi, Zhiyuan Sun, and Bang Liu. Deciphering digital detectives: Un- derstanding llm behaviors and capabilities in multi-agent mystery games.arXiv preprint arXiv:2312.00746, 2023. 39

  155. [162]

    Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, and Xiangnan He.β-dpo: Direct preference optimization with dynamicβ.Advances in Neural Information Processing Systems, 37:129944–129966, 2024

  156. [163]

    Autogen: Enabling next-gen llm applications via multi-agent conversation.arXiv preprint arXiv:2308.08155, 2023

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xi- aoyun Zhang, Shaokun Zhang, Jiale Liu, et al. Autogen: Enabling next-gen llm applications via multi-agent conversation.arXiv preprint arXiv:2308.08155, 2023

  157. [164]

    Wang, and Prateek Mittal

    Tong Wu, Ashwinee Panda, Jiachen T. Wang, and Prateek Mittal. Privacy-preserving in- context learning for large language models. InThe Twelfth International Conference on Learning Representations, 2024

  158. [165]

    From role-play to drama-interaction: An llm solution.arXiv preprint arXiv:2405.14231, 2024

    Weiqi Wu, Hongqiu Wu, Lai Jiang, Xingyuan Liu, Jiale Hong, Hai Zhao, and Min Zhang. From role-play to drama-interaction: An llm solution.arXiv preprint arXiv:2405.14231, 2024

  159. [166]

    Shall we team up: Exploring sponta- neous cooperation of competing llm agents

    Zengqing Wu, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian Inhyuk Kwon, Makoto Onizuka, Shaojie Tang, and Chuan Xiao. Shall we team up: Exploring sponta- neous cooperation of competing llm agents. InConference on Empirical Methods in Natural Language Processing, 2024

  160. [167]

    The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101, 2025

    Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101, 2025

  161. [168]

    Behavioral bias of vision-language models: A behavioral finance view

    Yuhang Xiao, yudilin, and Ming-Chang Chiu. Behavioral bias of vision-language models: A behavioral finance view. InICML 2024 Workshop on LLMs and Cognition, 2024

  162. [169]

    Text2reward: Automated dense reward function generation for rein- forcement learning

    Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Qian Luo, Victor Zhong, Yanchao Yang, and Tao Yu. Text2reward: Automated dense reward function generation for rein- forcement learning. InInternational Conference on Learning Representations (ICLR), 2024 (07/05/2024-11/05/202...

  163. [170]

    Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

    Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

  164. [171]

    How different ai chatbots behave? benchmarking large language models in behavioral economics games.arXiv preprint arXiv:2412.12362, 2024

    Yutong Xie, Yiyao Liu, Zhuang Ma, Lin Shi, Xiyuan Wang, Walter Yuan, Matthew O Jackson, and Qiaozhu Mei. How different ai chatbots behave? benchmarking large language models in behavioral economics games.arXiv preprint arXiv:2412.12362, 2024

  165. [172]

    Monte carlo tree search boosts reasoning via iterative pref- erence learning.arXiv preprint arXiv:2405.00451, 2024

    Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P Lillicrap, Kenji Kawaguchi, and Michael Shieh. Monte carlo tree search boosts reasoning via iterative pref- erence learning.arXiv preprint arXiv:2405.00451, 2024

  166. [173]

    Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration.arXiv preprint arXiv:2311.08562, 2023

    Lin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren, Zhen Dong, Kurt Keutzer, See Kiong Ng, and Jiashi Feng. Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration.arXiv preprint arXiv:2311.08562, 2023

  167. [175]

    Exploring large language models for communication games: An empirical study on werewolf.arXiv preprint arXiv:2309.04658, 2023

    Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. Exploring large language models for communication games: An empirical study on werewolf.arXiv preprint arXiv:2309.04658, 2023

  168. [176]

    Opencity: A scalable platform to simulate urban activities with massive llm agents.arXiv preprint arXiv:2410.21286, 2024

    Yuwei Yan, Qingbin Zeng, Zhiheng Zheng, Jingzhe Yuan, Jie Feng, Jun Zhang, Fengli Xu, and Yong Li. Opencity: A scalable platform to simulate urban activities with massive llm agents.arXiv preprint arXiv:2410.21286, 2024

  169. [177]

    Simschat: A customisable persona-driven role- playing agent.arXiv e-prints, pages arXiv–2406, 2024

    Bohao Yang, Dong Liu, Chen Tang, Chenghao Xiao, Kun Zhao, Chao Li, Lin Yuan, Guang Yang, Lanxiao Huang, and Chenghua Lin. Simschat: A customisable persona-driven role- playing agent.arXiv e-prints, pages arXiv–2406, 2024. 40

  170. [178]

    Llm- powered decentralized generative agents with adaptive hierarchical knowledge graph for co- operative planning.arXiv preprint arXiv:2502.05453, 2025

    Hanqing Yang, Jingdi Chen, Marie Siew, Tania Lorido-Botran, and Carlee Joe-Wong. Llm- powered decentralized generative agents with adaptive hierarchical knowledge graph for co- operative planning.arXiv preprint arXiv:2502.05453, 2025

  171. [179]

    Anatomy of an ai-powered malicious social botnet

    Kai-Cheng Yang and Filippo Menczer. Anatomy of an ai-powered malicious social botnet. arXiv preprint arXiv:2307.16336, 2023

  172. [180]

    Prevalence of low-credibility information on twitter during the covid-19 outbreak.arXiv preprint arXiv:2004.14484, 2020

    Kai-Cheng Yang, Christopher Torres-Lugo, and Filippo Menczer. Prevalence of low-credibility information on twitter during the covid-19 outbreak.arXiv preprint arXiv:2004.14484, 2020

  173. [181]

    Qiang Yang. Toward responsible ai: An overview of federated learning for user-centered privacy-preserving computing.ACM Transactions on Interactive Intelligent Systems (TiiS), 11(3-4):1–22, 2021

  174. [182]

    Safe- World: Geo-Diverse Safety Alignment.Advances in Neural Information Processing Systems, 37:128734–128768, January 2025

    Da Yin, Haoyi Qiu, Kung-Hsiang Huang, Kai-Wei Chang, and Nanyun Peng. Safe- World: Geo-Diverse Safety Alignment.Advances in Neural Information Processing Systems, 37:128734–128768, January 2025

  175. [183]

    Neeko: Leveraging dynamic lora for efficient multi-character role-playing agent.arXiv preprint arXiv:2402.13717, 2024

    Xiaoyan Yu, Tongxu Luo, Yifan Wei, Fangyu Lei, Yiming Huang, Hao Peng, and Liehuang Zhu. Neeko: Leveraging dynamic lora for efficient multi-character role-playing agent.arXiv preprint arXiv:2402.13717, 2024

  176. [184]

    Tokens-to-token vit: Training vision transformers from scratch on imagenet

    Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. InProceedings of the IEEE/CVF international conference on computer vision, pages 55...

  177. [185]

    Combin- ing large language models and crowdsourcing for hybrid human-ai misinformation detection

    Xia Zeng, David La Barbera, Kevin Roitero, Arkaitz Zubiaga, and Stefano Mizzaro. Combin- ing large language models and crowdsourcing for hybrid human-ai misinformation detection. InProceedings of the 47th International ACM SIGIR Conference on Research and Develop- ment in Info...

  178. [186]

    Token-level direct preference optimization.arXiv preprint arXiv:2404.11999, 2024

    Yongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang, Haifeng Zhang, and Jun Wang. Token-level direct preference optimization.arXiv preprint arXiv:2404.11999, 2024

  179. [187]

    Ex- ploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124, 2023

    Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. Ex- ploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124, 2023

  180. [188]

    Mutual theory of mind in human-ai collab- oration: An empirical study with llm-driven ai agents in a real-time shared workspace task

    Shao Zhang, Xihuai Wang, Wenhao Zhang, Yongshan Chen, Landi Gao, Dakuo Wang, Weinan Zhang, Xinbing Wang, and Ying Wen. Mutual theory of mind in human-ai collab- oration: An empirical study with llm-driven ai agents in a real-time shared workspace task. arXiv preprint arXiv:240...

  181. [189]

    Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37:132208–132237, 2024

    Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Arik. Chain of agents: Large language models collaborating on long-context tasks.Advances in Neural Information Processing Systems, 37:132208–132237, 2024

  182. [190]

    Competeai: Understanding the competition dynamics in large language model-based agents

    Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, and Xing Xie. Competeai: Understanding the competition dynamics in large language model-based agents. arXiv preprint arXiv:2310.17512, 2023

  183. [191]

    Aligning Large Language Mod- els for Faithful Integrity Against Opposing Argument, January 2025

    Yong Zhao, Yang Deng, See-Kiong Ng, and Tat-Seng Chua. Aligning Large Language Mod- els for Faithful Integrity Against Opposing Argument, January 2025

  184. [192]

    Does training with synthetic data truly protect privacy? InThe Thirteenth International Conference on Learning Representations, 2025

    Yunpeng Zhao and Jie Zhang. Does training with synthetic data truly protect privacy? InThe Thirteenth International Conference on Learning Representations, 2025

  185. [193]

    Cheating auto- matic LLM benchmarks: Null models achieve high win rates

    Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. Cheating auto- matic LLM benchmarks: Null models achieve high win rates. InThe Thirteenth International Conference on Learning Representations, 2025. 41

  186. [194]

    Chatgpt research group for optimizing the crystallinity of mofs and cofs.ACS Central Science, 9(11):2161–2170, 2023

    Zhiling Zheng, Oufan Zhang, Ha L Nguyen, Nakul Rampal, Ali H Alawadhi, Zichao Rong, Teresa Head-Gordon, Christian Borgs, Jennifer T Chayes, and Omar M Yaghi. Chatgpt research group for optimizing the crystallinity of mofs and cofs.ACS Central Science, 9(11):2161–2170, 2023

  187. [195]

    Larger and more instructable language models become less reli- able.Nature, 634(8032):61–68, 2024

    Lexin Zhou, Wout Schellaert, Fernando Mart ´ınez-Plumed, Yael Moros-Daval, C `esar Ferri, and Jos´e Hern´andez-Orallo. Larger and more instructable language models become less reli- able.Nature, 634(8032):61–68, 2024. 42

  188. [2024]

    Association for Computational Linguistics

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.