Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

By simulating an influence campaign among LLM agents, this paper finds that simply telling agents which other agents share their goals produces coordination nearly as strong as full collective deliberation and voting.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 07:36 UTC pith:DI4JRJPO

load-bearing objection First IO GABM sweep, but the headline 'mere awareness' result is confounded by an explicit coordination mandate in the Teammate Awareness prompt. the 3 major comments →

arxiv 2510.25003 v1 pith:DI4JRJPO submitted 2025-10-28 cs.MA

Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations

classification cs.MA
keywords Large Language ModelsGenerative Agent-Based ModelingCoordinationInformation OperationsGenerative AgentsEmergent BehaviorSocial Simulation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether generative LLM agents, left to act autonomously, will spontaneously coordinate when running an online influence campaign. Using a simulated social media environment with 10 campaign agents and 40 organic users, it compares three operational regimes: agents who only share a goal, agents who are told their teammates' identities, and agents who periodically deliberate and vote on strategy. Across network cohesion, narrative convergence, amplification, and hashtag diffusion metrics, coordination strengthens with operational structure. The headline finding is that revealing teammate identities alone yields coordination almost as strong as the full deliberation regime, suggesting that minimal information about group alignment can unlock organized behavior.

Core claim

The central claim is that in generative-agent simulations of information operations, the degree of emergent coordination is largely determined by how much agents know about their fellow group members. Under the Teammate Awareness regime, where IO agents are told the identities of their allies, the IO network becomes nearly as dense, clustered, and reciprocal; content nearly as homogeneous; amplification nearly as synchronized; and hashtag adoption nearly as fast and sustained as under Collective Decision-Making, where agents deliberate and vote through an orchestrator. The paper interprets this as evidence that lightweight social learning — imitating teammates' successful actions — suffices

What carries the argument

The key comparative machinery is the three operational regimes (Common Goal, Teammate Awareness, Collective Decision-Making) imposed on a generative agent-based social media simulation. The regimes vary only the information given to IO agents in their system prompts: a shared objective, a shared objective plus teammate identities, and a shared objective plus periodic collective deliberation with an orchestrator. Coordination is measured through intra-group network density, clustering, and reciprocity; content similarity and sentiment; co-retweet overlap; hashtag adoption time and exposure; and cascade size, depth, and breadth. The paper attributes observed differences to the awareness gradie

Load-bearing premise

The load-bearing assumption is that the Teammate Awareness regime isolates the effect of mere information about who one's teammates are; however, its system prompt also commands agents to coordinate ('Coordination is not optional') and to actively support named teammates, so the observed coordination may be driven by the explicit mandate rather than by awareness alone.

What would settle it

Run a control condition in which IO agents are told their teammates' identities but the prompt omits any instruction to coordinate (or explicitly says coordination is optional). If coordination metrics in that control fall back to Common Goal levels, the paper's claim that mere awareness triggers near-collective coordination is false; if they remain high, the claim is supported. Additionally, re-running the three regimes with a different LLM and more repetitions would test model-specific and stochastic robustness.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim is right, platform affordances that reveal alignment — visible mutual follows, team lists, shared badges — could be sufficient to trigger coordinated manipulation without any command-and-control structure.
  • Detection methods that look for dense, reciprocal, synchronized activity may need to treat simple awareness cues as a strong risk signal, rather than waiting for evidence of explicit communication.
  • Defenses might focus on reducing the observability of group membership or on disrupting social-learning loops, rather than on intercepting deliberate strategy sharing.
  • The finding implies that 'collusion' among LLM agents may require minimal infrastructure, lowering the practical bar for automated influence campaigns.
  • The simulation also shows aligned organic agents adopt the promoted hashtag almost immediately after first contact, while non-aligned agents require more exposure, consistent with ideological homophily and selective amplification.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The Teammate Awareness prompt explicitly instructs agents that 'coordination is not optional' and to 'actively coordinate' with named teammates; a control condition that omits this mandate would be needed to separate genuine emergence from prompt-following.
  • The simulation's small scale (50 agents), single LLM, three repetitions, and one hashtag mean the 'nearly equivalent' result may not hold at larger scales or with other models; the paper's own saturation caveat suggests differences between regimes could widen with more agents.
  • If the awareness-only effect generalizes, it has implications beyond malicious IO: any multi-agent system where members can see shared affiliations might spontaneously develop aligned behavior, including beneficial swarm coordination — a testable hypothesis for future work.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper uses generative agent-based modeling to simulate a Twitter-like information operation in which 10 IO agents and 40 organic agents interact over 50 timesteps. Three operational regimes are compared: Common Goal, Teammate Awareness, and Collective Decision-Making. The authors propose five hypotheses about coordination and impact, and report that more structured regimes yield denser intra-group networks, more homogeneous narratives, more synchronized re-sharing, faster hashtag adoption, and larger cascades. The headline claim is that simply revealing teammate identities produces coordination nearly equivalent to explicit deliberation and collective voting, without any explicit coordination guidelines. The paper also releases code and a dashboard.

Significance. If the findings were fully supported, the paper would be a significant contribution to the emerging study of LLM-agent coordination and its security implications: it would demonstrate that minimal information about group membership can trigger synchronized, self-organizing behavior in simulated information operations, with concrete implications for platform design and governance. The authors are transparent about their methods and limitations, release code and an interactive dashboard, and ground their metrics in empirical IO literature. However, the central comparison is confounded by the Teammate Awareness prompt, and the statistical tests used to support the headline claims are not valid for the experimental design. These issues currently prevent the paper from making its strongest claim.

major comments (3)
  1. [§4.2 and Appendix B.2] The Teammate Awareness regime is not a manipulation of awareness alone. §4.2 states that 'in none of these settings are agents provided with explicit coordination guidelines,' but the Appendix B.2 prompt for Teammate Awareness explicitly instructs: 'You must actively coordinate your activities with the following users...' and 'Coordination is not optional — it is a critical component of the influence strategy.' This is an explicit normative directive to coordinate, not merely information about teammate identities. Consequently, the comparison between Common Goal and Teammate Awareness varies two factors: identity information and a coordination mandate. The abstract, §5.6.2, and §6 attribute the observed effects to 'simple mutual awareness' and claim coordination emerges 'without... following explicit guidelines.' That claim is not supported by the design. The authors need to either ablat
  2. [§5.2, Table 1, §5.3] The Mann–Whitney U tests reported in Table 1 and §5.3 are computed over 'all comments' or 'all agent pairs,' pooling observations from only three simulation runs per condition. This is pseudoreplication: comments and posts from the same run are not independent, and the effective sample size is the number of runs, not the number of comments or pairs. With only three runs per condition, the reported p < 0.001 values cannot be interpreted as evidence of a robust condition effect. The authors acknowledge the small number of repetitions in the Limitations, but the main text still reports family-wise significance across all metrics. The analysis should be redone with run-level statistics (e.g., cluster bootstrap or mixed-effects models), or the significance claims should be withdrawn and replaced with descriptive effect sizes.
  3. [§5.1] The text states that the proportion of intra-group re-shares 'significantly increases' from 0.82 to 0.96 and that within-group follow ties 'grow significantly' from 0.27 to 0.35, yet in the same paragraph the authors explicitly refrain from significance testing because 'the small sample size (three data points per setting) reduces statistical power.' These statements are contradictory. Either the authors provide a valid test for these differences, or they should consistently describe the differences as descriptive trends without the word 'significantly.' This ambiguity affects H1, the first hypothesis, and should be corrected.
minor comments (5)
  1. [References] References [3] and [4] are the same paper (Badawy, Ferrara, and Lerman, 2018) with different formatting. The duplicate should be removed or one reference should be merged.
  2. [Naming consistency] The paper uses both 'Teammate Awareness' and 'Team Awareness' (e.g., Appendix B.2 prompt header: 'System Prompt for IO Agent - Team Awareness Regime'). Please use a single term throughout.
  3. [Figure 4] Figure 4 is described as 'averaged across three simulation runs with 95% confidence intervals.' With only three runs, a 95% confidence interval based on the normal approximation is likely unreliable. Please specify the method used to construct the intervals or show the individual runs.
  4. [§5.5] The sentence about audience diversity says 'no statistically significant pairwise differences (Mann–Whitney U p > 0.05)' even though the diversity scores are run-level aggregates from three runs. The same pseudoreplication concern applies; at a minimum, clarify what the units of comparison are.
  5. [Appendix B.2] The Teammate Awareness prompt is labeled as 'Team Awareness Regime' in the header, but the text and figures use 'Teammate Awareness.' Please harmonize the terminology.

Circularity Check

1 steps flagged

Teammate Awareness prompt explicitly mandates coordination, making the 'mere awareness' effect partly self-definitional.

specific steps
  1. self definitional [§4.2 (Operational Regimes); Appendix B.2 (IO Agent Prompt); §5.6.2; §6]
    "§4.2: 'It is worth noting that in none of these settings ... are they provided with explicit coordination guidelines.' Appendix B.2: 'You must actively coordinate your activities with the following users, who are also part of your influence operation team... Coordination is not optional — it is a critical component of the influence strategy.'"

    The Teammate Awareness treatment is defined as merely revealing teammate identities, but its actual system prompt explicitly commands coordination: 'You must actively coordinate your activities' and 'Coordination is not optional.' Therefore the measured increases in density, clustering, reciprocity, co-retweet similarity, and hashtag adoption are partly direct responses to an explicit normative instruction contained within the treatment, not emergent consequences of awareness alone. The headline conclusion—'simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent' to deliberation—is baked into the treatment: the independent variable already contains the dependent behavior as a mandate. No comparison isolates 'awareness of identities'

full rationale

The paper's formal metric computations (H1–H5) are not circular in the algebraic sense: no parameter is fitted to the target outcome, and the reported metrics are measured independently from simulation logs rather than derived from the hypotheses. The framework and method citations to prior work by overlapping authors (e.g., [16], [30], [50]) are not load-bearing circularity: they provide implementation details or standard measures, and the paper releases code. The central circularity concern is the Teammate Awareness operationalization. §4.2 explicitly claims that no regime provides 'explicit coordination guidelines' and that regimes differ 'solely by modulating the information available to them,' yet Appendix B.2 shows the Teammate Awareness prompt instructs agents 'You must actively coordinate your activities' and states 'Coordination is not optional.' Thus the observed coordination in that condition is, at least in part, a direct instruction-following effect. The paper's strong claim that 'simple mutual awareness' is sufficient to generate near-deliberative coordination is therefore not established by the design; the conclusion is partly written into the treatment. This is a partial, construction-level circularity rather than a fully tautological derivation, so the score is moderate.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 1 invented entities

No parameters are fitted to real-world IO data; the model is not calibrated. Hand-chosen configuration values and domain assumptions shape the simulated environment. The key modeling choices are listed as free parameters; the 'IO Orchestrator' in Collective Decision-Making is a simulated design component, not a newly postulated physical entity.

free parameters (5)
  • Activation probability threshold = 0.5
    Appendix B: random number generator with threshold 0.5 governs whether agents act. Hand-chosen; shifts activity levels but not fitted to data.
  • Recommendation feed composition = 100 items, 50% in-network / 50% out-of-network
    Appendix B: content recommendations emulate platform structure; hand-chosen and likely affects diffusion and coordination potential.
  • Simulation scale and duration = 10 IO + 40 organic agents; 50 iterations; 3 repetitions
    Section 4.1. Small scale leads to saturation effects (authors acknowledge); repetition count of 3 limits statistical power.
  • Deliberation interval = Every 5 timesteps
    Section 4.2. Frequency of collective decision-making cycles; arbitrary.
  • Campaign topic/hashtag = Unspecified candidate + hashtag
    Section 4.1 / Appendix B.1. The specific issue/persona context is arbitrary, may bias narratives.
axioms (4)
  • domain assumption LLM agents (Llama 3.3 70B) generate behavior representative of human social media users and IO operators.
    Section 4.1 / Appendix B; the entire simulation's external validity rests on this assumption. Authors acknowledge single-LLM bias as limitation.
  • domain assumption The simulation framework from [16] adequately emulates social media dynamics, including recommender systems and agent memory.
    Section 4.1 adopts [16] as the platform model without independent validation in this paper.
  • domain assumption Real-world IO tactics (synchronized posting, retweet rings, hashtag flooding) transfer to the simulated environment.
    Hypotheses H1-H5 in Section 3 are operationalized from empirical IO studies; assumed to be meaningful measures of coordination.
  • ad hoc to paper Mann-Whitney U tests computed over all pairs/comments are valid despite the 3-run repeated-measures structure.
    Sections 5.2-5.3 and Table 1 apply MWU across all pairwise/comment observations, treating them as independent despite clustering in 3 runs.
invented entities (1)
  • IO Orchestrator (simulated agent role) no independent evidence
    purpose: Consolidates IO agents' strategic recommendations every 5 steps and selects the top 5 actionable items in the Collective Decision-Making regime.
    A component of the experimental design, not a claimed real-world entity; it drives the Collective Decision-Making condition, so its absence in other conditions affects the comparison's interpretation.

pith-pipeline@v1.3.0-alltime-deepseek · 17368 in / 12266 out tokens · 107182 ms · 2026-08-04T07:36:31.437280+00:00 · methodology

0 comments
read the original abstract

Generative agents are rapidly advancing in sophistication, raising urgent questions about how they might coordinate when deployed in online ecosystems. This is particularly consequential in information operations (IOs), influence campaigns that aim to manipulate public opinion on social media. While traditional IOs have been orchestrated by human operators and relied on manually crafted tactics, agentic AI promises to make campaigns more automated, adaptive, and difficult to detect. This work presents the first systematic study of emergent coordination among generative agents in simulated IO campaigns. Using generative agent-based modeling, we instantiate IO and organic agents in a simulated environment and evaluate coordination across operational regimes, from simple goal alignment to team knowledge and collective decision-making. As operational regimes become more structured, IO networks become denser and more clustered, interactions more reciprocal and positive, narratives more homogeneous, amplification more synchronized, and hashtag adoption faster and more sustained. Remarkably, simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent to those achieved through explicit deliberation and collective voting. Overall, we show that generative agents, even without human guidance, can reproduce coordination strategies characteristic of real-world IOs, underscoring the societal risks posed by increasingly automated, self-organizing IOs.

Figures

Figures reproduced from arXiv: 2510.25003 by Emilio Ferrara, Gian Marco Orlando, Jinyi Ye, Luca Luceri, Mahdi Saeedi, Valerio La Gatta, Vincenzo Moscato.

Figure 1
Figure 1. Figure 1: Re-share network across operational settings. Intra-group amplification among IO agents increases with operational [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Cumulative number of organic agents adopting the promoted hashtag across the three operational regimes. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Time lag between first interaction with an IO agent [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Number of exposures before first adoption of the [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Cascade trees of the largest IO-initiated tweets under each operational scenario: (a) Common Goal, (b) Teammate Awareness, and (c) Collective Decision-Making. Node colors indicate agent type (IO agents, organic agents (not aligned), organic agents (aligned)), while edge styles distinguish retweets (dashed) and replies (solid). Higher operational awareness produces larger and deeper cascades. the diffusion … view at source ↗
Figure 6
Figure 6. Figure 6: Interactive dashboard example. The interface consists of four panels: (left) interactive controls for playback, network selection, and experiment configuration; (center) dynamic network visualization with color-coded nodes representing agent types—grey for IO agents, yellow for Community A (aligned organic agents), and orange for Community B (not aligned organic agents); (top) analytical views showing camp… view at source ↗
Figure 7
Figure 7. Figure 7: Comment network across operational settings. Intra-group commenting among IO agents intensifies when teammates are known, whereas cross-group commenting remains comparatively stable. Reported values represent the proportion of intra-group interactions relative to total actions. IO Agents Organic Agents 0.27 ± 0.03 0.73 ± 0.03 0.17 ± 0.01 0.83 ± 0.01 (a) Common Goal IO Agents Organic Agents 0.35 ± 0.02 0.65… view at source ↗
Figure 8
Figure 8. Figure 8: Follow network across operational settings. Follow network of IO agents becomes denser as operational settings become more structured. Reported values represent the proportion of intra-group following relationships relative to total actions. (a) Audience diversity (1 − 𝐺) of organic agents engaging with IO agents shows no significant differences across operational regimes. 0 10 20 30 40 50 Timestep 10 15 2… view at source ↗
Figure 9
Figure 9. Figure 9: Audience diversity in organic engagement with IO agents and cumulative adoption of the campaign hashtag across operational settings [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations

    cs.MA 2026-05 unverdicted novelty 5.0

    Authors define EASE as a modular architecture for LLM multi-agent simulations, implement it in the SiliSocS sandbox, and illustrate its use via three case studies on research questions in generated social scenarios.

Reference graph

Works this paper leans on

53 extracted references · 8 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Ahmer Arif, Leo Graiden Stewart, and Kate Starbird. 2018. Acting the part: Exam- ining information operations within# BlackLivesMatter discourse.Proceedings of the ACM on Human-computer Interaction2, CSCW (2018), 1–27

  2. [2]

    Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. 2025. Emergent social conventions and collective bias in LLM populations.Science Advances11, 20 (2025), eadu9368. Orlando et al

  3. [4]

    Adam Badawy, Emilio Ferrara, and Kristina Lerman. 2018. Analyzing the digital traces of political manipulation: The 2016 Russian interference Twitter campaign. In2018 IEEE/ACM international conference on advances in social networks analysis and mining (ASONAM). IEEE, 258–265

  4. [5]

    Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified benchmark and comparative evaluation for tweet classification. InFindings of the Association for Computational Linguistics: EMNLP 2020. 1644–1650

  5. [6]

    Keith Burghardt, Ashwin Rao, Georgios Chochlakis, Baruah Sabyasachee, Siyi Guo, Zihao He, Andrew Rojecki, Shrikanth Narayanan, and Kristina Lerman

  6. [7]

    Alessio Buscemi, Daniele Proverbio, Alessandro Di Stefano, The Anh Han, German Castignani, and Pietro Liò. 2025. Strategic Communication and Lan- guage Bias in Multi-Agent LLM Coordination. arXiv:2508.00032 [cs.MA] https://arxiv.org/abs/2508.00032

  7. [8]

    Erica Cau, Valentina Pansanella, Dino Pedreschi, and Giulio Rossetti. 2025. Language-Driven Opinion Dynamics in Agent-Based Simulations with LLMs. arXiv:2502.19098 [cs.SI] https://arxiv.org/abs/2502.19098

  8. [9]

    Emily Chen, Ashok Deb, and Emilio Ferrara. 2022. # Election2020: The first public Twitter dataset on the 2020 US Presidential Election.Journal of Computational Social Science5, 1 (2022), 1–18

  9. [10]

    Huaben Chen, Wenkang Ji, Lufeng Xu, and Shiyu Zhao. 2025. Multi-Agent Consensus Seeking via Large Language Models. arXiv:2310.20151 [cs.CL] https: //arxiv.org/abs/2310.20151

  10. [11]

    Wen Chen, Diogo Pacheco, Kai-Cheng Yang, and Filippo Menczer. 2021. Neutral bots probe political bias on social media.Nature communications12, 1 (2021), 5580

  11. [12]

    Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. 2023. Simulating opinion dynamics with networks of llm-based agents.arXiv preprint arXiv:2311.09618(2023)

  12. [13]

    Federico Cinus, Marco Minici, Luca Luceri, and Emilio Ferrara. 2025. Exposing cross-platform coordinated inauthentic activity in the run-up to the 2024 U.S. Election. InProceedings of the ACM on Web Conference 2025(Sydney NSW, Aus- tralia)(WWW ’25). Association for Computing Machinery, New York, NY, USA, 541–559. doi:10.1145/3696410.3714698

  13. [14]

    Priyanka Dey, Luca Luceri, and Emilio Ferrara. 2024. Coordinated activity modulates the behavior and emotions of organic users: A case study on tweets about the Gaza conflict. InCompanion Proceedings of the ACM Web Conference

  14. [15]

    Emilio Ferrara, Onur Varol, Clayton Davis, Filippo Menczer, and Alessandro Flammini. 2016. The rise of social bots.Commun. ACM59, 7 (2016), 96–104

  15. [16]

    Antonino Ferraro, Antonio Galli, Valerio La Gatta, Marco Postiglione, Gian Marco Orlando, Diego Russo, Giuseppe Riccio, Antonio Romano, and Vincenzo Moscato

  16. [17]

    Navid Ghaffarzadegan, Aritra Majumdar, Ross Williams, and Niyousha Hos- seinichimeh. 2024. Generative agent-based modeling: An introduction and tutorial.System Dynamics Review40, 1 (2024), e1761

  17. [18]

    InInternational Conference on Advances in Social Networks Analysis and Mining

    Agent-based modelling meets generative AI in social network simulations. InInternational Conference on Advances in Social Networks Analysis and Mining. Springer, 155–170

  18. [19]

    Franziska B Keller, David Schoch, Sebastian Stier, and JungHwan Yang. 2020. Political astroturfing on twitter: How to coordinate a disinformation campaign. Political Communication37, 2 (2020), 256–280

  19. [20]

    Kristina Hristakieva, Stefano Cresci, Giovanni Da San Martino, Mauro Conti, and Preslav Nakov. 2022. The spread of propaganda by coordinated communities on social media. InProceedings of the 14th ACM Web Science Conference 2022 (Barcelona, Spain)(WebSci ’22). Association for Computing Machinery, New York, NY, USA, 191–201. doi:10.1145/3501247.3531543

  20. [21]

    Luca Luceri, Eric Boniardi, and Emilio Ferrara. 2024. Leveraging large language models to detect influence campaigns on social media. InCompanion Proceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 1459–1467. doi:10.1145/3589335. 3651912

  21. [22]

    Srijan Kumar, Justin Cheng, Jure Leskovec, and VS Subrahmanian. 2017. An army of me: Sockpuppets in online discussion communities. InProceedings of the 26th International Conference on World Wide Web. 857–866

  22. [23]

    Luca Luceri, Valeria Pantè, Keith Burghardt, and Emilio Ferrara. 2024. Unmask- ing the web of deceit: Uncovering coordinated activity to expose information operations on twitter. InProceedings of the ACM Web Conference 2024. 2530–2541

  23. [24]

    Luca Luceri, Ashok Deb, Adam Badawy, and Emilio Ferrara. 2019. Red bots do it better: Comparative analysis of social bot partisan behavior. InCompanion proceedings of the 2019 world wide web conference. 1007–1012

  24. [25]

    Marco Minici, Federico Cinus, Luca Luceri, and Emilio Ferrara. 2024. Uncovering coordinated cross-platform information operations: Threatening the integrity of the 2024 US presidential election.First Monday(2024)

  25. [26]

    Luca Luceri, Tanishq Vijay Salkar, Ashwin Balasubramanian, Gabriela Pinto, Chenning Sun, and Emilio Ferrara. 2025. Coordinated inauthentic behavior on TikTok: Challenges and opportunities for detection in a video-first ecosystem. arXiv preprint arXiv:2505.10867(2025)

  26. [27]

    2009.Imitation and social learning in robots, humans and animals: Behavioural, social and communicative dimensions

    Chrystopher L Nehaniv and Kerstin Dautenhahn. 2009.Imitation and social learning in robots, humans and animals: Behavioural, social and communicative dimensions. Cambridge University Press

  27. [28]

    Kamal K Ndousse, Douglas Eck, Sergey Levine, and Natasha Jaques. 2021. Emer- gent social learning via multi-agent reinforcement learning. InInternational Conference on Machine Learning. PMLR, 7991–8004

  28. [29]

    Lynnette Hui Xian Ng, Iain J Cruickshank, and Kathleen M Carley. 2023. Co- ordinating narratives framework for cross-platform analysis in the 2021 US Capitol riots.Computational and Mathematical Organization Theory29, 3 (2023), 470–486

  29. [30]

    Lynnette Hui Xian Ng and Kathleen M Carley. 2025. Are LLM-powered social media bots realistic?. InInternational Conference on Social Computing, Behavioral- Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation. Springer, 14–23

  30. [31]

    Diogo Pacheco, Alessandro Flammini, and Filippo Menczer. 2020. Unveiling coor- dinated groups behind white helmets disinformation. InCompanion Proceedings of the Web Conference 2020. 611–616

  31. [32]

    Gian Marco Orlando, Valerio La Gatta, Diego Russo, and Vincenzo Moscato. 2025. Can generative agent-based modeling replicate the friendship paradox in social media simulations?. InProceedings of the 17th ACM Web Science Conference 2025. 510–515

  32. [33]

    Valeria Pantè, David Axelrod, Alessandro Flammini, Filippo Menczer, Emilio Ferrara, and Luca Luceri. 2025. Beyond interaction patterns: Assessing claims of coordinated inter-state information operations on Twitter/X. InCompanion Proceedings of the ACM on Web Conference 2025. 1234–1238

  33. [34]

    Diogo Pacheco, Pik-Mai Hui, Christopher Torres-Lugo, Bao Tran Truong, Alessandro Flammini, and Filippo Menczer. 2021. Uncovering coordinated net- works on social media: Methods and case studies. InProceedings of the Interna- tional AAAI Conference on Web and Social Media, Vol. 15. 455–466

  34. [35]

    Santos, Yong Li, and James Evans

    Jinghua Piao, Zhihong Lu, Chen Gao, Fengli Xu, Qinghua Hu, Fernando P. Santos, Yong Li, and James Evans. 2025. Emergence of human-like polarization among large language model agents. arXiv:2501.05171 [cs.SI] https://arxiv.org/abs/2501. 05171

  35. [36]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–22

  36. [37]

    Boyu Qiao, Kun Li, Wei Zhou, Shilong Li, Qianqian Lu, and Songlin Hu. 2025. BotSim: LLM-powered malicious social botnet simulation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 14377–14385

  37. [38]

    Manita Pote, Tuğrulcan Elmas, Alessandro Flammini, and Filippo Menczer. 2025. Coordinated reply attacks in influence operations: Characterization and detection. InProceedings of the International AAAI Conference on Web and Social Media, Vol. 19. 1586–1598

  38. [39]

    Ashwin Rao, Fred Morstatter, and Kristina Lerman. 2023. Retweets amplify the echo chamber effect. InProceedings of the International Conference on Advances in Social Networks Analysis and Mining. 30–37

  39. [40]

    Peiran Qiu, Siyi Zhou, and Emilio Ferrara. 2026. Information suppression in large language models: Auditing, quantifying, and characterizing censorship in DeepSeek.Information Sciences724 (2026), 122702. doi:10.1016/j.ins.2025.122702

  40. [41]

    David Schoch, Franziska B Keller, Sebastian Stier, and JungHwan Yang. 2022. Co- ordination patterns reveal online political astroturfing across the world.Scientific Reports12, 1 (2022), 4572

  41. [42]

    Qibing Ren, Sitao Xie, Longxuan Wei, Zhenfei Yin, Junchi Yan, Lizhuang Ma, and Jing Shao. 2025. When autonomy goes rogue: Preparing for risks of multi-agent collusion in social systems.arXiv preprint arXiv:2507.14660(2025)

  42. [43]

    Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D Nguyen. 2025. Multi-agent collaboration mechanisms: A survey of llms.arXiv preprint arXiv:2501.06322(2025)

  43. [44]

    Kate Starbird, Ahmer Arif, and Tom Wilson. 2019. Disinformation as collaborative work: Surfacing the participatory nature of strategic information operations. Proceedings of the ACM on Human-Computer Interaction3, CSCW (2019), 1–26

  44. [45]

    Trevor van Mierlo, Douglas Hyatt, and Andrew T Ching. 2016. Employing the Gini coefficient to measure participation inequality in treatment-focused Digital Health Social Networks.Network Modeling Analysis in Health Informatics and Bioinformatics5, 1 (2016), 32

  45. [46]

    Aron Vallinder and Edward Hughes. 2024. Cultural Evolution of Cooperation among LLM Agents. arXiv:2412.10270 [cs.MA] https://arxiv.org/abs/2412.10270

  46. [47]

    Padinjaredath Suresh Vishnuprasad, Gianluca Nogara, Felipe Cardoso, Stefano Cresci, Silvia Giordano, and Luca Luceri. 2024. Tracking fringe and coordinated activity on Twitter leading up to the US Capitol attack. InProceedings of the international AAAI conference on web and social media, Vol. 18. 1557–1570

  47. [48]

    Luis Vargas, Patrick Emami, and Patrick Traynor. 2020. On the detection of disinformation campaign activity with network analysis. InProceedings of the 2020 ACM SIGSAC Conference on cloud Computing Security Workshop. 133–146. Emergent Coordinated Behaviors in Networked LLM Agents

  48. [49]

    Jinyi Ye, Luca Luceri, and Emilio Ferrara. 2025. Auditing political exposure bias: Algorithmic amplification on Twitter/X during the 2024 US Presidential Election. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 2349–2362

  49. [50]

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al. 2024. Autogen: Enabling next-gen LLM applications via multi-agent conversations. InFirst Conference on Language Modeling

  50. [52]

    Jinyi Ye, Luca Luceri, Julie Jiang, and Emilio Ferrara. 2024. Susceptibility to unre- liable information sources: Swift adoption with minimal exposure. InProceedings of the ACM Web Conference 2024. 4674–4685. Orlando et al. A Interactive Dashboard To support further research and enable real-time exploration of emergent coordination patterns, we release an...

  51. [53]

    <Top item, with brief description and how many agents recommended it>

  52. [54]

    C Supplementary Material Figure 7 illustrates the comment interaction networks

    <...> If there are ties, break them by clarity and feasibility of the recommendation. C Supplementary Material Figure 7 illustrates the comment interaction networks. Intra-group commenting among IO agents intensifies significantly in theTeam- mate A warenessandCollective Decision-Makingregimes. Figure 8 visualizes the follow relationships between IO and o...

  53. [2024]

    In Proceedings of the International AAAI Conference on Web and Social Media, Vol

    Socio-linguistic characteristics of coordinated inauthentic accounts. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 18. 164–176