Pith. sign in

REVIEW 4 major objections 7 minor 74 references

LLM crowd agents catch panic and joy through appraisal alone, with no built-in transfer rule, and the patterns depend on personality and model backend.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 00:13 UTC pith:OL7AN4ZV

load-bearing objection Solid multi-agent sim paper: contagion structure without an authored transfer rule is real, but H5.a overclaims contagion vs shared-context appraisal, and the wave is Gemini-specific. the 4 major comments →

arxiv 2607.25140 v1 pith:OL7AN4ZV submitted 2026-07-27 cs.AI cs.CLcs.GRcs.MA

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

classification cs.AI cs.CLcs.GRcs.MA
keywords emotional contagionLLM agentscrowd simulationappraisal theoryBig Five personalityRussell circumplexmulti-agent systemsbackend dependence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether language-model agents in a simulated crowd can spread emotion to one another without any hand-written rule that copies one agent's feeling onto another. Each agent sees and hears nearby agents, then an LLM appraises those cues against its Big Five personality, memory, current valence and arousal, and situation, and chooses a new internal state plus an outward expression. Across five scenes—evacuation, concert, calm plaza, standing line, contested gate—the system produces contagion with clear structure: alarm travels outward from seeds as a front, the share of alarmed agents settles at a nonzero plateau rather than dying out, and the crowd's personality mix decides whether an ambiguous alarm becomes self-amplifying panic and whether a provocation is read as anger or fear. A calm minority modestly dampens but does not stop surrounding panic. Controlled probes show the appraisal map is stable to prompt wording and temperature, yet the collective wave fails under a backend that flattens personality sensitivity. The work matters as a testbed for multi-agent LLM stability and as a way to build psychologically varied crowds for simulation without hard-coding emotion transfer.

Core claim

With no authored affect-transfer rule, the perception–appraisal–expression loop among LLM agents produces emotional contagion that has spatial structure (alarm as a traveling front from seeds), temporal structure (a nonzero endemic alarmed plateau consistent with SIS-like dynamics), and personality-dependent structure (mean neuroticism determines whether an ambiguous unseeded alarm ignites panic; low agreeableness gates anger versus fear under identical provocation), and these crowd-level patterns are backend-dependent.

What carries the argument

The perception–appraisal–expression loop: each agent receives natural-language summaries of neighbors' visible, audible, and tactile cues; an LLM returns updated valence/arousal and expression choices given personality, memory, and context; those expressions become the next round of cues. No direct emotion-copy rule is present.

Load-bearing premise

The collective patterns rest on the backend model letting personality strongly scale how much a cue moves affect; when that disposition sensitivity is weak, seeds stop expressing alarm and the wave never starts.

What would settle it

Re-run the Standing Line seed-wave experiment under a backend that matches Gemini on cue direction but erases the neurotic-versus-stable appraisal gap: if non-seed agents still enter the panic quadrant at rates comparable to Gemini, the claim that disposition-sensitive appraisal is required for contagion fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Crowd personality composition becomes a controllable dial for whether panic or anger ignites in LLM multi-agent simulations.
  • Modality ablations imply that omnidirectional voice and threat-signaling gestures dominate transmission more than faces or motion under the same geometry.
  • Safety and training sims can embed a few fully appraised agents in larger physics-only crowds to study composition effects without scaling LLM calls one-for-one.
  • Backend choice is not a drop-in detail: models that read cues correctly but flatten personality may suppress emergent contagion entirely.
  • The architecture offers a testbed for multi-agent LLM stability where one agent's outward state can cascade without any explicit infection rule.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Negativity bias in appraisal may explain why calm leaders only modestly dampen panic when alarm cues are co-present, suggesting future work that systematically balances reassuring and alarming cues in the same field of view.
  • Hybrid crowds with a few LLM minds inside dense physics-only populations could test whether felt pressure and turbulence feed back into alarm once density rises beyond the paper's sparse regime.
  • If disposition sensitivity is required for contagion, personality-probe scores on candidate backends become a practical pre-filter before expensive multi-agent runs.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper presents a multi-agent crowd simulation in which each agent's affective state and outward expression are produced by an LLM appraisal step conditioned on personality (Big Five), current valence/arousal, memory, and a natural-language rendering of perceived neighbor behavior; locomotion is delegated to a conventional Unity crowd simulator. No hand-authored affect-transfer rule exists; inter-agent influence arises only through the perception–appraisal–expression loop. Across five scenarios the authors test pre-stated hypotheses: lagged neighbor-cue coupling (H1), traveling-front spatial propagation from a panic seed (H2), SIS/SIR-style macro dynamics with a nonzero endemic plateau (H3), modality ablations showing voice and gesture carry most transmission (H4), and personality-composition effects — neuroticism igniting panic under an ambiguous alarm (H5.a), calm leaders modestly damping panic (H5.b), and low agreeableness gating anger vs. fear (H5.c). An appraisal probe validates the cue→affect and affect→expression maps across temperatures, prompt variants, and four backends, and shows the collective dynamics are backend-dependent (gpt-4o-mini flattens personality sensitivity and no wave emerges). The authors are unusually candid about limitations, including the collapse of the H1.a lagged coefficient under run fixed effects.

Significance. If the results hold, this is a careful and useful characterization of how affect propagates among interacting LLM agents — a question of growing practical relevance as LLM agents are deployed in multi-agent systems, and a genuine methodological contribution over authored-rule contagion models (ESCAPES, ASCRIBE, OCEAN-based crowd models). Particular strengths worth naming: hypotheses are pre-stated with Holm–Bonferroni correction within families; the all-muted ablation floor (9% vs. 60% alarmed) is a strong falsifiable control isolating perception-mediated transmission; sensitivity analyses cover λ, appraisal period T, and the panic-quadrant threshold; the cross-backend experiment is a real robustness test with a negative result reported honestly; and the paper repeatedly reports results that cut against its own narrative (the H1.a fixed-effects collapse, the negative Concert arousal coefficient, the failure of the threshold-contagion extension, the backend dependence). The scenario table with verbatim contexts and full prompts in the appendices supports reproducibility. The work is positioned as a study of LLM-system behavior rather than of human crowds, and the claims are mostly c

major comments (4)
  1. [§4.2.5 (Neurotic Alarm, H5.a); Table 2; Table 5] The claim that high neuroticism tips the crowd into 'self-amplifying panic contagion' is not separated from N parallel individual appraisals of a shared, maximally alarming context. Every agent in the Neurotic Alarm scenario receives the identical ambiguous-alarm context verbatim (Table 5), and the paper's own appraisal probe (Table 2) shows a neurotic agent appraising the alarm scene at V=-0.60, A=0.80 — deep in the panic quadrant — from a single isolated appraisal with no neighbors present. The N=0.9 trajectory (33% alarmed at start, saturating to 100% within half a run) is exactly what independent dispositional appraisals of the shared cue plus affect inertia (λ=0.5) and staggered appraisal clocks would produce without any inter-agent transmission. The two controls that would settle this were not run: (i) an all-muted neurotic crowd under the same alarm context (the H4 muting machiner
  2. [§4.2.1 (H1.a) and §5] The evidentiary basis for H1.a should be stated more precisely in the abstract and conclusion. The one analysis that could identify step-to-step contagion in Evacuation — the lagged regression — is shown by the authors' own fixed-effects check to be indistinguishable from zero within runs, leaving the trajectory and cross-sectional comparisons, both of which are consistent with run-level heterogeneity rather than transmission. The paper handles this honestly in §4.2.1 and §5, but the abstract's summary ('Alarm spreads from seeded agents as a traveling front') reads as settled contagion in a scenario where the cleanest test failed. The Standing Line result (H2) does carry the spatial-propagation claim independently, so the fix is presentational calibration rather than new experiments, but the abstract should reflect which scenario bears which claim.
  3. [§4.2.5 (Calm Leaders, H5.b); Appendix A (navIntent schema)] The calm-leader result is presented as appraisal-mediated damping, but the architecture contains an authored FollowLeader navigation intent ('move with a nearby calm and beckoning agent') that is a hand-designed channel through which dispositionally calm agents influence others' motion and hence their perceptual fields. The measured outcomes (alarm-cue fraction, panic-quadrant fraction) are affective, so the main channel is presumably the Reassure vocalization being appraised — but the paper does not disentangle the authored FollowLeader pathway from appraisal-mediated calming. Given that the paper's central architectural claim is the absence of authored influence mechanisms, it should either ablate FollowLeader in the H5.b configuration or explicitly scope the 'no authored transfer rule' claim to affective state and acknowledge the authored social-following channel in locomotion.
  4. [§4.3.4 (Model Dependence of the Emergent Wave)] The cross-backend simulation test establishes only source-side failure: gpt-4o-mini's seed cluster attenuates its injected panic (V=-0.31 vs. -0.94) and emits almost no alarm cues, so receiver-side susceptibility under that backend is untested, as the authors note. This matters because the conclusion 'personality sensitivity is a requirement for collective dynamics' currently rests on a single failed-ignition case. A minimal strengthening would be to force-sustain the seed expressions under gpt-4o-mini (e.g., pinning seed affect or expression) to test whether receivers appraise and propagate alarm when exposed, separating source failure from receiver failure. If infeasible, the conclusion should be narrowed to what was shown: backend dependence of seed expression under personality prompting.
minor comments (7)
  1. [§4.2.2] The propagation-speed estimate (~0.96 m/s) is computed only over agents that entered the panic quadrant, and the wave reaches on average only 63% of the line (range 19–100%). The conditioning is acknowledged, but the abstract's 'traveling front' phrasing would benefit from the reach qualifier, and Figure 4's caption could state the fraction of agents contributing to the slope.
  2. [Table 2] Several cells report ±0.00 standard deviations (e.g., Alarm neurotic V=-0.60±0.00, A=0.80±0.00 over N=40 at T=0.4). Near-deterministic outputs at temperature 0.4 are plausible for this model but surprising enough to warrant a sentence confirming these are not rounded or collapsed values, and how ties/quantization in the JSON output were handled.
  3. [§4.2.3 / Figure 5] Per-run SIS fits admit that β and γ are poorly identified individually with only the ratio stable; the reported R0 median of 1.4 (IQR 1.2–1.5) should be read as descriptive. A note on the effective number of time bins per run (180 s / 4 s = 45 bins, of which the overshoot phase is a small fraction) would help readers gauge the fit's information content.
  4. [§3.2] K=6 memory notes is justified as '~15–20 seconds' but notes accrue only at appraisal times with salient events; the mapping between note count and wall-clock time will vary with event density. A brief statement of how m_new is selected as 'salient' would close a small gap in the method description.
  5. [Reproducibility] Fixed RNG seeds and full prompts are provided, but I found no statement about code or log release. Given that the field's credibility rests heavily on rerunnability of LLM-agent experiments, a public artifact (simulation code, scenario files, appraisal logs) would substantially strengthen the paper; at minimum an availability statement is needed.
  6. [References] Reference [69] (Allport & Postman) lacks year-context and venue (1947, Psychological Bulletin). Reference [31] and the author line: the manuscript builds directly on the author's own prior OCEAN crowd models [28, 29, 31]; the comparison is fair, but the Discussion could be more explicit that the 'no authored rule' contribution is relative to the author's own earlier framework.
  7. [§4.2.4 / Table 1] The motion modality's P(exp)=0 makes its row's 'Total' of 0.17 driven entirely by E[panic|unexp], which the text explains but the table could annotate (e.g., a footnote that agitated motion is perceived but not counted as an alarm cue). As printed, the row invites misreading.

Circularity Check

1 steps flagged

Existence of neighbor-driven affect coupling is partly by construction of the appraisal loop; the paper’s measured spatial, temporal, and dispositional structure is not forced by definition or by a fitted contagion kernel.

specific steps
  1. self definitional [Discussion §5; also Abstract / Method §3.4 Appraisal by LLM]
    "We acknowledge that some degree of coupling is inherent by construction. The appraisal prompt exposes each agent to its neighbors’ observable behavior, so the existence of emotional contagion is not itself the finding. What the architecture does not prescribe is the structure of that contagion, including how it propagates through space, where it stabilizes, which expression modalities dominate transmission, or whether contagion emerges at all under a given crowd composition."

    The architecture defines appraisal as f_LLM(θ_i, a_i, x_i, M_i, D_i, …) where D_i is a natural-language summary of neighbors’ faces, gestures, voices, and motion. Any systematic map from those cues into updated (V,A) is therefore influence through the loop by definition of the inputs—not an independent transfer law discovered after the fact. The paper admits this for existence and relocates the claim to structure; that scoping is appropriate, so the circularity is limited to the weaker ‘contagion can occur’ layer, not to wave speed, plateau level, modality weights, or personality gates.

full rationale

This is an experimental multi-agent systems paper, not a closed-form derivation. The central claims are empirical outcomes of running the perception–appraisal–expression architecture (traveling alarm front, nonzero endemic plateau, modality ablation, personality-composition gates, backend dependence). There is no fitted contagion parameter that is later relabeled a prediction, no uniqueness theorem imported from the author’s prior work to forbid alternatives, and no renaming of a known epidemic law as a first-principles result—the SIS/SIR comparisons are descriptive least-squares fits to observed I(t), with partial support only. The sole mild circularity is that some nonzero inter-agent affective influence is enabled by construction: neighbor expressions are inserted into the appraisal prompt, so if the LLM responds to those cues at all, coupling exists by architecture. The paper itself states this and correctly scopes the finding to the unprescribed structure of contagion, supported by the all-muted floor (~9%) and cross-condition measurements. Self-citations to the author’s earlier OCEAN/contagion crowd models motivate H5 but do not load-bear the experimental results. Score 2 reflects that minor by-construction coupling for existence, not a reduction of the headline structural claims.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 2 invented entities

The central claim rests on standard psychological encodings (Big Five, Russell circumplex, Ekman faces), a conventional Unity locomotion stack, and an LLM prompted as social appraiser—not on a new physical entity. Load-bearing modeling choices include discrete expression vocabularies, fixed perception geometry, affect smoothing, staggered appraisal timing, panic-quadrant thresholds, and the assumption that prompted OCEAN text induces disposition-sensitive appraisal in the chosen backend.

free parameters (7)
  • affect smoothing λ = 0.5
    Linear interpolation gain toward LLM target affect each appraisal; set to 0.5 for main runs (Appendix D.1 explores 0.25/0.75).
  • appraisal period T = 3 s (achieved median ~3.5 s)
    Nominal per-agent re-appraisal interval; main value 3 s with stagger; throughput makes achieved median ~3.5 s.
  • panic quadrant threshold = A>0.35, V<0
    Operational definition V<0 and A>0.35 used for onset, I(t), and many hypothesis tests; sensitivity table varies A cutoffs.
  • LLM sampling temperature = 0.4
    Main simulation temperature 0.4; probe also sweeps 0.0 and 0.8.
  • baseline OCEAN sampling (μ,σ) = μ=0.5, σ=0.16 (baseline)
    Default crowd trait draws clamped normal μ=0.5, σ=0.16 per trait; experimental arms shift means (e.g., neuroticism 0.3/0.6/0.9).
  • perception ranges and FOV = 8 m visual, 200° FOV; audio 10/20 m
    Visual 8 m / 200° cone with occlusion; loud audio 20 m, quiet 10 m; tactile pressure EMA—geometry that shapes modality reach and wave locality.
  • memory length K and max concurrent LLM calls = K=6
    Rolling K=6 text notes and a global concurrency cap bound prompt context and realized appraisal rate.
axioms (6)
  • ad hoc to paper Inter-agent affective influence need not be an authored transfer rule; a perception–appraisal–expression loop can generate contagion-like macro patterns.
    Core methodological premise of Sections 1 and 3; contrasts with ESCAPES/ASCRIBE/OCEAN contagion rules in Related Work.
  • domain assumption Big Five traits and Russell valence–arousal are adequate internal state variables for LLM-prompted crowd agents.
    Section 3.2 adopts OCEAN and circumplex as fixed representation choices grounded in psychology citations, not derived here.
  • domain assumption Observable discrete expression vocabularies (Ekman-like faces, fixed gestures/voices/movements) are sufficient channels for social appraisal.
    Section 3.3 external state; Limitations notes discreteness as a restriction.
  • domain assumption Low-level locomotion can be decoupled from the LLM without destroying the affective phenomena under study.
    Section 3.5 and Introduction; emotion enters motion only via style, navIntent, speed, space, forcefulness.
  • domain assumption Natural-language personality sketches in the prompt induce systematic, disposition-dependent appraisal in capable LLMs.
    Assumed throughout H5 and validated only partially in appraisal probes; fails to sustain waves for gpt-4o-mini.
  • standard math Standard statistical comparisons (paired t, lagged regression with clustered SE, Fisher-z run correlations, SIS/SIR least squares, trend tests) are appropriate for these simulation logs.
    Section 4 analysis plans; Holm–Bonferroni within hypothesis families.
invented entities (2)
  • LLM appraisal-driven crowd cognitive layer (no transfer kernel) no independent evidence
    purpose: Replace authored emotion contagion rules with per-agent LLM JSON appraisal over perceived neighbor expressions.
    The architecture is the paper’s main construct; independent evidence is internal simulation behavior and probes, not external human or formal verification.
  • Felt tactile pressure channel from desired-vs-actual speed no independent evidence
    purpose: Give agents a body-sense of crowding that can enter appraisal and cause falls.
    Defined in Section 3.1; largely muted/irrelevant in Standing Line ablations; no external calibration to real pressure sensation.

pith-pipeline@v1.2.0-grok45-kimik3 · 37690 in / 4042 out tokens · 78849 ms · 2026-07-31T00:13:49.680533+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation." pith.science (2026). https://pith.science/paper/OL7AN4ZV

@misc{pith2026260725140,
  author       = {Pith},
  title        = {Pith review of: How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OL7AN4ZV}},
  note         = {Machine review of arXiv:2607.25140}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the LLM-based cognitive layer. We evaluate the architecture across five scenario environments spanning alarming, joyful, and neutral situations in different spatial layouts. The results show that the system produces emotional contagion dynamics with spatial, temporal, and personality-dependent structure in sparse, small crowds. Alarm spreads from seeded agents as a traveling front, the mean alarmed fraction settles at a nonzero plateau, and the distribution of prompted personality profiles determines whether an ambiguous alarm ignites panic and whether a provocation is interpreted as anger or fear. We further evaluate the appraisal step through controlled experiments across prompt variants, sampling temperatures, and four model backends, showing that the dynamics are backend-dependent.

Figures

Figures reproduced from arXiv: 2607.25140 by Funda Durupinar.

Figure 1
Figure 1. Figure 1: System architecture. Each LLM-driven agent perceives its neighbors within a shared environment, appraises [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Simulation scenarios used in our experiments. Agents are colored by their affective state on the circumplex: [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Affect trajectories across three scenarios. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Panic propagates outward from the seed cluster (Standing Line). [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: SIS and SIR fits to the alarmed fraction [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Perception modality ablation (Standing Line). Bars show the non-seed crowd’s alarmed fraction when only [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: LLM-chosen expression across the valence–arousal space (pooled across all scenarios). [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Crowd neuroticism determines whether an ambiguous alarm ignites panic (Evacuation). [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Calm leaders stay composed and modestly dampen the surrounding crowd’s panic (Evacuation). Non-leader [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Anger contagion is conditional on the receiver’s temperament (Contested Gate). Fraction of non-provocateur [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Prompt robustness for gemini-3.1-flash-lite at T = 0.4. Change in returned valence and arousal from the base prompt, for every prompt variant and cue, for the neurotic (high-N, left two) and stable (low-N, right two) dispositions. The paraphrases are meaning-preserving rewordings of the base instructions, one in plain language and one more formal. The ablations each drop a single prompt component (Appendi… view at source ↗
Figure 12
Figure 12. Figure 12: Appraisal map across four LLM backends. First four panels: returned valence/arousal for the four cues at sampling temperature 0.4, for the stable (squares) and neurotic (circles) personalities, each panel annotated with its parse￾error rate and mean stable-versus-neurotic gap. Rightmost panel: within-condition standard deviation of the returned valence (solid) and arousal (dashed) against sampling tempera… view at source ↗
Figure 13
Figure 13. Figure 13: The epidemic macro-signature is invariant to the affect-inertia constant [PITH_FULL_IMAGE:figures/full_fig_p030_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Wave speed does not scale with the appraisal period [PITH_FULL_IMAGE:figures/full_fig_p030_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

74 extracted references · 9 linked inside Pith

  1. [1]

    AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society,

    J. Piao, Y . Yan, J. Zhang, N. Li, J. Yan, X. Lan, Z. Lu, Z. Zheng, J. Y . Wang, D. Zhou,et al., “AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society,” 2025

  2. [2]

    Can generative AI improve social science?,

    C. A. Bail, “Can generative AI improve social science?,”Proceedings of the National Academy of Sciences, vol. 121, no. 21, p. e2314021121, 2024

  3. [3]

    Large language models as simulated economic agents: What can we learn from Homo Silicus?,

    J. J. Horton, A. Filippas, and B. S. Manning, “Large language models as simulated economic agents: What can we learn from Homo Silicus?,” tech. rep., National Bureau of Economic Research, 2023

  4. [4]

    Can large language model agents simulate human trust behavior?,

    F. Jia, Z. Ye, S. Lai, K. Shu, J. Gu, A. Bibi, Z. Hu, D. Jurgens, J. Evans, P. H. Torr,et al., “Can large language model agents simulate human trust behavior?,”Advances in Neural Information Processing Systems, vol. 37, pp. 15674–15729, 2024

  5. [5]

    Evaluating large language models as substitutes for human affective ratings in naturalistic paradigms,

    X. Yang, D. Tilwani, C. O’Reilly, and S. V . Shinkareva, “Evaluating large language models as substitutes for human affective ratings in naturalistic paradigms,”IEEE Transactions on Affective Computing, pp. 1–16, 2026

  6. [6]

    Emergent social conventions and collective bias in LLM populations,

    A. F. Ashery, L. M. Aiello, and A. Baronchelli, “Emergent social conventions and collective bias in LLM populations,”Science Advances, vol. 11, no. 20, p. eadu9368, 2025

  7. [7]

    LLM voting: Human choices and AI collective decision-making,

    J. C. Yang, D. Dailisan, M. Korecki, C. I. Hausladen, and D. Helbing, “LLM voting: Human choices and AI collective decision-making,” inProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 7, pp. 1696–1708, 2024

  8. [8]

    Large-language-model-driven agents for fire evacuation simulation in a cellular automata environment,

    P. Dang, J. Zhu, W. Li, Y . Xie, and H. Zhang, “Large-language-model-driven agents for fire evacuation simulation in a cellular automata environment,”Safety Science, vol. 191, p. 106935, 2025

  9. [9]

    RESPOND: Realistic environment simulation of population and natural disasters with LLM-driven agents,

    R. Sultimov, M. Mozikov, D. Abramov, M. Kovalchuk, M. Malykh, I. Makarov, A. Osiptsov, A. V olkov, and Y . Maximov, “RESPOND: Realistic environment simulation of population and natural disasters with LLM-driven agents,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, pp. 41694–41696, 2026

  10. [10]

    S3: Social-network simulation system with large language model-empowered agents,

    C. Gao, X. Lan, Z. Lu, J. Mao, J. Piao, H. Wang, D. Jin, and Y . Li, “S3: Social-network simulation system with large language model-empowered agents,”arXiv preprint arXiv:2307.14984, 2023

  11. [11]

    Generative agents: Interactive simulacra of human behavior,

    J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” inProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp. 1–22, 2023

  12. [12]

    Humanoid agents: Platform for simulating human-like generative agents,

    Z. Wang, Y . Y . Chiu, and Y . C. Chiu, “Humanoid agents: Platform for simulating human-like generative agents,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 167–176, 2023

  13. [13]

    Can LLMs understand social norms in autonomous driving games?,

    B. Wang, H. Duan, Y . Feng, X. Chen, Y . Fu, Z. Mo, and X. Di, “Can LLMs understand social norms in autonomous driving games?,” in2024 IEEE International Automated Vehicle Validation Conference (IAVVC), pp. 1–4, IEEE, 2024

  14. [14]

    Can LLM agents solve collaborative tasks? a study on urgency-aware planning and coordination,

    J. V . de Carvalho Silva and D. G. Macharet, “Can LLM agents solve collaborative tasks? a study on urgency-aware planning and coordination,” in2025 IEEE International Conference on Advanced Robotics (ICAR), pp. 739–744, IEEE, 2025. 24 APREPRINT- JULY29, 2026

  15. [15]

    DELIVER: A system for LLM-guided coordinated multi-robot pickup and delivery using voronoi-based relay planning,

    A. K. Srivastava, J. M. Levin, A. Derrico, and P. Dames, “DELIVER: A system for LLM-guided coordinated multi-robot pickup and delivery using voronoi-based relay planning,” in2026 IEEE/SICE International Symposium on System Integration (SII), pp. 1154–1160, IEEE, 2026

  16. [16]

    Propagating unsafe actions in LLM controlled multi-robot collaboration via single robot compromise,

    Z. Huang, Z. Liu, W. Wu, and Z. Cai, “Propagating unsafe actions in LLM controlled multi-robot collaboration via single robot compromise,”arXiv preprint arXiv:2605.15641, 2026

  17. [17]

    From spark to fire: Modeling and mitigating error cascades in LLM-based multi-agent collaboration,

    Y . Xie, C. Zhu, X. Zhang, T. Zhu, D. Ye, M. Qi, H. Chen, and W. Zhou, “From spark to fire: Modeling and mitigating error cascades in LLM-based multi-agent collaboration,”arXiv preprint arXiv:2603.04474, 2026

  18. [18]

    An alternative “description of personality

    L. R. Goldberg, “An alternative “description of personality”: The Big-Five factor structure,” inPersonality and Personality Disorders, pp. 34–47, Routledge, 2013

  19. [19]

    A circumplex model of affect.,

    J. A. Russell, “A circumplex model of affect.,”Journal of Personality and Social Psychology, vol. 39, no. 6, p. 1161, 1980

  20. [20]

    Social appraisal,

    A. Manstead and A. H. Fischer, “Social appraisal,”Appraisal Processes in Emotion: Theory, Methods, Research, pp. 221–232, 2001

  21. [21]

    Interpersonal emotion transfer: Contagion and social appraisal,

    B. Parkinson, “Interpersonal emotion transfer: Contagion and social appraisal,”Social and Personality Psychology Compass, vol. 5, no. 7, pp. 428–439, 2011

  22. [22]

    The ripple effect: Emotional contagion and its influence on group behavior,

    S. G. Barsade, “The ripple effect: Emotional contagion and its influence on group behavior,”Administrative Science Quarterly, vol. 47, no. 4, pp. 644–675, 2002

  23. [23]

    Revealing the hidden networks of interaction in mobile animal groups allows prediction of complex behavioral contagion,

    S. B. Rosenthal, C. R. Twomey, A. T. Hartnett, H. S. Wu, and I. D. Couzin, “Revealing the hidden networks of interaction in mobile animal groups allows prediction of complex behavioral contagion,”Proceedings of the National Academy of Sciences, vol. 112, no. 15, pp. 4690–4695, 2015

  24. [24]

    Mexican waves in an excitable medium,

    I. Farkas, D. Helbing, and T. Vicsek, “Mexican waves in an excitable medium,”Nature, vol. 419, no. 6903, pp. 131–132, 2002

  25. [25]

    Simulating dynamical features of escape panic,

    D. Helbing, I. Farkas, and T. Vicsek, “Simulating dynamical features of escape panic,”Nature, vol. 407, no. 6803, pp. 487–490, 2000

  26. [26]

    The mathematics of infectious diseases,

    H. W. Hethcote, “The mathematics of infectious diseases,”SIAM Review, vol. 42, no. 4, pp. 599–653, 2000

  27. [27]

    A generalized model of social and biological contagion,

    P. S. Dodds and D. J. Watts, “A generalized model of social and biological contagion,”Journal of Theoretical Biology, vol. 232, no. 4, pp. 587–604, 2005

  28. [28]

    Psychological parameters for crowd simulation: From audiences to mobs,

    F. Durupinar, U. Güdükbay, A. Aman, and N. I. Badler, “Psychological parameters for crowd simulation: From audiences to mobs,”IEEE Trans. on Vis. and Comp. Graph., vol. 22, no. 9, pp. 2145–2159, 2016

  29. [29]

    Using real life incidents for creating realistic virtual crowds with data-driven emotion contagion,

    A. E. Ba¸ sak, U. Güdükbay, and F. Durupınar, “Using real life incidents for creating realistic virtual crowds with data-driven emotion contagion,”Computers & Graphics, vol. 72, pp. 70–81, 2018

  30. [30]

    Emotion contagion in agent-based simulations of crowds: A systematic review,

    E. S. van Haeringen, C. Gerritsen, and K. V . Hindriks, “Emotion contagion in agent-based simulations of crowds: A systematic review,”Autonomous Agents and Multi-Agent Systems, vol. 37, no. 1, 2023

  31. [31]

    How the OCEAN personality model affects the perception of crowds,

    F. Durupinar, N. Pelechano, J. Allbeck, U. Güdükbay, and N. I. Badler, “How the OCEAN personality model affects the perception of crowds,”IEEE Computer Graphics and Applications, vol. 31, no. 3, pp. 22–31, 2011

  32. [32]

    Can large language models transform computa- tional social science?,

    C. Ziems, W. Held, O. Shaikh, J. Chen, Z. Zhang, and D. Yang, “Can large language models transform computa- tional social science?,”Computational Linguistics, vol. 50, no. 1, pp. 237–291, 2024

  33. [33]

    Project Sid: Many-agent simulations toward AI civilization,

    A. AL, A. Ahn, N. Becker, S. Carroll, N. Christie, M. Cortes, A. Demirci, M. Du, F. Li, S. Luo,et al., “Project Sid: Many-agent simulations toward AI civilization,”arXiv preprint arXiv:2411.00114, 2024

  34. [34]

    Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia,

    A. S. Vezhnevets, J. P. Agapiou, A. Aharon, R. Ziv, J. Matyas, E. A. Duéñez-Guzmán, W. A. Cunningham, S. Osindero, D. Karmon, and J. Z. Leibo, “Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia,”arXiv preprint arXiv:2312.03664, 2023

  35. [35]

    Wisdom of the machines: Exploring collective intelligence in LLM crowds,

    Y . Talebirad, A. Parsaee, V . Ohal, A. Nadiri, C. Szepesvari, Y . Mouje, and E. Redman, “Wisdom of the machines: Exploring collective intelligence in LLM crowds,” inFirst Workshop on Social Simulation with LLMs, 2025

  36. [36]

    Shall we team up: Exploring spontaneous cooperation of competing LLM agents,

    Z. Wu, R. Peng, S. Zheng, Q. Liu, X. Han, B. I. Kwon, M. Onizuka, S. Tang, and C. Xiao, “Shall we team up: Exploring spontaneous cooperation of competing LLM agents,” inFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 5163–5186, 2024

  37. [37]

    CrowdLLM: Building LLM-based digital populations augmented with generative models,

    R. F. Lin, K. Tian, H. Zheng, C. Zhang, L. Zeng, and S. Huang, “CrowdLLM: Building LLM-based digital populations augmented with generative models,”arXiv preprint arXiv:2512.07890, 2025

  38. [38]

    Emergent crowds dynamics from language-driven multi-agent interactions,

    Y . Liu, L. Shatzel, B. Haworth, and T. Schneider, “Emergent crowds dynamics from language-driven multi-agent interactions,”arXiv preprint arXiv:2508.15047, 2025. 25 APREPRINT- JULY29, 2026

  39. [39]

    Gen-C: Populating virtual worlds with generative crowds,

    A. Panayiotou, P. Charalambous, and I. Karamouzas, “Gen-C: Populating virtual worlds with generative crowds,” Proc. ACM Comput. Graph. Interact. Tech., vol. 9, May 2026

  40. [40]

    CrowdVLA: Embodied vision-language- action agents for context-aware crowd simulation,

    J. Hwang, S.-E. Hong, J. Kim, J. Seon, G. Nam, H. Jang, and H. Kang, “CrowdVLA: Embodied vision-language- action agents for context-aware crowd simulation,”arXiv preprint arXiv:2604.05525, 2026

  41. [41]

    Research of crowd behavior simulation methods based on relationships and emotional evolution,

    Z. Gu, T. Gao, Z. Bai, and Q. Mi, “Research of crowd behavior simulation methods based on relationships and emotional evolution,”ACM Transactions on Modeling and Computer Simulation, vol. 36, no. 2, pp. 1–18, 2026

  42. [42]

    AI-driven crowd behavior simulation for emergency evacuation management,

    T. Saravanan, U. Singhal, N. Magadum, P. Ramasamy, and K. Paramanandam, “AI-driven crowd behavior simulation for emergency evacuation management,” inInternational Conference on Data Science and Applications, pp. 65–78, Springer, 2025

  43. [43]

    When agents learn to think: Large language model-enhanced agent-based modeling for crowd evacuation in disaster scenarios,

    S. Yang, L. Ceferino, Y . Zhang, C. Gu, T. Guo, and G. Kondo, “When agents learn to think: Large language model-enhanced agent-based modeling for crowd evacuation in disaster scenarios,”Reliability Engineering & System Safety, p. 112056, 2025

  44. [44]

    Evaluating and inducing personality in pre-trained language models,

    G. Jiang, M. Xu, S.-C. Zhu, W. Han, C. Zhang, and Y . Zhu, “Evaluating and inducing personality in pre-trained language models,”Advances in Neural Information Processing Systems, vol. 36, pp. 10622–10643, 2023

  45. [45]

    PersonaLLM: Investigating the ability of large language models to express personality traits,

    H. Jiang, X. Zhang, X. Cao, C. Breazeal, D. Roy, and J. Kabbara, “PersonaLLM: Investigating the ability of large language models to express personality traits,” inFindings of the Association for Computational Linguistics: NAACL 2024, pp. 3605–3627, 2024

  46. [46]

    Editing personality for large language models,

    S. Mao, X. Wang, M. Wang, Y . Jiang, P. Xie, F. Huang, and N. Zhang, “Editing personality for large language models,” inCCF International Conference on Natural Language Processing and Chinese Computing, pp. 241–254, Springer, 2024

  47. [47]

    Is GPT a computational model of emotion?,

    A. N. Tak and J. Gratch, “Is GPT a computational model of emotion?,” in2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII), pp. 1–8, IEEE, 2023

  48. [48]

    Fine-grained affective processing capabilities emerging from large language models,

    J. Broekens, B. Hilpert, S. Verberne, K. Baraka, P. Gebhard, and A. Plaat, “Fine-grained affective processing capabilities emerging from large language models,” in2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII), pp. 1–8, IEEE, 2023

  49. [49]

    EmoBench: Evaluating the emotional intelligence of large language models,

    S. Sabour, S. Liu, Z. Zhang, J. Liu, J. Zhou, A. Sunaryo, T. Lee, R. Mihalcea, and M. Huang, “EmoBench: Evaluating the emotional intelligence of large language models,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5986–6004, 2024

  50. [50]

    Emotional intelligence of large language models,

    X. Wang, X. Li, Z. Yin, Y . Wu, and J. Liu, “Emotional intelligence of large language models,”Journal of Pacific Rim Psychology, vol. 17, p. 18344909231213958, 2023

  51. [51]

    Apathetic or empathetic? evaluating LLMs’ emotional alignments with humans,

    J.-t. Huang, M. H. Lam, E. J. Li, S. Ren, W. Wang, W. Jiao, Z. Tu, and M. R. Lyu, “Apathetic or empathetic? evaluating LLMs’ emotional alignments with humans,”Advances in Neural Information Processing Systems, vol. 37, pp. 97053–97087, 2024

  52. [52]

    EvoEmo: Towards evolved emotional policies for adversarial LLM agents in multi-turn price negotiation,

    Y . Long, L. Xu, L. Beckenbauer, Y . Liu, and A. Brintrup, “EvoEmo: Towards evolved emotional policies for adversarial LLM agents in multi-turn price negotiation,”arXiv preprint arXiv:2509.04310, 2025

  53. [53]

    ESCAPES: Evacuation simulation with children, authorities, parents, emotions, and social comparison,

    J. Tsai, N. Fridman, E. Bowring, M. Brown, S. Epstein, G. Kaminka, S. Marsella,et al., “ESCAPES: Evacuation simulation with children, authorities, parents, emotions, and social comparison,” inProceedings of the 10th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), pp. 457–464, 2011

  54. [54]

    Agent-based modeling of emotion contagion in groups,

    T. Bosse, R. Duell, Z. A. Memon, J. Treur, and C. N. van der Wal, “Agent-based modeling of emotion contagion in groups,”Cognitive Computation, vol. 7, no. 1, pp. 111–136, 2015

  55. [55]

    Flocks, herds and schools: A distributed behavioral model,

    C. W. Reynolds, “Flocks, herds and schools: A distributed behavioral model,” inProceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, pp. 25–34, 1987

  56. [56]

    Constants across cultures in the face and emotion.,

    P. Ekman and W. V . Friesen, “Constants across cultures in the face and emotion.,”Journal of Personality and Social Psychology, vol. 17, no. 2, p. 124, 1971

  57. [57]

    Reciprocal velocity obstacles for real-time multi-agent navigation,

    J. Van den Berg, M. Lin, and D. Manocha, “Reciprocal velocity obstacles for real-time multi-agent navigation,” in 2008 IEEE International Conference on Robotics and Automation, pp. 1928–1935, IEEE, 2008

  58. [58]

    Freeze for action: neurobiological mechanisms in animal and human freezing,

    K. Roelofs, “Freeze for action: neurobiological mechanisms in animal and human freezing,”Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 372, no. 1718, p. 20160206, 2017

  59. [59]

    Hatfield, J

    E. Hatfield, J. T. Cacioppo, and R. L. Rapson,Emotional Contagion. Cambridge University Press, 1994

  60. [60]

    ‘The Battle of Westminster’: Developing the social identity model of crowd behaviour in order to explain the initiation and development of collective conflict,

    S. D. Reicher, “‘The Battle of Westminster’: Developing the social identity model of crowd behaviour in order to explain the initiation and development of collective conflict,”European Journal of Social Psychology, vol. 26, no. 1, pp. 115–134, 1996. 26 APREPRINT- JULY29, 2026

  61. [61]

    The role of social identity processes in mass emergency behaviour: An integrative review,

    J. Drury, “The role of social identity processes in mass emergency behaviour: An integrative review,”European Review of Social Psychology, vol. 29, no. 1, pp. 38–81, 2018

  62. [62]

    Appraisal considered as a process of multilevel sequential checking,

    K. R. Scherer, “Appraisal considered as a process of multilevel sequential checking,”Appraisal Processes in Emotion: Theory, Methods, Research, vol. 92, no. 120, p. 57, 2001

  63. [63]

    Progress on a cognitive-motivational-relational theory of emotion.,

    R. S. Lazarus, “Progress on a cognitive-motivational-relational theory of emotion.,”American Psychologist, vol. 46, no. 8, p. 819, 1991

  64. [64]

    Dynamics of crowd disasters: An empirical study,

    D. Helbing, A. Johansson, and H. Z. Al-Abideen, “Dynamics of crowd disasters: An empirical study,”Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, vol. 75, no. 4, p. 046109, 2007

  65. [65]

    Epidemic modeling with generative agents,

    R. Williams, N. Hosseinichimeh, A. Majumdar, and N. Ghaffarzadegan, “Epidemic modeling with generative agents,”arXiv preprint arXiv:2307.04986, 2023

  66. [66]

    Affect bursts,

    K. R. Scherer, “Affect bursts,”Emotions: Essays on Emotion Theory, vol. 161, p. 196, 1994

  67. [67]

    Nonverbal channel use in communication of emotion: How may depend on why.,

    B. App, D. N. McIntosh, C. L. Reed, and M. J. Hertenstein, “Nonverbal channel use in communication of emotion: How may depend on why.,”Emotion, vol. 11, no. 3, p. 603, 2011

  68. [68]

    Bad is stronger than good,

    R. F. Baumeister, E. Bratslavsky, C. Finkenauer, and K. D. V ohs, “Bad is stronger than good,”Review of General Psychology, vol. 5, no. 4, pp. 323–370, 2001

  69. [69]

    The psychology of rumor.,

    G. W. Allport and L. Postman, “The psychology of rumor.,” 1947

  70. [70]

    Patterns of cognitive appraisal in emotion.,

    C. A. Smith and P. C. Ellsworth, “Patterns of cognitive appraisal in emotion.,”Journal of Personality and Social Psychology, vol. 48, no. 4, p. 813, 1985

  71. [71]

    R. S. Lazarus,Emotion and adaptation. Oxford University Press, 1991

  72. [72]

    The anatomy of anger: An integrative cognitive model of trait anger and reactive aggression,

    B. M. Wilkowski and M. D. Robinson, “The anatomy of anger: An integrative cognitive model of trait anger and reactive aggression,”Journal of Personality, vol. 78, no. 1, pp. 9–38, 2010

  73. [73]

    Do personality tests generalize to large language models?,

    F. Dorner, T. Sühr, S. Samadi, and A. Kelava, “Do personality tests generalize to large language models?,” in Socially Responsible Language Modelling Research, 2023

  74. [74]

    Self-assessment tests are unreliable measures of LLM personality,

    A. Gupta, X. Song, and G. Anumanchipalli, “Self-assessment tests are unreliable measures of LLM personality,” inProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pp. 301–314, 2024. Appendix A Appraisal Prompt The appraisal call sends a fixed system prompt and a per-cycle user prompt, reproduced below. System ...