{"id":"d3ca62d8-e2aa-4e1e-9dbb-00a3f338dd39","arxiv_id":"2506.13783","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"In three simulation runs, LLM agents primed with a swine flu article reduced social activity compared with controls, but the effect is confounded by prompt instructions and no inferential statistics are reported.","lead":"Generative agents in a simulated town who read a news article about a swine flu outbreak spent less time in cafes and parks, held fewer conversations, and skipped a planned party more often than agents who saw no disease news. The study tests whether large language model agents can serve as experimental subjects for behavioral immune system research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The disease-threat prompt (A.2) explicitly tells agents to consider how the news shapes social interactions and injects a 'neighbors unwell' cue that the no-threat prompt (A.3) lacks, so the reported sociality drop may be prompt compliance, not an emergent threat response.","rationale":"The reader's weakest assumption is the same one I would flag: the observed sociality drop may be an artifact of the manipulation prompt rather than an emergent response to disease threat. I agree with the rejection because internal validity is the foundation of the paper's claim; if the threat condition differs in instruction and injected cues, the causal attribution is not identified. One correction to the reader's specifics: the no-threat prompt (A.3) also grants 'flexibility to add, remove, or adjust scheduled activities,' so that feature is not unique. The unique confounds are the explicit directive to reason about 'social interactions, and daily activities' and the added 'neighbors had seemed unwell' sentence. These are serious enough to block the central claim. The proposed yoked-instruction control would settle the matter. I credit the paper for reporting the prompts transparently and for including a noninfectious-disease condition, but that condition's prompt is undocumented and the comparison is not matched, so it does not rescue the causal claim. With only three runs and no statistical tests, the 'significant' language is also unsupported. Hence REJECT remains appropriate.","tokens_in":22618,"tokens_out":8256,"duration_ms":86570,"concrete_test":"Run a yoked-instruction control. Give the no-threat condition the disease-threat prompt template (A.2) with the swine flu article replaced by a neutral article (e.g., a local technology fair) and with the 'Reading the news reminded ... neighbors had seemed unwell' sentence removed. Keep all other system changes identical. Run at least three matched pairs using the same seeds. If this control shows sociality reductions comparable to the original disease-threat condition, the effect is attributable to the prompt instruction, not to disease threat; if it tracks the original no-threat baseline, the concern is mitigated.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is causal: swine flu news reduced agents' social engagement. The manipulation, however, is confounded. The disease-threat prompt (Section A.2) asks the agent to 'carefully think step-by-step about how these together might shape ... social interactions, and daily activities,' and then appends the sentence 'Reading the news reminded [agent] that a few neighbors had seemed unwell lately.' The no-threat prompt (A.3) contains neither element; it asks only for an updated status from prior state. Both prompts allow schedule adjustments, so schedule flexibility is not the distinguishing feature, but the explicit instruction to reason about social consequences and the injected local-sickness cue are unique to the threat condition. These are exactly the features that would induce the measured reductions (party attendance, third-place time, conversation count) even without any genuine disease-avoidance process. The type 2 diabetes condition does not resolve the confound. Its full prompt is not shown (only the article, A.4), so we cannot determine whether it included the 'neighbors unwell' sentence and the social-consequences instruction. If it did, the null result would suggest the instruction alone is insufficient; if it did not, the condition is not a proper control. Moreover, the diabetes article is not matched for emotional valence or perceived severity. With n=3 and no inferential statistics, the 'significant' language is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a generative agent-based modeling (GABM) study adapted from Park et al. (2023), in which 25 LLM-driven agents in a simulated town either read a news article about a swine flu (H1N1) epidemic (disease-threat condition) or received no such news (no-threat condition). Across three independent runs, the authors report that disease-threat agents showed lower party attendance, less time in third places, fewer steps, lower conversation initiation probability, fewer total conversations, and higher self-reported disease avoidance, with agents' interview responses attributing these changes to infection concerns. A single additional run included a noninfectious-disease condition (type 2 diabetes) that closely resembled the no-threat condition on most measures. The paper interprets these outcomes as evidence that disease threat reduces sociality in generative agents and as support for the behavioral immune system framework, arguing that GABM can complement existing methods in social psychology.","tokens_in":22969,"tokens_out":2392,"duration_ms":26566,"significance":"If the central causal claim were sound, the paper would offer a compelling demonstration that GABM can test evolved psychological mechanisms such as the behavioral immune system, with the added strengths of transparent prompts, run-level data in Appendix B, and a within-simulation baseline (February 13) against which February 14 changes are compared. The authors also make constructive methodological improvements to the Park et al. sandbox, including structured outputs and a probabilistic conversation-initiation mechanism, and they provide detailed qualitative interview excerpts that enrich the quantitative results. However, the manuscript's central inference is currently undermined by a prompt confound that is visible in its own appendix, and by the absence of any inferential statistics for an n of only three paired runs. Because the confound directly targets the manipulation rather than an auxiliary detail, the reported 'significant' reductions cannot be distinguished from compliance with the experimental instructions. The contribution, as it stands, is therefore a proof-of-concept of the simulation architecture rather than a valid experimental test of the stated hypothesis.","major_comments":[{"comment":"The disease-threat prompt (A.2) instructs agents to 'carefully think step-by-step about how these together might shape ... social interactions, and daily activities' and to 'add, remove, or adjust scheduled activities,' and it appends the sentence 'Reading the news reminded [agent] that a few neighbors had seemed unwell lately.' The no-threat prompt (A.3) contains neither the social-consequences instruction nor the unwell-neighbors cue. Both prompts allow schedule adjustment, so the observed reductions in party attendance, third-place time, and conversation frequency may largely reflect the agents following the explicit instruction to reconsider their social plans, rather than an emergent disease-avoidance process. The paper's claim in Section 3.1.1 that 'the news about swine flu clearly reduced social engagement' is therefore unsupported by the present manipulation.","section":"A.2 vs. A.3, Section 2.2, Section 3.1.1"},{"comment":"All between-condition comparisons rest on three paired runs, yet the paper reports no standard errors, confidence intervals, or formal tests. Phrases such as 'significantly reduced' (Abstract, Section 3.2.1) and 'clearly reduced' (Section 3.1.1) are not backed by any inferential statistic. Given the small n and the absence of variability information, the magnitudes reported (e.g., party attendance 1.33 vs. 7.33; third-place time 30.11 vs. 112.65 minutes) should be presented as descriptive differences, with all claims of statistical or practical significance explicitly justified or removed.","section":"Section 3, Tables 3 and 4"},{"comment":"The validity check using type 2 diabetes does not resolve the prompt confound. The full prompt for the noninfectious-disease condition is not shown; only the news article text (A.4) is given. Without the full prompt, readers cannot determine whether that condition also included the 'neighbors unwell' sentence or the instruction to reason about social interactions. If it did, the authors would need to explain why the instruction alone produced no behavioral change; if it did not, the condition is not a matched control for the disease-threat prompt. In addition, the diabetes article is not matched to the flu article on emotional valence, perceived severity, or personal relevance, so any null effect is open to alternative interpretations.","section":"Section 2.4 and A.4"},{"comment":"The conversation-initiation outcome is computed as a per-encounter frequency, yet the paper compares aggregate percentages across conditions without accounting for the number and composition of encounters. Because disease-threat agents moved less and spent less time in third places, they presumably had fewer encounters overall, and the comparison does not adjust for encounter volume or for the mix of familiar versus unfamiliar agents across conditions. The claim that agents were '12.3 percentage points less likely to initiate conversations' may therefore conflate reduced opportunity with reduced willingness. A per-agent encounter-to-conversation ratio with its distribution across agents and runs would be needed to support the stated interpretation.","section":"Section 3.2.1 and Table 3"}],"minor_comments":[{"comment":"The notation 'p= 0.1×score' omits the subscript on p and is inconsistent with the later use of p_i; please align the notation throughout the section.","section":"Section 2.3.2, Eq. (1)"},{"comment":"The text alternates between 'Hobbs Cafe' and 'Hobb's Cafe' (e.g., Section 2.3.4 interview question). Please standardize the name of the establishment.","section":"Section 3.1.1 and Appendix B.1"},{"comment":"The description of the behavioral immune system cites several key works, but the activation cues listed (e.g., 'disgusting images or unpleasant odors') are not all used in the study; consider trimming or explicitly connecting each cue to the manipulation.","section":"Section 1.2 and references"},{"comment":"The noninfectious-disease condition is reported only for 'the first run,' but Table 3 lists a full 'Run 1 (Noninfectious-Disease Condition)' column without indicating whether this is the sole run; the table's header should make the n=1 explicit to avoid implying three runs.","section":"Section 3.5 and Table 3"},{"comment":"The limitations paragraph appropriately notes LLMs' WEIRD biases and hallucination risks, but it does not acknowledge the more immediate limitation that the manipulation prompt explicitly requested the behavior under study; this should be addressed in the limitations discussion.","section":"Section 4, Discussion"}],"recommendation":"reject","confidential_remarks":"The paper's own Appendix A.2 provides direct evidence of the central confound: the disease-threat prompt explicitly instructs agents to alter their social interactions and schedules, which makes the reported effect a compliance artifact rather than an emergent phenomenon. This is not a case of disagreement with the field's consensus; it is an internal inconsistency between the manipulation and the causal claim. I would advise the editor that, absent a redesign with matched prompts across conditions and inferential statistics, the manuscript cannot support its abstract claims, regardless of how suggestive the qualitative excerpts appear. If the authors re-run with a properly matched control prompt, the work could become a valuable proof-of-concept."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is the first to put the behavioral immune system in a generative agent-based model, and that framing is worth something. The methods appendix is unusually open: full prompts, run-level tables, interview transcripts, heatmaps. The outcome measures (party attendance, third-place time, conversation counts) are concrete and the diabetes control is a sensible idea. The interview responses give a qualitative consistency check. So there is real scaffolding here, not just a prompt-and-pray result.\n\nThe problem is that the main comparison is not clean. The disease-threat prompt (A.2) tells the agent to 'carefully think step-by-step about how these together might shape... social interactions, and daily activities' and then appends 'Reading the news reminded [agent] that a few neighbors had seemed unwell lately.' The no-threat prompt (A.3) contains neither of these elements. That means the manipulation bundles a disease news story with an explicit instruction to think about social consequences and a concrete local cue that someone nearby is sick. The no-threat condition is not just a neutral version of the same task; it is a different task. The observed drop in sociality could easily be demand compliance rather than emergent disease-avoidance. The diabetes condition does not resolve this because its full prompt is not shown, so you cannot tell whether it matched the threat condition on those two features. If it did not, it is not a proper control; if it did, the result would be more informative but still unproven.\n\nThere is also a statistical overclaim. The paper says the effect was 'significant' and 'clearly reduced' social engagement, but there are only three runs per condition, no standard errors, no confidence intervals, and no inferential test. With n=3, the word 'significant' carries no statistical meaning. The descriptive differences are large, and some pattern is evident, but the language outruns the evidence.\n\nI would not cite this result as evidence about the behavioral immune system. But I would send it to peer review rather than desk reject, because the topic is timely, the method is being applied in a novel domain, and the flaws are fixable with a matched prompt design, a shown control prompt, and at least a few more runs. The authors clearly know the relevant literature and have done careful descriptive work. The confound is central, so I would expect heavy revision.\n\nMy take: this deserves a serious referee, not because the current claim stands, but because the raw material is usable and the questions it raises about designing controlled GABM experiments are exactly what a reviewer should push on.","headline":"The headline finding is likely a prompt artifact: the disease-threat condition explicitly instructs agents to consider social consequences and injects a neighborhood sickness cue that the control lacks, so the causal claim is not supported as-is.","tokens_in":23431,"tokens_out":2261,"would_cite":false,"duration_ms":27465,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reading a swine flu news story made LLM agents reduce their social engagement in a simulated town.","keywords":["behavioral immune system","generative agents","large language models","agent-based modeling","disease avoidance","sociality","simulation","infectious disease threat"],"falsifier":"A decisive test would run the disease-threat condition with the news article unchanged but remove the planning-prompt sentences instructing agents to think about how the news shapes their social interactions and to add, remove, or adjust scheduled activities, and also remove the note that reading the news reminded them that neighbors had seemed unwell. If party attendance and third-place time then return to no-threat levels, the reported withdrawal is a product of the prompt structure rather than an emergent behavioral immune response; if the withdrawal persists, the prompt instructions are not necessary for the effect.","tokens_in":22458,"feed_emoji":"🦠","tokens_out":7115,"duration_ms":66335,"temperature":0.7,"pith_summary":"This paper asks whether the mere perception of an infectious disease threat changes how socially engaged people are, and it tests the question with synthetic people rather than human subjects. In a simulated town of 25 large-language-model agents, the researchers compared agents who read a local newspaper article about a swine flu outbreak with agents who received no such news. The outbreak-news agents attended the town's Valentine's Day party far less often, spent much less time in cafes and parks, took fewer steps, and had fewer conversations; in interviews, they cited avoiding infection as the reason. A control condition that substituted a noninfectious disease story (type 2 diabetes) did not produce the same withdrawal, which the authors read as evidence that the agents distinguished infectious from noninfectious threats. If the finding is robust, it would extend the behavioral immune system account (the idea that organisms avoid disease risk before infection) to generative agents and suggest that LLM-based simulations can serve as experimental testbeds for social psychology.","feed_headline":"Disease-threat news makes a simulated town go quiet","feed_subtitle":"After swine flu news, LLM agents skipped parties, spent less time in cafes, and talked less, a pattern read as disease avoidance.","key_machinery":"The central machinery is generative agent-based modeling (GABM) running on the Smallville sandbox, a simulated town of 25 agents whose behavior is driven by large language models through a memory-reflection-planning-action cycle. The experimental trigger is a single local news article inserted into the agents' evening planning prompt; the same prompt is what asks agents to think step-by-step about how the news might shape their thoughts, emotions, social interactions, and daily activities, and explicitly permits them to add, remove, or adjust scheduled activities. Sociality is measured through spatial behavior (party attendance, third-place time, steps) and conversational patterns, with conversation initiation converted from an LLM-assigned likelihood score into a Bernoulli trial. This machinery matters because it is the channel through which a short news text is claimed to propagate into emergent, measurable changes in individual and community-level social behavior.","core_discovery":"On the paper's own terms, the central discovery is that a disease-threat news prime reliably reduces sociality among generative agents, and that agents themselves articulate disease-avoidance motivations for their behavior. Across three independent simulation runs, agents who read about a swine flu epidemic attended the Valentine's Day party at an average of 1.33 attendees versus 7.33 in the no-threat condition, spent 73.3% less time in third places, took 22.3% fewer steps, initiated conversations 12.3 percentage points less often, and engaged in 48.4% fewer conversations. Interview responses attributed these changes to concerns about infection, and a case study showed the cafe owner postponing the party and adopting hygiene practices. The noninfectious-disease control condition produced social engagement closely resembling the no-threat condition, supporting the paper's claim that the withdrawal was specific to infectious disease risk rather than any health-related news.","pith_inferences":["A sharper control than the authors used would give the no-threat agents a neutral or unrelated news article; if such a story also suppressed sociality, the reported effect would be a generic response to news rather than disease-specific avoidance, a possibility the diabetes comparison does not fully rule out.","The planning prompt's explicit instruction to consider how the news shapes social interactions and to adjust scheduled activities makes it plausible that the effect partly reflects instruction-following; future runs that strip those instructions from the prompt while keeping the article could separate emergence from compliance.","The current simulations do not model actual disease transmission, so the town's reduced sociality is a behavioral response without epidemiological consequences; coupling this agent behavior to an infection model could test whether the observed withdrawal would flatten an epidemic curve.","If the behavioral-immune-system pattern in LLM agents is stable across models and prompt variants, researchers could use the platform to explore how different threat framings or policy messages shift collective behavior, an extension beyond the paper's two-day design."],"forward_implications":["If the reported effect is real, a single news article can shift the social behavior of an entire simulated community, giving researchers a low-cost way to probe how threat information cascades through social networks.","The specificity check implies that LLM-driven agents can represent the difference between infectious and noninfectious health threats, so the same platform could test other pathogen cues such as disgust stimuli, sickness symptoms, or crowding.","The stronger drop in conversations with unfamiliar agents than familiar agents suggests disease threat selectively suppresses interaction with strangers, a prediction the paper identifies as a candidate for human research.","Because the agents' avoidance behavior emerges from natural-language planning rather than hand-coded rules, GABM could complement traditional rule-based agent-based models in studying pandemic-related behavior."],"supporting_citations":[{"why":"Provides the Smallville simulation environment and the generative-agent memory-planning architecture on which the experiment is built.","marker":"Park et al. 2023"},{"why":"Supplies the experimental paradigm and the swine flu news text used as the disease-threat prime.","marker":"Huang et al. 2011"},{"why":"Defines the behavioral immune system as a disease-avoidance mechanism, the theoretical target of the simulations.","marker":"Schaller & Duncan 2007"},{"why":"Reviews evidence that disease cues produce avoidant social behavior, motivating the study's predictions.","marker":"Ackerman et al. 2018"},{"why":"Provides the Fundamental Social Motives Inventory used to measure disease-avoidance and affiliation motives in agent interviews.","marker":"Neel et al. 2016"},{"why":"Shows LLM-guided agents can adjust quarantine decisions in epidemic models, positioning the current study's focus on emergent sociality.","marker":"Williams et al. 2023"},{"why":"Introduces the comparison-condition logic for showing disease-threat responses are specific to infectious risk.","marker":"Faulkner et al. 2004"},{"why":"Extends that comparison-condition logic by using non-disease threats, grounding the diabetes control design.","marker":"Wang & Ackerman 2019"}],"fun_headline_variants":["Flu news makes AI agents socially distant","LLM agents read about flu, then skip the party","Infectious headlines drive AI townsfolk to isolate","Disease-threat news prompts LLM agents to go quiet","Simulated citizens avoid cafes after outbreak news"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that the reduction in social activity is an emergent reaction to the disease-threat news rather than a direct consequence of the planning prompt telling agents to think about how the news affects their social interactions and to adjust their plans accordingly.","fun_headline_variants_meta":{"raw":{"variants":["Flu news makes AI agents socially distant","LLM agents read about flu, then skip the party","Infectious headlines drive AI townsfolk to isolate","Disease-threat news prompts LLM agents to go quiet","Simulated citizens avoid cafes after outbreak news"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1406,"prompt_tokens":878,"completion_tokens":528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":454}},"tokens_in":494,"tokens_out":528,"duration_ms":6685,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:55:05.585214+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would run the disease-threat condition with the news article unchanged but remove the planning-prompt sentences instructing agents to think about how the news shapes their social interactions and to add, remove, or adjust scheduled activities, and also remove the note that reading the news reminded them that neighbors had seemed unwell. If party attendance and third-place time then return to no-threat levels, the reported withdrawal is a product of the prompt structure rather than an emergent behavioral immune response; if the withdrawal persists, the prompt instructions are not necessary for the effect.","supporting_citations":[{"cited_title":"Y., Sedlovskaya, A., Ackerman, J","cited_arxiv_id":null,"evidence_quote":"Supplies the experimental paradigm and the swine flu news text used as the disease-threat prime."},{"cited_title":"and Duncan, L","cited_arxiv_id":null,"evidence_quote":"Defines the behavioral immune system as a disease-avoidance mechanism, the theoretical target of the simulations."},{"cited_title":"M., Hill, S","cited_arxiv_id":null,"evidence_quote":"Reviews evidence that disease cues produce avoidant social behavior, motivating the study's predictions."},{"cited_title":"T., White, A","cited_arxiv_id":null,"evidence_quote":"Provides the Fundamental Social Motives Inventory used to measure disease-avoidance and affiliation motives in agent interviews."},{"cited_title":"H., and Duncan, L","cited_arxiv_id":null,"evidence_quote":"Introduces the comparison-condition logic for showing disease-threat responses are specific to infectious risk."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends that comparison-condition logic by using non-disease threats, grounding the diabetes control design."}],"review_version":1}