{"id":"dac954ee-c694-475e-8ec2-78c7c39b0f99","arxiv_id":"2411.11581","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"OASIS lets one million LLM agents interact on simulated X and Reddit platforms and reproduces information spreading, polarization, and herd effects at scale.","lead":"OASIS is a simulator that runs up to one million AI agents on realistic mockups of X and Reddit, with agents that post, follow, like, and comment. The team reports reproducing familiar online behaviors such as rumor spreading, opinion polarization, and herding, and observes that larger simulated crowds give more varied and helpful answers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The scale claim ('larger agent group scale leads to more helpful and diverse opinions') is confounded: in §3.4.2/F.4.2 agent count, injected post count (30/300/3000), and cache size (50/500/5000) all change together, so the reported scale effects may be content-volume/exposure effects.","rationale":"The engineering contribution is substantial: an open-sourced, modular simulator that reaches one million agents, with ablations and reproducibility runs, deserves credit. My concern is targeted at the causal scale claim, not at the replication of information spreading, polarization, or herd effects at fixed scale, which are supported by comparisons to real-world datasets, modulo the acknowledged depth and RecSys mismatches. The reader identified the same confounding in §3.4.2/F.4.2, and I agree it is the most load-bearing weakness. A conditional verdict is appropriate: the paper should either add the control experiments or soften the abstract's causal claim to a correlational observation that larger simulations generate more content. The concern is not an internal inconsistency or a disagreement with consensus; it is an under-identified experimental contrast that can be resolved by additional runs.","tokens_in":26564,"tokens_out":5878,"duration_ms":57714,"concrete_test":"Run a 2x2 factorial on the counterfactual herd-effect setup of §3.4.2/F.4.2: (N=100, 30 posts/step, cache 50), (N=10,000, 30 posts/step, cache 50), (N=100, 3,000 posts/step, cache 5,000), and (N=10,000, 3,000 posts/step, cache 5,000), with all other settings identical. If the down-treated vs. control disagree-score gap tracks post volume or cache size rather than N, Finding 5 is an exposure-volume artifact. To settle Finding 4, repeat the polarization scale experiment at N=196 and N=10,196 while subsampling generated users' comments so each core user sees the same number of comments per time step; if helpfulness and diversity no longer improve with N, the scale claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central scientific novelty beyond the engineering artifact is the scale claim (abstract; Findings 4 and 5). The evidence for it does not isolate population size. In the counterfactual herd-effect experiment, Section F.4.2 states that as agents increase from 100 to 1,000 to 10,000, the controlled user creates 30, 300, or 3,000 posts per time step and the recommendation cache holds 50, 500, or 5,000 posts. Thus the down-treated disagree-score gap shown in §3.4.2, the basis for Finding 5, can be explained by more down-treated posts being injected and more of them being visible, not by a larger agent population. A larger cache also means each activated agent samples from a larger pool, changing exposure independently of group size. For Finding 4, the same 196 core users are evaluated as the synthetic population grows; the number of comments and recommendations they receive grows with N, so the reported increase in diversity and helpfulness could arise from sampling more responses under unchanged per-agent behavior. The paper reports no condition that varies N while holding post volume, injection rate, and cache size fixed, nor one that varies volume at fixed N. Without such a control, the headline observation that 'larger agent group scale leads to more enhanced group dynamics and more diverse and helpful opinions' is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OASIS, an LLM-agent-based social media simulator with an environment server, recommendation systems, a 21-action agent module, a time engine, and a scalable inference backend. The authors claim generalizability across X and Reddit, demonstrate simulations with up to one million agents, and report replications of information spreading, group polarization, and herd effects. They also report two scale findings: that larger agent populations yield more helpful and diverse opinions (Finding 4) and that larger groups amplify dislike-driven herd effects (Finding 5).","tokens_in":26873,"tokens_out":6280,"duration_ms":61050,"significance":"If the claims hold, OASIS is a valuable open-source platform for studying social-media dynamics at a scale that prior LLM-agent simulators have not reached. The modular five-component architecture, the open-source release, and the 1M-agent demonstration on 24 A100s are genuine engineering contributions, and the repeated propagation runs in Appendix F.2.2 are a commendable reproducibility check. However, the empirical validation is weakened by confounded scale experiments and by a propagation fit with roughly 30% normalized RMSE and a systematically lower depth curve. The scale findings are the paper's headline novelty, so the confounds are load-bearing.","major_comments":[{"comment":"The scale comparison for the counterfactual herd effect varies the number of agents (100, 1k, 10k) together with the number of posts injected per time step (30, 300, 3k) and the recommendation cache size (50, 500, 5k). Because activated agents sample 5 posts from the cache (F.4.2), the probability of seeing a down-treated post increases with both injection rate and cache size, so the monotonic down-treated disagree-score gap in Figure 8 can be explained by exposure volume rather than population size. Finding 5 requires a control that varies N while holding post volume and cache size fixed, or varies volume at fixed N.","section":"§3.4.2, F.4.2"},{"comment":"Finding 4 compares the same 196 core users under population sizes 196, 10,196, and 100,196, but no condition isolates N from the amount of content the core users receive. As the surrounding population grows, the number of comments, likes, and recommendations directed at the 196 core users grows as well; the observed increase in diversity and helpfulness could therefore reflect a larger response sample under unchanged per-agent behavior. A matched condition that fixes the post volume and recommendation exposure per core user is needed to support the causal claim that 'larger agent group scale leads to ... more diverse and helpful agents' opinions.'","section":"§3.4.1, Fig. 7"},{"comment":"Finding 1 states that OASIS replicates the information-spreading process, but the reported normalized RMSE is about 30% and the depth curve is systematically lower than the real data for most of the observation window. Since 'closely replicate' is a headline claim of the paper, this level of fit needs a quantitative threshold or a benchmark comparison (e.g., against a rule-based baseline) before the claim is supported. Reporting per-metric NRMSE and clearly separating the depth limitation from the scale/breadth fit would also help.","section":"§3.3.1, Fig. 4"}],"minor_comments":[{"comment":"The text contains an unresolved cross-reference, 'as described in Section ??'; please fix it.","section":"§3.5"},{"comment":"The caption begins 'TThe figure shows...'; the extra 'T' should be removed.","section":"Fig. 9 caption"},{"comment":"The sentence defining the variables repeats 'yi_simu, yi_simu' where the second variable should presumably be 'yi_real'.","section":"§F.2.1, Eq. (7)"},{"comment":"The protagonist's name is spelled inconsistently as both 'Halen' and 'Helen' across the text and figure.","section":"§3.3.1 and Fig. 5"},{"comment":"There are minor typos: 'Appenix' instead of 'Appendix' and 'Helpfullness' instead of 'Helpfulness'.","section":"§3.2 and F.3.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is publishable in principle as a systems and demonstration paper, but the headline scale findings need to be re-run with proper controls or substantially softened. The editor may want to ensure that the revision is reviewed by someone with experimental-design expertise in multi-agent simulation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real news here is the artifact. OASIS is an open-source, modular LLM-agent simulator that runs a million agents across X-like and Reddit-like environments, with dynamic follow networks, a recommendation system, a time engine, and a 21-action space. That combination is new relative to S3, HiSim, AgentScope, and AgentTorch, and the release could genuinely seed a lot of follow-up work in computational social science. The ablations on RecSys and temporal features are also reasonably done, and they at least attempt reproducibility with repeated runs on selected topics. Credit where it's due: this is a serious engineering contribution.\n\nThe science is softer, and the softest spot is exactly the one the stress-test note flags. Findings 4 and 5 claim that larger agent populations produce more diverse, helpful, and herd-prone behavior. But in Sections 3.4.1 and F.4.2, agent count is varied together with injected post count (30/300/3000) and recommendation cache size (50/500/5000). So the observed effects can be explained by more content and broader exposure, not by population size alone. The paper reports no condition that fixes volume while varying N, or vice versa. That means the headline claim, repeated in the abstract, is not actually established. This is a fixable flaw, but it is load-bearing, and the prose overstates what the experiments show.\n\nTwo other concerns, in proportion. The herd effect is partially built into the environment: agents see like/dislike counts and the Reddit RecSys ranks by hot score, so herding is arguably an artifact of the observation model rather than an emergent social behavior. And the one-million-agent misinformation comparison uses a hand-set TF-IDF threshold with no statistical test, so that section reads as a demonstration rather than a finding. The information-spreading fit at roughly 30% normalized RMSE with consistently lower depth is honest, and the paper admits the depth discrepancy, which is good, but it does mean the replication claim is qualitative.\n\nThe paper's self-citations are fine; nothing there bothers me. The limitations section is also unusually candid about RecSys simplifications and scalability costs.\n\nBottom line: this paper deserves a serious referee and likely publication after major revision, but the revision needs to add proper controls for the scale experiments, or explicitly reframe Findings 4 and 5 as observations about scaling the whole system, not about population size. I'd send it to review.","headline":"A genuinely useful open-source simulator for million-agent social media studies, but the headline scale findings are confounded and need control conditions before they can be stated as results.","tokens_in":751,"tokens_out":734,"would_cite":true,"duration_ms":18756,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims OASIS, a modular LLM-agent simulator, scales to one million users and reproduces information spreading, group polarization, and herd effects on X and Reddit, while larger populations yield more diverse and helpful…","keywords":["agent-based simulation","large language models","social media","information propagation","group polarization","herd effect","misinformation","scalability"],"falsifier":"Run OASIS with a fixed agent count of 1,000 while sweeping the number of posts injected per time step (30, 300, 3000) and the recommendation cache size (50, 500, 5000). If the diversity and helpfulness of the 196 core users' opinions rise with content volume alone, the paper's 'larger group' conclusion is called into question. Alternatively, fix content volume and cache size while sweeping agent count from 196 to 10,196; if helpfulness and diversity do not rise, the population-size claim fails.","tokens_in":26365,"feed_emoji":"🤖","tokens_out":5244,"duration_ms":45848,"temperature":0.7,"pith_summary":"OASIS is a modular simulator for social media platforms that replaces rule-based agents with large language models and scales to a million users. The paper's central claim is that this simulator is general enough to reproduce established social phenomena—information spreading and group polarization on X, and herd effects on Reddit—while being scalable enough to support large-population experiments. It further claims that the number of agents qualitatively changes the simulation: with more agents, the opinions reaching a fixed set of core users become more diverse and more helpful, and herd responses to disliked content become stronger. If these claims hold, OASIS offers social scientists a controllable testbed for questions that are impractical or unethical to run on real platforms.","feed_headline":"One million AI agents re-create social media dynamics","feed_subtitle":"OASIS reproduces information spreading, polarization, and herd effects, and links scale to opinion quality.","key_machinery":"The load-bearing components are the five modules of OASIS. The Environment Server stores users, posts, comments, and relations in a database. The RecSys controls what each agent sees: on X it ranks in-network posts by likes and out-of-network posts by recency, follower impact, and TwHIN-BERT cosine similarity of interests; on Reddit it ranks by the platform's hot-score formula. The Agent Module gives each LLM-backed user a memory and 21 action types, with chain-of-thought reasoning. The Time Engine activates each agent according to a 24-dimensional hourly activity profile, and the Scalable Inferencer distributes inference requests across GPUs asynchronously. A large-scale user-generation algorithm builds networks that preserve the scale-free structure of real social graphs. Together these parts make the simulator platform-adaptable and population-scalable.","core_discovery":"On its own terms, the paper demonstrates that a single architecture with five components—an environment server, platform-specific recommendation systems, an LLM-based agent module, a time engine, and a scalable inference layer—can simulate up to one million agents across X and Reddit. Using real-world data from Twitter15/16, Reddit, and counterfactual posts, OASIS reproduces real propagation trends with a normalized RMSE around 30%, produces increasing group polarization over 80 time steps (more strongly with uncensored models), and replicates the human herd effect, while finding that agents are more conformist than humans on down-treated content. In scale experiments, the same 196 core users receive more diverse and more helpful opinions when the surrounding population grows from 196 to 10,196 and then to 100,196 agents, and herd effects on counterfactual posts appear only above a certain group size.","pith_inferences":["The paper's scale-ablation design increases post-injection volume and recommendation cache size together with agent count, so the reported gains in diversity and helpfulness may be driven by more content or broader exposure rather than by population size per se; an editorially proposed control experiment would hold content volume fixed while varying only the number of agents.","The finding that agents herd on dislikes but not likes, unlike humans, could reflect the LLM's prior about negative signals, and would be worth testing across different base models to see whether the asymmetry is a property of the simulator or of the particular model.","OASIS's architecture lends itself to studying interventions—modifying the recommendation algorithm or content visibility—in a way that would be hard to test on real platforms, a direction the paper notes as untapped potential."],"forward_implications":["Social-science findings replicated in small groups could be re-tested at population scale, where the paper claims emergent behavior appears only with enough agents.","Researchers can swap modules to compare platform designs, since OASIS supports both interest-based (X) and hot-score-based (Reddit) recommendation systems.","The scale effect on helpfulness suggests that simulations with hundreds of agents may underestimate the quality and variety of opinions a core user receives.","The million-agent misinformation experiment shows a tractable path for studying how false content spreads relative to official news, with misinformation generating more related posts."],"supporting_citations":[{"why":"Supplies the real-world information propagation phenomenon and evaluation metrics (scale, depth, breadth) that the X experiments replicate.","marker":"(Vosoughi et al., 2018)"},{"why":"Twitter15 rumor dataset provides propagation instances and user data for the X information-spreading simulations.","marker":"(Liu et al., 2015)"},{"why":"Twitter16 dataset complements Twitter15 for the propagation experiments.","marker":"(Ma et al., 2016)"},{"why":"Randomized human social-influence experiment whose design OASIS replicates for the Reddit herd-effect comparison.","marker":"(Muchnik et al., 2013)"},{"why":"Provides the Reddit hot-score ranking algorithm implemented in OASIS's RecSys.","marker":"(Salihefendic, 2015)"},{"why":"TwHIN-BERT, the socially pretrained embedding model used for interest-based post recommendations on X.","marker":"(Zhang et al., 2023)"},{"why":"Safe RLHF benchmark whose criteria GPT-4o-mini uses to judge helpfulness of agent opinions.","marker":"(Dai et al., 2023)"},{"why":"Counterfactual content dataset used to test herd effects on misinformation-like posts.","marker":"(Meng et al., 2022)"},{"why":"Generative agents paradigm that informs the memory, reasoning, and time-step simulation approach.","marker":"(Park et al., 2023)"}],"fun_headline_variants":["OASIS runs million-agent simulations of X and Reddit","Scale fuels diversity and herd effects in OASIS sims","Million-agent OASIS reproduces polarization and herd","OASIS: from 196 to a million agents, opinions expand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that a larger agent population improves opinion diversity and helpfulness assumes that the concurrent increases in the number of posts injected per time step and the size of the recommendation cache are not the actual causes of those improvements.","fun_headline_variants_meta":{"raw":{"variants":["OASIS runs million-agent simulations of X and Reddit","Scale fuels diversity and herd effects in OASIS sims","Million-agent OASIS reproduces polarization and herd","OASIS: from 196 to a million agents, opinions expand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1850,"prompt_tokens":1017,"completion_tokens":833,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":761}},"tokens_in":633,"tokens_out":833,"duration_ms":8236,"temperature":1.0,"reasoning_tokens":761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:20:51.906139+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run OASIS with a fixed agent count of 1,000 while sweeping the number of posts injected per time step (30, 300, 3000) and the recommendation cache size (50, 500, 5000). If the diversity and helpfulness of the 196 core users' opinions rise with content volume alone, the paper's 'larger group' conclusion is called into question. Alternatively, fix content volume and cache size while sweeping agent count from 196 to 10,196; if helpfulness and diversity do not rise, the population-size claim fails.","supporting_citations":[],"review_version":1}