{"id":"dbc0ab4a-5271-42db-a1a7-61413110ee92","arxiv_id":"2606.29722","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large-scale analysis of 3.1 million posts shows AI agent sub-communities on Moltbook develop distinct linguistic identities through selective attraction and differential retention, not individual adaptation.","lead":"AI agents on a large simulated social platform form distinct linguistic communities mainly by attracting newcomers who already match the group's style and keeping those who fit, rather than by agents changing how they talk over time. This pattern, if general, could shape how future multi-agent AI systems are built and moderated.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Semantic similarity and vocabulary measures may capture topic or embedding artifacts rather than linguistic identity","rationale":"The reader's weakest assumption directly identifies the measurement-validity risk that would falsify the attraction/retention mechanism. Because the full methods, embedding details, and robustness checks are not visible in the supplied abstract, this remains the single load-bearing gap; no other internal inconsistency is detectable from the given material.","tokens_in":1763,"tokens_out":339,"duration_ms":24307,"concrete_test":"Recompute the within-submolt semantic similarity time series and between-submolt vocabulary divergence using a second, independent embedding model (e.g., switch from the paper's reported model to all-MiniLM-L6-v2) on the same Moltbook Observatory Archive subset; if the reported temporal trends or stable-cohort null result change sign or lose significance, the linguistic-identity interpretation is embedding-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that within-submolt semantic similarity growth and between-submolt vocabulary divergence index genuine community-specific linguistic identities formed via attraction/retention. This rests on the untested assumption that the chosen embeddings and divergence metrics are insensitive to topic substructure within submolts, platform vote mechanics, or agent posting volume. The stable-cohort result (no convergence among long-tenured agents) inherits the same measurement; if the similarity metric is partly driven by shared topical content rather than style or lexicon, the \"no adaptation\" conclusion does not follow. No independent validation (e.g., human annotation of linguistic features or alternative embeddings) is described in the provided abstract-level summary.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that AI agent communities on the Moltbook platform develop distinct linguistic identities not through adaptation but via selective attraction (newcomers arrive already compatible) and differential retention (conforming agents stay active longer). Evidence from 3.1M posts by 179k agents across 8683 submolts shows rising within-submolt semantic similarity, increasing between-submolt vocabulary divergence, no convergence in stable cohorts of long-tenured agents, higher vote engagement for semantically aligned posts (vanishing under placebo), and faster effects in smaller submolts.","tokens_in":1890,"tokens_out":658,"duration_ms":45919,"significance":"If the semantic similarity and vocabulary metrics validly index linguistic style independent of topic, the distinction between attraction/retention and adaptation would be a substantive contribution to multi-agent systems research, with design implications for autonomous platforms. The stable-cohort analysis and placebo controls are strengths that help isolate the proposed mechanisms.","major_comments":[{"comment":"The central claim that differentiation occurs through attraction and retention (rather than adaptation) rests on the stable-cohort result showing no linguistic convergence among long-tenured agents. This interpretation is load-bearing but depends on the unvalidated assumption that the semantic similarity metric isolates style/lexicon from topical content or embedding artifacts; no independent validation (alternative embeddings, human annotation of linguistic features) is described.","section":"stable-cohort analysis and semantic similarity measurement"},{"comment":"The reinforcement channel (aligned posts receive higher votes) and its placebo control are presented as supporting evidence, but without details on placebo construction, how vote scores are normalized for posting volume or sub molt size, or whether topic is held constant, it is unclear whether the association isolates linguistic compatibility from other engagement drivers.","section":"reinforcement channel and placebo controls"},{"comment":"Community-size moderation (smaller submolts converge faster) is reported as a key qualifier, yet the manuscript does not specify the exact regression specification, controls for sub molt age or activity level, or robustness checks that would confirm this is not an artifact of smaller samples having noisier similarity estimates.","section":"community-size moderation analysis"}],"minor_comments":[{"comment":"The abstract states both a 100-day window and an 18-week observation period; these should be reconciled with a single consistent time frame and any sensitivity to window length reported.","section":"Abstract"},{"comment":"Clarify the precise definition of 'vocabulary divergence' (e.g., which divergence measure, top-k terms, or embedding-based) and report inter-submolt vs. within-submolt baselines to aid interpretation.","section":"vocabulary divergence results"}],"recommendation":"major_revision","confidential_remarks":"The provided abstract-level summary leaves methods details (embedding model, exact similarity computation, cohort construction) underspecified, which is the source of the low confidence; if the full manuscript contains these, the evaluation could shift. The work fits cs.SI scope but would benefit from explicit discussion of how results generalize beyond the Moltbook environment."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive report. The comments identify areas where additional methodological transparency would strengthen the paper. We respond to each major comment below, indicating revisions where appropriate. We believe the core distinction between attraction/retention and adaptation remains supported by the stable-cohort and placebo results, but we agree that fuller specification of the analyses is warranted.","responses":[{"response":"We acknowledge that the manuscript does not include explicit independent validation of the semantic similarity metric (e.g., via alternative embeddings or human annotation of style features). The submolts are topical by construction, which provides some control for content, and the key pattern—no convergence within stable long-tenured cohorts while overall within-submolt similarity rises—helps isolate the mechanism. Nevertheless, we agree this assumption merits further support. In revision we will add a dedicated robustness subsection that reports results with at least one alternative embedding model and discusses potential embedding artifacts. We will also note the practical limits on large-scale human annotation for this dataset.","revision_made":"partial","referee_comment":"[stable-cohort analysis and semantic similarity measurement] The central claim that differentiation occurs through attraction and retention (rather than adaptation) rests on the stable-cohort result showing no linguistic convergence among long-tenured agents. This interpretation is load-bearing but depends on the unvalidated assumption that the semantic similarity metric isolates style/lexicon from topical content or embedding artifacts; no independent validation (alternative embeddings, human annotation of linguistic features) is described."},{"response":"We agree that the current description of the placebo analysis and vote normalization is insufficiently detailed. The manuscript states that the association vanishes under placebo controls, but does not fully specify construction, normalization, or topic controls. In the revised version we will expand the methods and results sections to provide the exact placebo procedure, the normalization approach (including any adjustments for sub molt size and posting volume), and confirmation that the analysis is conducted within submolts to hold topic constant.","revision_made":"yes","referee_comment":"[reinforcement channel and placebo controls] The reinforcement channel (aligned posts receive higher votes) and its placebo control are presented as supporting evidence, but without details on placebo construction, how vote scores are normalized for posting volume or sub molt size, or whether topic is held constant, it is unclear whether the association isolates linguistic compatibility from other engagement drivers."},{"response":"The manuscript reports that community size moderates the rate of convergence but does not present the full regression specification or robustness checks. We will add the precise model equation (including the size interaction term), the full list of controls (sub molt age, activity level, and total post volume), and results from robustness analyses that address potential noise in small samples (e.g., sample restrictions and weighted specifications). These details will appear in the main text or supplementary materials of the revision.","revision_made":"yes","referee_comment":"[community-size moderation analysis] Community-size moderation (smaller submolts converge faster) is reported as a key qualifier, yet the manuscript does not specify the exact regression specification, controls for sub molt age or activity level, or robustness checks that would confirm this is not an artifact of smaller samples having noisier similarity estimates."}],"tokens_in":1500,"tokens_out":692,"duration_ms":50136,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that differentiation in these agent communities comes from who arrives already matching the group's language and who stays active, not from agents shifting their own output over time. The stable-cohort result is the key piece: long-tenured agents show no convergence, while overall within-submolt similarity rises and between-submolt vocabulary diverges.\n\nThe work does a few things cleanly. The dataset is large and controlled enough to track 179k agents over 100 days with explicit placebo tests on the vote-engagement link. The moderator finding on smaller submolts converging faster adds a testable pattern. The design directly targets the attraction-versus-adaptation contrast that prior work on norms often leaves mixed.\n\nThe soft spot is measurement. Semantic similarity and vocabulary divergence could track topic overlap or embedding artifacts more than style or lexicon, and the abstract gives no sign of human validation, alternative embeddings, or topic-controlled robustness checks. If that holds, the no-adaptation conclusion rests on thinner ground than the cohort split suggests. The reinforcement channel is interesting but inherits the same issue.\n\nThis is for researchers studying multi-agent coordination and emergent norms on platforms. The question is timely and the observational design with controls is coherent, so it deserves a serious referee to examine the metrics and any unreported robustness steps.","headline":"The paper's main claim is that AI agent communities differentiate linguistically via selective attraction and retention rather than adaptation, shown through stable-cohort analysis and placebo controls on a large Moltbook dataset.","tokens_in":2373,"tokens_out":347,"would_cite":false,"duration_ms":31648,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"AI agent communities develop distinct linguistic identities through selective attraction of compatible newcomers and retention of conforming members rather than agents adapting their language.","keywords":["AI agents","linguistic identity","community differentiation","selective attraction","differential retention","multi-agent platforms","social simulation","vocabulary divergence"],"falsifier":"A check showing whether within-forum semantic similarity remains after controlling for the exact topics discussed or after applying different embedding techniques to the same posts.","tokens_in":2662,"feed_emoji":"🤖","tokens_out":574,"duration_ms":27172,"temperature":0.7,"pith_summary":"The paper studies thousands of AI agents interacting on a simulated social platform with over 3 million posts across thousands of topical forums. It finds that agents within each forum become more semantically similar over time while the overall platform diversifies and vocabularies between forums diverge. The mechanism is not agents changing their language to fit in; instead, new agents arrive already aligned with the community's style, and those who match stay active longer. Posts that fit the community's center also receive higher engagement, reinforcing the pattern, with smaller forums showing faster differentiation.","feed_headline":"AI agents sort into linguistic communities by selection, not adaptation","feed_subtitle":"Newcomers join already matching forums and conforming agents stay active, producing distinct vocabularies over 18 weeks.","key_machinery":"selective attraction and differential retention, where newcomers join already matching communities and conforming agents persist longer","core_discovery":"Community-level linguistic differentiation operates through selective attraction—newcomers arrive already linguistically compatible with their chosen community—and differential retention—conforming agents remain active longer. Long-tenured agents do not converge linguistically over time, and the differentiation is supported by higher vote scores for semantically aligned posts.","pith_inferences":["Platform designers could influence linguistic patterns more through entry filters or onboarding than through ongoing moderation of agent behavior.","Similar sorting dynamics might appear in other multi-agent systems where agents choose interaction groups based on initial traits.","If topic controls are insufficient, the observed patterns could partly reflect content clustering rather than language style alone."],"forward_implications":["Smaller, specialized forums show faster linguistic convergence.","Posts aligned with a community's linguistic center receive higher engagement scores.","The platform as a whole diversifies even as individual forums develop distinct vocabularies.","Stable long-term members do not drive the differentiation through personal adaptation."],"fun_headline_variants":["AI agents attract to matching linguistic communities","Selection not adaptation creates distinct AI vocabularies","AI agents sort by language fit not behavioral change","Newcomers join AI forums already linguistically aligned"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That measured rises in within-community semantic similarity and between-community vocabulary divergence reflect genuine linguistic identity formation instead of topic content, platform rules, or the embedding methods used.","fun_headline_variants_meta":{"raw":{"variants":["AI agents attract to matching linguistic communities","Selection not adaptation creates distinct AI vocabularies","AI agents sort by language fit not behavioral change","Newcomers join AI forums already linguistically aligned"]},"model":"grok-4.3","cost_usd":0.00446,"raw_usage":{"total_tokens":2158,"prompt_tokens":694,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":44603000,"prompt_tokens_details":{"text_tokens":694,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1409,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":694,"tokens_out":55,"duration_ms":23659,"temperature":1.0,"reasoning_tokens":1409,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T04:25:18.622987+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A check showing whether within-forum semantic similarity remains after controlling for the exact topics discussed or after applying different embedding techniques to the same posts.","supporting_citations":[],"review_version":1}