{"id":"f66514c0-6d03-4455-abab-a224db29a1c2","arxiv_id":"2512.12441","paper_version":5,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In Facebook hate groups, concentrated posting activity is tied to higher engagement; centralized Islamophobic groups show uniform narratives while centralized anti-Semitic groups show diverse ones.","lead":"This study examines ten years of Facebook hate groups and finds that groups where a few users post most of the content get more engagement, and that centralized Islamophobic groups use more uniform messaging while centralized anti-Semitic groups use more varied messaging. It matters because it offers evidence-based entry points for platforms and policymakers trying to disrupt online hate.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3/RQ4 compare 1,820 group-month observations from 24 groups as if independent; without group-level clustering, the headline framing-diversity contrast may not survive.","rationale":"The reader's conditional verdict identifies the right soft spot. The paper's most novel and policy-relevant claim is not the engagement finding—which is consistent with prior work and protected by controls—but the ideological contrast in framing diversity. That contrast is built on repeated monthly observations from very few groups. Independent resampling of group-months will overstate precision whenever observations within a group are correlated; here they almost certainly are, since narrative style and leadership are sticky over time. The appendix checks threshold choice but does not address the dependence structure. A cluster bootstrap or mixed model is the standard fix. If the contrast survives, the paper's headline is much stronger; if not, the paper should be read as reporting an association in need of replication with more groups. I therefore agree with the reader's weakest assumption and see no reason to move the verdict away from CONDITIONAL.","tokens_in":29282,"tokens_out":5506,"duration_ms":56344,"concrete_test":"Run a group-level cluster bootstrap for the central RQ3 comparisons: resample the 24 groups with replacement, keep all monthly observations for sampled groups, and recompute the centralized-vs-decentralized framing-homogeneity Mann-Whitney test separately for Islamophobic and Anti-Semitic groups, with Bonferroni correction across the 12 content-homogeneity tests. If the cluster-bootstrap 95% CI for either median difference includes zero or the adjusted p exceeds 0.004, the headline framing-diversity contrast is unsupported. As a complementary check, fit a mixed-effects model with a random intercept for group on framing Gini.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing claim is the RQ3 contrast in Fig. 4A: centralized Islamophobic groups have higher framing homogeneity than decentralized ones (median 0.55 vs 0.52; p=2.05e-4), while centralized Anti-Semitic groups have lower framing homogeneity (median 0.51 vs 0.57; p<1e-6). These p-values come from bootstrapped Mann-Whitney U tests on 1,820 group-month observations from only 24 groups (Methodology: 'Answering RQ3 and RQ4'; Appendix: 'Additional Analysis Notes'). The bootstrap is not described as clustered by group, and no group-level random effects are reported. Because consecutive months from the same group are strongly autocorrelated and groups differ in stable ways—size, topic focus, moderation history—the effective independent sample size is far below 1,820; for Islamophobic groups it is at most 5 groups. If within-group autocorrelation is substantial, the reported p-values are inflated and the qualitative contrast—centralized Islamophobic groups more uniform, centralized anti-Semitic groups more diverse—could disappear. This is structurally distinct from RQ2, which at least includes controls for serial correlation, making the engagement result less exposed. The appendix checks median vs mean threshold sensitivity but does not address the dependence structure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies longitudinal Facebook data from 24 hate groups (995,716 posts, July 2014–June 2024) related to the Israel–Palestine conflict. It defines 'participation structure' via the monthly Gini coefficient of user posting activity and addresses four research questions: (1) differences in centralization across ideologies; (2) association between centralization and next-month engagement; (3) association between centralization and narrative framing/topic homogeneity; (4) relation between centralization and inter-group homophily. Negative binomial regressions support RQ2: centralization is positively associated with future engagement across ideologies. Bootstrapped Mann–Whitney tests support the RQ3/RQ4 contrasts, including an ideology-dependent pattern in which centralized Islamophobic groups are more homogeneous and centralized Anti-Semitic groups are more diverse.","tokens_in":29639,"tokens_out":6275,"duration_ms":64745,"significance":"If the RQ2 result holds, it provides longitudinal evidence for the club-goods/leadership account in an online hate context and contrasts with leaderless-resistance expectations. The paper also offers a comparative cross-ideology design, a large dataset, and transparent release of code and aggregated statistics. The narrative contrast in RQ3 is novel and practically relevant. The main statistical concern—unclustered inference over repeated group-month observations—limits the strength of the RQ3/RQ4 conclusions and must be addressed before the headline claims are accepted.","major_comments":[{"comment":"The bootstrapped Mann–Whitney U tests pool 1,820 group-month observations from only 24 groups (5 Islamophobic, 14 Anti-Semitic, 5 Other-hate; Table A1). Consecutive months from the same group are not independent, and the bootstrap is not described as clustered by group; no group-level random effects are reported. For the RQ3 headline results (centralized vs. decentralized Islamophobic framing homogeneity: median 0.55 vs 0.52, p=2.05e-4; Anti-Semitic: 0.51 vs 0.57, p<1e-6), within-group autocorrelation can inflate these p-values substantially. Please re-run at the group level (e.g., cluster bootstrap by group, or mixed-effects model with random group intercepts) and report effective sample sizes.","section":"Answering RQ3 and RQ4; Appendix 'Additional Analysis Notes'"},{"comment":"The eight-frame taxonomy is claimed to cover 'all cases in our Facebook and StormFront samples' after Grounded Theory coding. The same taxonomy defines the label space of the classifier applied to those posts, so coverage is ensured by construction; the classifier cannot detect frames outside the taxonomy. This makes the homogeneity contrasts in RQ3 partly a test of the coding scheme's fit. Please validate the taxonomy on independently labeled data or show that an independent coding scheme reproduces the same centralized-vs-decentralized differences.","section":"Methodology 'Classification of Narrative Frames'; Table A4"},{"comment":"Annotators selected all applicable frames, but the analysis assigns each post only the top-ranked frame. If many posts contain multiple frames, single-label Gini homogeneity can be an artifact of the ranking/threshold rule. Please report a multi-label robustness check or justify why single-label assignment is appropriate.","section":"Methodology 'Inferring Narrative Framing and Topic'; Codebook"}],"minor_comments":[{"comment":"The caption says only statistically significant factors are shown, but the text describes faded markers for coefficients that are neither statistically nor practically significant. Clarify the inclusion rule.","section":"Fig. 3 caption"},{"comment":"The column '#Groups' actually counts group-month observations, not unique groups. Relabel it as '#Group-months' to avoid ambiguity.","section":"Table A2"},{"comment":"For each bootstrapped Mann–Whitney U test, state the number of bootstrap resamples and the resampling unit (group-month vs. group). This is essential for assessing the inference.","section":"Appendix 'Additional Analysis Notes'"},{"comment":"The frame name 'Moral Justification' is inconsistent with the taxonomy's 'Justification of Violence.' Make terminology consistent throughout.","section":"Table A4"},{"comment":"The R^2 values for negative binomial regressions need a qualifier (e.g., McFadden pseudo-R^2); the current labels are ambiguous.","section":"Results RQ2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely salvageable: the RQ2 engagement result appears sound, while the RQ3/RQ4 headline contrasts require reanalysis with cluster-robust inference. I would not reject based on the current evidence, but the paper should not be accepted before the repeated-measures structure is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look but treat the headline as provisional. The genuinely new asset is the dataset: ten years of Facebook posts from 24 hate groups (995k posts), with a careful distinction between Islamophobic and anti-Semitic groups. The RQ2 result—centralization predicts next-month engagement, with controls for serial correlation and current activity—is plausible and consistent with earlier work. The ideological contrast in RQ3, centralized Islamophobic groups more uniform vs centralized anti-Semitic groups more diverse, is interesting and if true would matter.\n\nWhat the paper does well: RQ2 is specified with negative binomial regressions, VIF checks, and outlier sensitivity. The authors are transparent about limitations and provide code and aggregated statistics. The framing taxonomy is grounded and the classifier F1s are decent.\n\nThe soft spot is exactly where the reader put it. RQ3/RQ4 compare 1,820 group-month observations drawn from only 24 groups, and the Mann-Whitney tests treat those observations as independent. The appendix describes bootstrapped tests but does not describe clustering by group. With 5 Islamophobic groups, the effective sample size for the key contrast is tiny. The p-values (p=2e-4, p<1e-6) are not believable without a block bootstrap or group-level model. The contrast may survive, but the current analysis doesn't show it. This is the load-bearing novel claim, so it needs fixing before publication.\n\nTwo smaller issues: the classifier evaluation is under-specified (test sample counts don't match a standard split of the 940 annotated samples, and no error bars are reported—the authors admit this). And the augmentation with Llama-2 synthetic data could introduce artifacts, though the evaluation suggests decent quality.\n\nWho is this for? Researchers working on online hate group dynamics and platform governance, and anyone interested in repeated-measures pitfalls. It deserves a serious referee, but with a clear request to re-analyze RQ3/RQ4 with group-level clustering or block bootstrap, and to temper the intervention language accordingly. I'd send it to review rather than desk reject; the dataset and RQ2 are real contributions.","headline":"Useful dataset and solid RQ2, but the headline RQ3 contrast rests on unclustered tests over 24 groups.","tokens_in":30124,"tokens_out":2750,"would_cite":true,"duration_ms":29940,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that online hate groups sustain higher engagement when a few dominant users produce most of the content, and that this concentration affects narrative diversity in opposite ways for Islamophobic versus anti-Semitic groups.","keywords":["participation structure","online hate groups","engagement","narrative framing","Gini coefficient","Islamophobia","anti-Semitism","homophily"],"falsifier":"Recompute the centralized-versus-decentralized framing/topic homogeneity and homophily comparisons using cluster-robust standard errors or group-level random effects; if the differences (currently p<10^-6) become non-significant, the narrative contrast is unsupported.","tokens_in":29149,"feed_emoji":"📈","tokens_out":4164,"duration_ms":42312,"temperature":0.7,"pith_summary":"This paper argues that the internal participation structure of online hate groups—whether a few members dominate posting or activity is spread widely—is a key driver of how active these communities stay and what they say. Analyzing ten years of Facebook posts from 24 anti-Semitic, Islamophobic, and other-hate groups, it finds that groups with higher posting centralization attract significantly more likes, comments, and shares in the following month. It also finds an ideological split in messaging: centralized Islamophobic groups become more uniform in their narrative frames and topics, while centralized anti-Semitic groups become more diverse. These results matter because they point to different levers for disrupting different kinds of hate communities: removing central figures may work for coordinated Islamophobic networks, while anti-Semitic networks may require broader, less targeted interventions.","feed_headline":"Hate groups run by a few loud posters get more engagement","feed_subtitle":"A decade of Facebook data links concentrated posting to higher activity—and reveals opposite narrative styles in Islamophobic vs anti-Semiti","key_machinery":"Participation structure is operationalized through the Gini coefficient of per-user posting counts, classifying groups as centralized or decentralized by whether the monthly Gini exceeds the median. Narrative content is coded into an eight-frame extremism taxonomy (Us vs. Them, Heroic, Dehumanization/Demonization, Victimization, Justification of Violence, Legitimacy, Imminent War/Crisis, Religious) using a fine-tuned language model, and content homogeneity is measured as the Gini coefficient of the distribution of posts across frames and topics. Inter-group connectivity is measured with a weighted homophily index based on shared users. These measurements together allow the paper to connect '","core_discovery":"The paper's central claim is that participation centralization—measured as the Gini coefficient of user posting activity—predicts next-month engagement in hate groups across all ideologies studied. In regression models, a one-standard-deviation increase in centralization is associated with increases in engagement of 0.39 for Islamophobic, 0.28 for anti-Semitic, 1.24 for other-hate, and 0.72 overall, all significant at p<0.001. The paper also reports a second finding: centralization relates to narrative uniformity in opposite directions by ideology. Centralized Islamophobic groups show significantly higher framing and topic homogeneity than decentralized ones, while centralized anti-Semitic g","pith_inferences":["If the centralization-engagement link is causal, platform changes that reduce the visibility of a few prolific posters (e.g., per-user posting caps) could lower hate-group engagement without deplatforming entire groups.","The finding that centralized anti-Semitic groups are more diverse might reflect 'curated chaos': leaders amplify multiple frames to appeal to broader coalitions; this could be tested by comparing frame diversity before and after the removal of top accounts.","The homophily asymmetry suggests an asymmetric risk: Islamophobic conspiracy content remains siloed (echo chamber escalation), while anti-Semitic tropes may seed general-purpose misinformation networks; future work could trace cross-ideological user sharing patterns.","The Gini-based measure of 'leadership' conflates authority with volume; a testable extension is to validate against identifiable admin roles or verified accounts."],"forward_implications":["Tactics to disrupt Islamophobic groups should focus on removing or sidelining central posting actors, since their concentration sustains engagement.","For anti-Semitic groups, engagement is sustained even when content is diverse and leadership is less coordinated, so broader strategies addressing both prominent figures and grassroots participants are needed.","Because content homogeneity is associated with lower engagement, groups that keep a single dominant frame or topic tend to lose resonance; diversity may be a deliberate or emergent engagement strategy.","Anti-Semitic groups' lower homophily means their narratives are more likely to reach outside-the-ideology audiences, so cross-ideological bridges deserve monitoring.","The October 2023 conflict escalation coincided with shifts toward more diverse framing in both ideologies, suggesting external events can restructure narrative landscapes even in stable participation structures."],"fun_headline_variants":["Centralized hate groups see higher engagement on Facebook","Loud leaders drive engagement in online hate groups","Hate group centralization boosts engagement, shapes narratives","Centralized hate groups: more engagement, but opposite narrative styles","Leader-driven hate groups: more engagement, distinct narratives"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The narrative and homophily contrasts rest on treating 1,820 monthly group observations from only 24 groups as statistically independent; if activity within each group is autocorrelated month to month, the reported p-values (e.g., p<10^-6) would be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Centralized hate groups see higher engagement on Facebook","Loud leaders drive engagement in online hate groups","Hate group centralization boosts engagement, shapes narratives","Centralized hate groups: more engagement, but opposite narrative styles","Leader-driven hate groups: more engagement, distinct narratives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00031,"raw_usage":{"total_tokens":1615,"prompt_tokens":763,"completion_tokens":852,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":776}},"tokens_in":507,"tokens_out":852,"duration_ms":9543,"temperature":1.0,"reasoning_tokens":776,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:38:01.751810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the centralized-versus-decentralized framing/topic homogeneity and homophily comparisons using cluster-robust standard errors or group-level random effects; if the differences (currently p<10^-6) become non-significant, the narrative contrast is unsupported.","supporting_citations":[],"review_version":1}