{"id":"b9e3e36b-178a-4388-a39f-d3d949485027","arxiv_id":"2505.07646","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"The study reports Granger-causal associations between toxicity and structural divergence for some ideologically opposed Twitter clusters, including within the anti-mandate camp.","lead":"This paper tracks how Twitter users' retweet behavior and the toxicity of their posts evolved during the Covid vaccination and Ukraine war debates. It finds clusters of users that diverge structurally, and in a few cases, toxicity in one cluster appears to predict structural changes in another.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Granger-causality evidence in Table 2 may be an artifact of overlapping 7-day windows; a non-overlapping reanalysis is required before the central polarization claim can be accepted.","rationale":"The reader's weakest_assumption concerned the ideological-term filter. That is a legitimate construct-validity question, but it is not the most load-bearing point: the qualitative interpretation is supported by hashtag log-odds and sampled posts, and the filter affects cluster composition rather than the validity of the causal test. The Granger results are the only quantitative support for the headline temporal-association claim, and they are computed on series with severe induced autocorrelation due to overlapping windows. Overlapping windows are common for smoothing, but using them as if they were independent observations in a Granger test invalidates the p-values. This is a standard statistical correctness issue, not a disagreement with the field's consensus. The proposed non-overlapping reanalysis is inexpensive and decisive: if it reproduces the three significant pairs, the central claim stands with stronger support; if not, the paper should be revised to reframe its conclusions as qualitative and exploratory. I therefore keep a conditional verdict, with the condition now placed on the Granger inference rather than only on the term filter.","tokens_in":14858,"tokens_out":8535,"duration_ms":91440,"concrete_test":"Recompute the full Section 4.4 Granger analysis using non-overlapping 7-day windows (disjoint weeks, yielding roughly 71 weekly observations for the Covid sample and 44 for the Ukraine sample) with the same detrending, lag search, and Bonferroni threshold. If C1,C4, C3,C4, and U1,U4 are no longer significant at the corrected threshold, the reported temporal associations are artifacts of the overlapping-window construction. A complementary placebo: simulate independent daily series with common event shocks, aggregate them by overlapping windows, and measure the false-positive rate of the original procedure.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim—that toxicity and structural dissimilarity are temporally associated for C1,C4 and C3,C4 in the Covid sample and U1,U4 in the Ukraine sample—rests on the Granger tests in Section 4.4 (Table 2). The input series are not independent observations: Section 3.1 constructs a new 7-day window for each calendar day, so consecutive windows share six of seven days. Every per-cluster toxicity and structural-dissimilarity series is therefore a moving average of daily data. Granger causality uses an autoregressive F-test whose null distribution assumes no such induced autocorrelation; with overlapping windows, an event on one day enters both the current and next window for both variables, so lagged toxicity can appear to 'predict' structure with no real temporal order. The rolling OLS detrending in Section 3.3 removes trends but not this moving-average dependence, and no Newey-West, HAC, or block-bootstrap correction is reported. The Bonferroni threshold does not fix the invalid null distribution. The qualitative motifs in Figures 2 and 3 remain suggestive, but the quantitative evidence for polarization dynamics—including the within-camp C3,C4 result—is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript develops a dynamic, time-resolved approach to measuring polarization in Twitter debates about Covid-19 vaccination and the Ukraine war. Retweet behavior is embedded in a sequence of overlapping 7-day windows, reduced by SVD, and clustered with HDBSCAN; cluster-level time series of structural dissimilarity and Perspective-API toxicity are then analyzed with Granger causality. The paper reports significant temporal associations for cluster pairs C1,C4 and C3,C4 in the Covid sample and U1,U4 in the Ukraine sample, interpreting these as evidence of dynamic polarization, including polarization within a nominally aligned anti-mandate camp. The qualitative reading of hashtags and exemplary posts is used to label clusters and to contextualize the statistical findings.","tokens_in":15157,"tokens_out":4150,"duration_ms":43433,"significance":"If the Granger results survive the statistical concerns raised below, the paper would make a useful contribution by moving beyond static polarization measures and by proposing that polarization can be observed as a dynamic, time-dependent process, potentially even within a single ideological camp. The analysis pipeline is transparent and the two datasets are large and topically relevant, which are strengths. The paper also gives credit where due to the complexity of toxicity measurement across languages and dialects in its Limitations section. However, the central quantitative claims rest on Granger p-values that are currently invalidated by the overlapping-window construction and by the unadjusted search over lags; these issues must be addressed before the temporal-association results can be considered established.","major_comments":[{"comment":"The Granger tests are applied to time series built from overlapping 7-day windows, with one window starting on each calendar day. Each observation is therefore a moving average of the previous seven days, which induces strong serial dependence by construction. The standard Granger F-test assumes observations are sampled in a way that does not create spurious autocorrelation; with overlapping windows, a single event enters both the predictor and outcome series in multiple adjacent windows, so cross-variable predictability can appear without any true temporal ordering. The rolling OLS detrending described in §3.3 removes trends but not this overlap-induced moving-average dependence, and no Newey-West, HAC, or block-bootstrap correction is reported. A reanalysis with non-overlapping windows, or an explicit correction for the induced dependence, is required before the p-values in Table 2 can support the paper's central polarization claims.","section":"§3.1 and §4.4, Table 2"},{"comment":"For each test the authors search over lags up to 155 (Covid) or 85 (Ukraine) observations and report the minimum p-value over that search; for example, Table 2 reports C1,C2 toxicity-to-structure at lag 51, C1,C4 at lag 43, and U4,U5 at lag 83. Because the minimum over a large grid is reported without correcting for the number of lags tried, the quoted p-values are not valid as probabilities of a false positive under the null, even before the Bonferroni correction across cluster pairs. The authors should either fix lags a priori, apply a correction over the lag grid, or use a data-driven lag-selection procedure with post-selection inference. This issue affects every entry in Table 2, including the headline C1,C4, C3,C4, and U1,U4 results.","section":"§4.4, Table 2"},{"comment":"The aggregate toxicity Tc is defined as 1 - product over i in c of [1 - Fi·Ti], but Fi is never defined in the manuscript. The reader cannot tell whether Fi is a post frequency weight, a normalizer, or something else, and the reproducibility of the toxicity time series depends on this quantity. Please define Fi explicitly and state how posts with duplicate text, missing scores, or multiple authors are handled when forming the cluster-level series.","section":"§3.2, Eq. (3)"},{"comment":"The Covid sample is truncated to the first 500 observations after inspecting the series, described as a 'conservative approach' to avoid the low-activity final 231 observations. Because this clipping decision is made on the same data that are then tested, it is a post hoc selection that can alter the Granger results, and the paper reports no sensitivity analysis (for example, results on the full series, alternative cutoffs, or a pre-registered criterion). Please justify the cutoff on a priori grounds or show that the substantive findings do not depend on this choice.","section":"§2 and §4.4"},{"comment":"The ideological engagement filter is based on a term list (nazism, holocaust, holodomor, etc.) and on per-sample thresholds (28 posts for Ukraine, 7 for Covid) taken from prior work, but the manuscript does not validate that users meeting this criterion are representative of the ideological camps in the two debates. If the term list selects a narrow or unusual subset of users, the clusters and the resulting Granger tests may reflect a specific subpopulation rather than the broader polarization dynamics. Please provide robustness evidence (for example, comparing cluster structure and key Granger results with and without the filter, or showing construct validity by contrasting included and excluded users' hashtags and sharing patterns).","section":"§2, Table 1"}],"minor_comments":[{"comment":"Equation (2) defines structural dissimilarity as a negative cosine similarity; the en dash is presumably a minus sign, but as written the quantity can take negative values, which is an unusual dissimilarity measure. Please clarify the intended definition (e.g., 1 - cos) and its range.","section":"Eq. (2)"},{"comment":"The sentence 'we test for temporal dependence between our variables in both directions, up to a lag of 155 observations in the Covid sample and 85 observations in the Ukraine sample' appears immediately after the clipping description; please state how the maximum lags were chosen and whether the lag grid is in days or in window steps.","section":"§4.4"},{"comment":"The captions mention shaded areas and a gray dashed vertical line for the clipping point, but the figures themselves would benefit from a legend or explicit annotation identifying which curve is toxicity and which is structural dissimilarity, as well as the meaning of the shaded regions.","section":"Figures 2 and 3"},{"comment":"There are several typographical issues, including 'Unversity' in the OSoMe reference, 'inquerie' instead of 'inquiry' in §4.4, and the duplicated 'and and' in §4.3; these should be corrected in a final pass.","section":"References and text"}],"recommendation":"major_revision","confidential_remarks":"The paper is interesting and the qualitative interpretation is careful, but the core Granger evidence is not yet statistically valid because of the overlapping-window autocorrelation and the unadjusted lag search. These are fixable with a reanalysis, so I do not recommend rejection, but the authors should be asked to rerun the temporal analyses with non-overlapping windows or a valid correction, and to address the post hoc clipping and the undefined Fi before the manuscript can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper deserves a serious referee, but the quantitative core needs to be redone. What is genuinely new is the attempt to track polarization as a moving process: windowed SVD to follow cluster positions in retweet space, paired with per-cluster toxicity, and then Granger causality to look for temporal dependencies. The within-camp polarization finding (C3/C4 in the Covid sample, both anti-mandate) is the most interesting idea, and the qualitative reading of clusters is careful and grounded in example tweets. The datasets are large and the methods are described transparently. That is real credit. The soft spot is load-bearing. The Granger tests in Table 2 use time series built from overlapping 7-day windows, with a new window each calendar day. That makes each observation a moving average of the previous six days. The Granger null assumes no such induced autocorrelation, and the paper reports no Newey-West, HAC, or block-bootstrap correction. An event on one day enters several consecutive windows for both variables, so lagged relationships can appear without any true temporal ordering. The Bonferroni threshold only corrects for the number of cluster pairs and directions, not for the lag search, which goes up to 155 lags with the minimum p-value reported per test. That is multiple testing on top of the autocorrelation problem, so the significant pairs in Table 2 are not trustworthy as evidence for polarization dynamics. The other issues are smaller but worth noting: Eq. 3 leaves Fi undefined; the ideological-term filter is ad hoc and borrowed from prior work; and clipping the Covid sample to the first 500 observations after seeing the data is post hoc, though it is at least in a conservative direction. None of these alone would sink the paper, and the descriptive time-series motifs in Figures 2 and 3 may still be informative. But the central claim rests on the Granger results, and those are not yet established. Who gets value? Social-media polarization researchers who want a template for dynamic measurement, and methodologists who want an example of how overlapping windows can quietly invalidate a causal test. I would not cite it as evidence for polarization dynamics until the Granger analysis is rerun on non-overlapping windows (or with a proper autocorrelation-robust procedure) and the lag search is accounted for. Send it out to a statistically careful reviewer, but the expected verdict should be major revision, not acceptance.","headline":"The paper's descriptive dynamic-polarization setup is worth a look, but its central Granger-causality evidence is undermined by overlapping windows and an unadjusted lag search.","tokens_in":763,"tokens_out":977,"would_cite":false,"duration_ms":29084,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hostility and drifting retweet patterns are temporally linked inside both the Covid-vaccine and Ukraine-war Twitter debates, including between two camps that should be allies.","keywords":["polarization dynamics","Granger causality","toxicity","retweet networks","Covid-19 vaccination debate","Ukraine war debate","principal component analysis","HDBScan clustering"],"falsifier":"Re-run the entire pipeline with the ideological-engagement filter replaced by an independent measure, for example a user's inferred ideological position from the set of politicians and news outlets they follow, and check whether the Bonferroni-significant Granger relationships for C1–C4, C3–C4, and U1–U4 still appear. If they disappear, the key results are an artifact of the historical-term filter rather than of polarization in the broader debate.","tokens_in":14691,"feed_emoji":"⚔️","tokens_out":6846,"duration_ms":58915,"temperature":0.7,"pith_summary":"Polarization is usually read off a static network snapshot, but this paper argues that it is a moving process: users' retweeting priorities shift over time, and hostility between camps can be both a cause and an effect of those shifts. Analysing large Twitter corpora on Covid-19 vaccination and the Ukraine war, the authors reduce retweet behaviour to a few principal components, cluster users into ideological camps, and track each camp's weekly centroid position and the toxicity of its posts. Using time-series tests of temporal precedence, they find that hostility in one cluster predicts structural drift between clusters, and structural drift predicts hostility, for the American pro-mandate and anti-mandate Covid clusters, for two anti-mandate clusters that should be allies, and for a French-speaking pro-Russia cluster and a pro-Ukraine cluster in the Ukraine data. If these relationships are real, polarization is not just a fixed split but a dynamic, multi-process phenomenon that can even occur within a single ideological camp.","feed_headline":"Toxicity can predict splits inside a Twitter camp","feed_subtitle":"Granger-style tests link toxic posts to drifting retweet patterns in Covid and Ukraine debates.","key_machinery":"The machinery has four linked parts. First, retweet behaviour is encoded in per-week incidence matrices and reduced by singular value decomposition, leaving a low-dimensional information diffusion space in which each user has a position. Second, HDBScan clusters users in that space into ideological camps, and each cluster's weekly centroid gives a time series of structural position. Third, structural dissimilarity between two clusters is defined as the negative cosine of their centroids, and a cluster's toxicity as the probability that a randomly engaged reader encounters a toxic post, computed from an automated toxicity-scoring model, the Perspective API. Fourth, Granger causality tests ask whether lagged values of one de-trended series improve predictions of another; a significant test means toxicity and structural distance are temporally linked, not merely correlated. The central load-bearing object is the Granger test between cluster toxicity and structural dissimilarity, because it converts evolving positions and hostility into evidence about polarization dynamics.","core_discovery":"On the paper's own terms, the discovery is that toxicity and structural divergence in retweet behaviour are temporally coupled for specific cluster pairs in both debates. In the Covid sample, the ideologically opposed American clusters C1 and C4 show Granger causality in both directions between their combined toxicity and their structural dissimilarity; unexpectedly, so do the aligned anti-mandate clusters C3 and C4, which exhibit the strongest structural dissimilarity of any Covid pair. In the Ukraine sample, the French-speaking pro-Russian cluster U1 and the pro-Ukrainian cluster U4 show significant Granger relationships for all four pairwise tests. The paper interprets these results as evidence that affective hostility and network separation reinforce each other over time, and that ideological alignment does not prevent internal polarization when commitment to a shared framework is uneven.","pith_inferences":["Inference: a testable extension the authors do not run is to apply the same pipeline to a debate without a major external shock; if the significant Granger links vanish, the detected dynamics may be driven by news events rather than intrinsic inter-group hostility.","Inference: the French U1–U4 pair points to a transnational, pan-European cleavage that the paper only partially interprets; a natural next step is to check whether the same temporal coupling appears in French-language-only retweet networks, independent of the English-language discourse.","Inference: if within-camp polarization is real, then models of echo chambers should treat each side as an internally differentiated set of publics with potentially conflicting commitment levels, not as a single bloc.","Inference: because the significant pairs are also the pairs with the largest structural dissimilarity, one might infer that toxicity becomes temporally coupled to structure mainly once groups have already drifted apart; this ordering hypothesis could be tested by comparing Granger results across early and late windows."],"forward_implications":["Polarization can be measured as a continuous time series from publicly visible retweet structure, without needing to predefine who the political influencers are.","Ideological allies are not automatically stable: the C3–C4 result implies camps can polarize internally as one faction drifts toward more extreme or conspiracy-laden content.","Affective hostility and structural separation can drive each other, since significant Granger relationships run in both directions for the key cluster pairs.","The same measurement approach can be applied to any debate with sustained retweeting, including conflicts beyond the two studied here.","Treating polarization as static may miss the moments when camps actually form, split, or realign, because those are exactly the periods where the time series diverge."],"supporting_citations":[{"why":"Provides the Granger causality test used to establish temporal dependence between toxicity and structural dissimilarity.","marker":"Granger 1969"},{"why":"Supplies the Covid-19 vaccination Twitter dataset analysed in the paper.","marker":"DeVerna et al. 2021"},{"why":"Supplies the Ukraine war Twitter dataset analysed alongside the Covid data.","marker":"Social Media at Indiana Unversity"},{"why":"Source of the secondary query terms used to restrict users to ideologically engaged sets.","marker":"Axelrod, Kim, and Paolillo 2024"},{"why":"Provides the HDBScan clustering algorithm used to identify user camps in the reduced retweet space.","marker":"Malzer and Baum 2020"},{"why":"Provides the Perspective API used to compute per-post toxicity scores.","marker":"Google"},{"why":"Informs the hypothesis that affective states play a role in polarization dynamics.","marker":"N. Druckman et al. 2020"}],"fun_headline_variants":["Toxic tweets predict splits inside aligned Twitter camps","Granger tests link toxicity to retweet drift in debates","Covid and Ukraine debates show toxicity precedes internal splits","Twitter analysis: toxicity and retweet separation are entangled","Even ideologically aligned camps split under toxic pressure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that users who repeatedly post about historical terms like nazism, holocaust, genocide, or communism are the ideologically engaged users whose retweeting defines the relevant camps; if those terms select a narrow or unrepresentative subset, the clusters and their inferred polarization would not represent the broader vaccine or Ukraine debates.","fun_headline_variants_meta":{"raw":{"variants":["Toxic tweets predict splits inside aligned Twitter camps","Granger tests link toxicity to retweet drift in debates","Covid and Ukraine debates show toxicity precedes internal splits","Twitter analysis: toxicity and retweet separation are entangled","Even ideologically aligned camps split under toxic pressure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000977,"raw_usage":{"total_tokens":4154,"prompt_tokens":954,"completion_tokens":3200,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":3125}},"tokens_in":570,"tokens_out":3200,"duration_ms":21313,"temperature":1.0,"reasoning_tokens":3125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:10:22.389088+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the entire pipeline with the ideological-engagement filter replaced by an independent measure, for example a user's inferred ideological position from the set of politicians and news outlets they follow, and check whether the Bonferroni-significant Granger relationships for C1–C4, C3–C4, and U1–U4 still appear. If they disappear, the key results are an artifact of the historical-term filter rather than of polarization in the broader debate.","supporting_citations":[],"review_version":1}