Pith. sign in

REVIEW 5 major objections 4 minor 9 references

'Congratulations, morons': Dynamics of Toxicity and Interaction Polarization in the Covid Vaccination and Ukraine War Twitter Debates

T0 review · 5 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Hostility and drifting retweet patterns are temporally linked inside both the Covid-vaccine and Ukraine-war Twitter debates, including between two camps that should be allies.

desk verdict The paper's descriptive dynamic-polarization setup is worth a look, but its central Granger-causality evidence is undermined by overlapping windows and an unadjusted lag search. read the letter →

arxiv 2505.07646 v1 pith:NNQZQVSF submitted 2025-05-12 cs.SI physics.soc-ph

classification cs.SIphysics.soc-ph
keywords polarizationdynamicsGrangercausalitytoxicityretweetnetworksCovid-19vaccinationdebateUkrainewarprincipalcomponentanalysisHDBScanclustering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Polarization is usually read off a static network snapshot, but this paper argues that it is a moving process: users' retweeting priorities shift over time, and hostility between camps can be both a cause and an effect of those shifts. Analysing large Twitter corpora on Covid-19 vaccination and the Ukraine war, the authors reduce retweet behaviour to a few principal components, cluster users into ideological camps, and track each camp's weekly centroid position and the toxicity of its posts. Using time-series tests of temporal precedence, they find that hostility in one cluster predicts structural drift between clusters, and structural drift predicts hostility, for the American pro-mandate and anti-mandate Covid clusters, for two anti-mandate clusters that should be allies, and for a French-speaking pro-Russia cluster and a pro-Ukraine cluster in the Ukraine data. If these relationships are real, polarization is not just a fixed split but a dynamic, multi-process phenomenon that can even occur within a single ideological camp.

What carries the argument

The machinery has four linked parts. First, retweet behaviour is encoded in per-week incidence matrices and reduced by singular value decomposition, leaving a low-dimensional information diffusion space in which each user has a position. Second, HDBScan clusters users in that space into ideological camps, and each cluster's weekly centroid gives a time series of structural position. Third, structural dissimilarity between two clusters is defined as the negative cosine of their centroids, and a cluster's toxicity as the probability that a randomly engaged reader encounters a toxic post, computed from an automated toxicity-scoring model, the Perspective API. Fourth, Granger causality tests ask whether lagged values of one de-trended series improve predictions of another; a significant test means toxicity and structural distance are temporally linked, not merely correlated. The central load-bearing object is the Granger test between cluster toxicity and structural dissimilarity, because it converts evolving positions and hostility into evidence about polarization dynamics.

What would settle it

Re-run the entire pipeline with the ideological-engagement filter replaced by an independent measure, for example a user's inferred ideological position from the set of politicians and news outlets they follow, and check whether the Bonferroni-significant Granger relationships for C1–C4, C3–C4, and U1–U4 still appear. If they disappear, the key results are an artifact of the historical-term filter rather than of polarization in the broader debate.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that toxicity and structural divergence in retweet behaviour are temporally coupled for specific cluster pairs in both debates. In the Covid sample, the ideologically opposed American clusters C1 and C4 show Granger causality in both directions between their combined toxicity and their structural dissimilarity; unexpectedly, so do the aligned anti-mandate clusters C3 and C4, which exhibit the strongest structural dissimilarity of any Covid pair. In the Ukraine sample, the French-speaking pro-Russian cluster U1 and the pro-Ukrainian cluster U4 show significant Granger relationships for all four pairwise tests. The paper interprets these results as evidence that affective hostility and network separation reinforce each other over time, and that ideological alignment does not prevent internal polarization when commitment to a shared framework is uneven.

Load-bearing premise

The paper assumes that users who repeatedly post about historical terms like nazism, holocaust, genocide, or communism are the ideologically engaged users whose retweeting defines the relevant camps; if those terms select a narrow or unrepresentative subset, the clusters and their inferred polarization would not represent the broader vaccine or Ukraine debates.

Editorial extensions

If this is right

  • Polarization can be measured as a continuous time series from publicly visible retweet structure, without needing to predefine who the political influencers are.
  • Ideological allies are not automatically stable: the C3–C4 result implies camps can polarize internally as one faction drifts toward more extreme or conspiracy-laden content.
  • Affective hostility and structural separation can drive each other, since significant Granger relationships run in both directions for the key cluster pairs.
  • The same measurement approach can be applied to any debate with sustained retweeting, including conflicts beyond the two studied here.
  • Treating polarization as static may miss the moments when camps actually form, split, or realign, because those are exactly the periods where the time series diverge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: a testable extension the authors do not run is to apply the same pipeline to a debate without a major external shock; if the significant Granger links vanish, the detected dynamics may be driven by news events rather than intrinsic inter-group hostility.
  • Inference: the French U1–U4 pair points to a transnational, pan-European cleavage that the paper only partially interprets; a natural next step is to check whether the same temporal coupling appears in French-language-only retweet networks, independent of the English-language discourse.
  • Inference: if within-camp polarization is real, then models of echo chambers should treat each side as an internally differentiated set of publics with potentially conflicting commitment levels, not as a single bloc.
  • Inference: because the significant pairs are also the pairs with the largest structural dissimilarity, one might infer that toxicity becomes temporally coupled to structure mainly once groups have already drifted apart; this ordering hypothesis could be tested by comparing Granger results across early and late windows.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The manuscript develops a dynamic, time-resolved approach to measuring polarization in Twitter debates about Covid-19 vaccination and the Ukraine war. Retweet behavior is embedded in a sequence of overlapping 7-day windows, reduced by SVD, and clustered with HDBSCAN; cluster-level time series of structural dissimilarity and Perspective-API toxicity are then analyzed with Granger causality. The paper reports significant temporal associations for cluster pairs C1,C4 and C3,C4 in the Covid sample and U1,U4 in the Ukraine sample, interpreting these as evidence of dynamic polarization, including polarization within a nominally aligned anti-mandate camp. The qualitative reading of hashtags and exemplary posts is used to label clusters and to contextualize the statistical findings.

Significance. If the Granger results survive the statistical concerns raised below, the paper would make a useful contribution by moving beyond static polarization measures and by proposing that polarization can be observed as a dynamic, time-dependent process, potentially even within a single ideological camp. The analysis pipeline is transparent and the two datasets are large and topically relevant, which are strengths. The paper also gives credit where due to the complexity of toxicity measurement across languages and dialects in its Limitations section. However, the central quantitative claims rest on Granger p-values that are currently invalidated by the overlapping-window construction and by the unadjusted search over lags; these issues must be addressed before the temporal-association results can be considered established.

major comments (5)
  1. [§3.1 and §4.4, Table 2] The Granger tests are applied to time series built from overlapping 7-day windows, with one window starting on each calendar day. Each observation is therefore a moving average of the previous seven days, which induces strong serial dependence by construction. The standard Granger F-test assumes observations are sampled in a way that does not create spurious autocorrelation; with overlapping windows, a single event enters both the predictor and outcome series in multiple adjacent windows, so cross-variable predictability can appear without any true temporal ordering. The rolling OLS detrending described in §3.3 removes trends but not this overlap-induced moving-average dependence, and no Newey-West, HAC, or block-bootstrap correction is reported. A reanalysis with non-overlapping windows, or an explicit correction for the induced dependence, is required before the p-values in Table 2 can support the paper's central polarization claims.
  2. [§4.4, Table 2] For each test the authors search over lags up to 155 (Covid) or 85 (Ukraine) observations and report the minimum p-value over that search; for example, Table 2 reports C1,C2 toxicity-to-structure at lag 51, C1,C4 at lag 43, and U4,U5 at lag 83. Because the minimum over a large grid is reported without correcting for the number of lags tried, the quoted p-values are not valid as probabilities of a false positive under the null, even before the Bonferroni correction across cluster pairs. The authors should either fix lags a priori, apply a correction over the lag grid, or use a data-driven lag-selection procedure with post-selection inference. This issue affects every entry in Table 2, including the headline C1,C4, C3,C4, and U1,U4 results.
  3. [§3.2, Eq. (3)] The aggregate toxicity Tc is defined as 1 - product over i in c of [1 - Fi·Ti], but Fi is never defined in the manuscript. The reader cannot tell whether Fi is a post frequency weight, a normalizer, or something else, and the reproducibility of the toxicity time series depends on this quantity. Please define Fi explicitly and state how posts with duplicate text, missing scores, or multiple authors are handled when forming the cluster-level series.
  4. [§2 and §4.4] The Covid sample is truncated to the first 500 observations after inspecting the series, described as a 'conservative approach' to avoid the low-activity final 231 observations. Because this clipping decision is made on the same data that are then tested, it is a post hoc selection that can alter the Granger results, and the paper reports no sensitivity analysis (for example, results on the full series, alternative cutoffs, or a pre-registered criterion). Please justify the cutoff on a priori grounds or show that the substantive findings do not depend on this choice.
  5. [§2, Table 1] The ideological engagement filter is based on a term list (nazism, holocaust, holodomor, etc.) and on per-sample thresholds (28 posts for Ukraine, 7 for Covid) taken from prior work, but the manuscript does not validate that users meeting this criterion are representative of the ideological camps in the two debates. If the term list selects a narrow or unusual subset of users, the clusters and the resulting Granger tests may reflect a specific subpopulation rather than the broader polarization dynamics. Please provide robustness evidence (for example, comparing cluster structure and key Granger results with and without the filter, or showing construct validity by contrasting included and excluded users' hashtags and sharing patterns).
minor comments (4)
  1. [Eq. (2)] Equation (2) defines structural dissimilarity as a negative cosine similarity; the en dash is presumably a minus sign, but as written the quantity can take negative values, which is an unusual dissimilarity measure. Please clarify the intended definition (e.g., 1 - cos) and its range.
  2. [§4.4] The sentence 'we test for temporal dependence between our variables in both directions, up to a lag of 155 observations in the Covid sample and 85 observations in the Ukraine sample' appears immediately after the clipping description; please state how the maximum lags were chosen and whether the lag grid is in days or in window steps.
  3. [Figures 2 and 3] The captions mention shaded areas and a gray dashed vertical line for the clipping point, but the figures themselves would benefit from a legend or explicit annotation identifying which curve is toxicity and which is structural dissimilarity, as well as the meaning of the shaded regions.
  4. [References and text] There are several typographical issues, including 'Unversity' in the OSoMe reference, 'inquerie' instead of 'inquiry' in §4.4, and the duplicated 'and and' in §4.3; these should be corrected in a final pass.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the Granger tests use measured time series; the only self-citation is a non-load-bearing ideological-term filter.

full rationale

The central derivation is not circular. Cluster centroids and structural dissimilarity D_{c1,c2} = -cos(mu_c1, mu_c2) are computed from the retweet-matrix PCA (Eqs. 1 and 2), and cluster toxicity T_c is independently computed via the Perspective API (Eqs. 3 and 4). The Granger tests in Section 4.4 compare these measured, time-indexed quantities; no parameter is fitted to the Granger outcome, and no significant result holds by construction. The ideological labels attached to clusters are post hoc interpretations of hashtags and example posts from the same data, so they do not force the statistical association between toxicity and structural dissimilarity. The only self-citation is the secondary-query term list drawn from Axelrod, Kim, and Paolillo (2024) to define 'ideologically engaged' users; this term list shapes the sample but is not a fitted parameter and is not equivalent to the target result. The within-camp C3,C4 result is therefore an empirical finding rather than a definitional artifact. A separate methodological concern is that the overlapping 7-day windows induce autocorrelation that may affect the Granger null distribution, but that is a statistical-validity issue, not circularity.

Assumptions & free parameters 7 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new theoretical entities. Its central claim rests on a set of hand-chosen parameters for PCA, clustering, and inclusion thresholds, plus domain assumptions about retweet and toxicity measures. The ideological term filter is the most consequential ad hoc element.

free parameters (7)
  • Number of PCs retained = 4
    Chosen by scree and PC plot inspection (Section 4.1).
  • Window SVD dimensions = 30
    Retained per window before sample PCA (Section 3.1).
  • Distance threshold for clustering = 10
    Euclidean distance from origin to initialize HDBSCAN (Section 3.1).
  • Toxicity retweet threshold = 10
    Only posts retweeted at least 10 times are scored for toxicity (Section 3.2).
  • Ideological engagement threshold = 28 Ukraine / 7 Covid
    Minimum posts with ideological terms to include a user (Section 2).
  • Granger lag search range = up to 155 (Covid) / 85 (Ukraine)
    Lags are searched and the minimum p-value is reported (Section 4.4).
  • De-trending window = preceding month
    Rolling OLS window for residualization (Section 3.3).
assumptions (3)
  • domain assumption Retweet behavior reflects political identity
    Authors cite Conover et al. 2011; this underlies the use of retweet structure as a measure of political polarization.
  • domain assumption Perspective API toxicity scores are a valid proxy for affective polarization
    Toxicity is used as the affect measure; the paper acknowledges language and dialect biases in limitations.
  • ad hoc to paper The ideological term list indexes relevant ideological frameworks
    The terms in Table 1 are drawn from the authors' prior work and are not validated against external benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of 'Congratulations, morons': Dynamics of Toxicity and Interaction Polarization in the Covid Vaccination and Ukraine War Twitter Debates." pith.science (2026). https://pith.science/paper/NNQZQVSF

@misc{pith2026250507646,
  author       = {Pith},
  title        = {Pith review of: 'Congratulations, morons': Dynamics of Toxicity and Interaction Polarization in the Covid Vaccination and Ukraine War Twitter Debates},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NNQZQVSF}},
  note         = {Machine review of arXiv:2505.07646}
}
read the original abstract

The existence of polarization and echo chambers has been noted in social media discussions of public concern such as the Covid-19 pandemic, foreign election interference, and regional conflicts. However, measuring polarization and assessing the manner in which polarization contributes to partisan behavior is not always possible to evaluate with static network or affect measurements. To address this, we conduct an analysis of two large Twitter datasets collected around Covid-19 vaccination and the Ukraine war to investigate polarization in terms of the evolution in influencer preferences and toxicity of post contents. By reducing retweet behavior in each sample to several key dimensions, we identify clusters that reflect ideological preferences, along with geographic or linguistic separation for some cases. By tracking the central retweet tendency of these clusters over time, we observe differences in the relative position of ideologically unaligned clusters compared to aligned ones, which we interpret as reflecting polarization dynamics in the information diffusion space. We then measure the toxicity of posts and test if toxicity in one cluster can be temporally dependent on its structural closeness to (or toxicity of) another. We find evidence of ideological opposition among clusters of users in both samples, and a temporal association between toxicity and structural divergence for at least two ideologically opposed clusters in our samples. These observations support the importance of analyzing polarization as a multifaceted dynamic phenomenon where polarization dynamics may also manifest in unexpected ways such as within a single ideological camp.

Figures

Figures reproduced from arXiv: 2505.07646 by the authors.

Figure 1
Figure 1. Time series for Covid (left) and Ukraine (right) samples. Retweet counts are depicted with solid lines while the number of unique users is shown with the dashed curves. To account for differences in scale, each plot has separate y-axes for each variable. Both samples began with queries to the Twitter API starting with limited set of topic-relevant [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Time series plots for each pair of clusters in the Covid sample. Blue curves show the structural dissimilarity for those clusters as defined in equation (2) while red curves show the combined toxicity for the two clusters as defined in equation (4). The shaded area hugging each curve corresponds to the difference between residuals from OLS fits and observations smoothed using a Gaussian kernel (σ = 3). Each curve is… view at source ↗
Figure 3
Figure 3. Time series plots for each pair of clusters in the Ukraine sample. Blue curves show the structural dissimilarity for those clusters as defined in equation (2) while red curves show the combined toxicity for the two clusters as defined in equation (4). The shaded area hugging each curve corresponds to the difference between residuals from OLS fits and observations smoothed using a Gaussian kernel (σ = 3). Each curve … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Left: Principal Component plots for the four retained components in the Covid sample. Color encodes users’ cluster membership. Right: Bar plot showing cluster sizes. Colors in bar plot map to clusters in the PC plots. Cluster C1 similarly stands out in [PITH_FULL_IMAG…
Figure 5
Figure 5. Figure 5: Hashtags with the top log-odds ratios per cluster (x-axis) for Covid (left) and Ukraine (right) samples. The log-odds ratio of a hashtag is computed in terms of the prevalence of that hashtag in a cluster compared to that hashtag’s prevalence among users in all other c…
Figure 6
Figure 6. Figure 6: Left: Principal Component plots for the four retained components in the Ukraine sample. Color encodes users’ cluster membership. Right: Bar plot showing cluster sizes. Colors in bar plot map to clusters in the PC plots. help Ukraine. The U5 cluster is rather more inwar…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

9 extracted references · 4 canonical work pages

  1. [6]

    Science Advances 9 (9)

    Quantifying ideological polarization on a network using generalized euclidean distance. Science Advances 9 (9). https://doi.org/10.1126/sciadv.abq2044. Iyengar, Shanto, Gaurav Sood, and Yphtach Lelkes

  2. [9]

    Nature Human Behaviour 5 (28-38)

    Affective polarization, local contexts and public opinion in america. Nature Human Behaviour 5 (28-38). Public Affairs, DOJ Office of. U.s. citizens convicted of conspiring to act as illegal agents of the russian government. https://www. justice.gov/archives/opa/pr/us-citizens-convicted-conspiring-act-illegal-agents-russian-government. Accessed: 2025-04-2...

  3. [1947]

    The Annals of Mathematical Statistics 18 (1): 50–60

    On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other. The Annals of Mathematical Statistics 18 (1): 50–60. https://doi.org/10.1214/aoms/1177730491. https://doi.org/10.1214/aoms/1177730491. McCoy, Jennifer, Tahmina Rahman, and Murat Somer

  4. [1969]

    Econometrica 37 (3)

    Investigating causal relations by econometric models and cross-spectral methods. Econometrica 37 (3). https://doi.org/10.2307/1912791. Hohmann, Marilena, Karel Devriendt, and Michele Coscia

  5. [2015]

    American Political Science Review 109 (1)

    Quantifying social media’s political space: estimating ideology from publicly revealed preferences on facebook. American Political Science Review 109 (1). ISSN: 0003-0554. https://doi.org/10.1017/ s0003055414000525. 18 D.S. Axelrod et al. Conover, Michael D, Bruno Gonçalves, Jacob Ratkiewicz, Alessandro Flammini, and Filippo Menczer. 2011a. Predicting the...

  6. [2018]

    Proceedings of the National Academy of Sciences (PNAS) 115 (37)

    Exposure to opposing views on social media can increase political polarization. Proceedings of the National Academy of Sciences (PNAS) 115 (37). https://doi.org/10.1073/pnas.1804840115. Axelrod, David, Sangyeon Kim, and John Paolillo

  7. [2020]

    In 2020 IEEE international conference on multisensor fusion and integration for intelligent systems (MFI)

    A hybrid approach to hierarchical density-based cluster selection. In 2020 IEEE international conference on multisensor fusion and integration for intelligent systems (MFI). IEEE, September. https: //doi.org/10.1109/mf i49285.2020.9235263. Mann, H. B., and D. R. Whitney

  8. [2023]

    Nature Human Behaviour 7:904–916

    Political polarization of news media and influencers on twitter in the 2016 and 2020 us presidential elections. Nature Human Behaviour 7:904–916. Google. Build with perspective api. https://developers.perspectiveapi.com. Accessed: 2025-02-10. Granger, C. W. J

Show all 9 references
  1. [2024]

    arXiv: 2406.16175 [cs.SI]

    The persistence of contrarianism on twitter: mapping users’ sharing habits for the ukraine war, covid-19 vaccination, and the 2022 midterm elections. arXiv: 2406.16175 [cs.SI]. Barberá, Pablo

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.