Pith. sign in

REVIEW 3 major objections 5 minor 76 references

Involvement drives complexity of language in online debates

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that in online debates, involvement drives linguistic complexity: users who post more negative or offensive content, and to a lesser degree those with partisan views or questionable reliability, use more complex language.

desk verdict A competent descriptive pipeline whose abstract outruns its own Table S1, with a tweet-count confound undermining the headline offensiveness-complexity claim. read the letter →

arxiv 2506.22098 v1 pith:ZVIYSTD2 submitted 2025-06-27 cs.CL cs.CYphysics.soc-ph

classification cs.CLcs.CYphysics.soc-ph
keywords languagecomplexityYule'sKonlinediscourseTwitter/Xpoliticalpolarizationoffensivenesssentimentanalysisinfluencernetworks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines roughly 1.66 million English tweets from more than 3,000 influential accounts debating COVID-19, COP26, and the Russia-Ukraine war, scoring each account with three metrics: Yule's K (lexical richness), the gzip compression ratio (repetitiveness), and the Flesch reading index (readability). Its central claim is that complexity is driven by involvement rather than expertise: individuals write more complex language than organizations, partisan accounts more than neutral ones, questionable-sourcing accounts more than reliable ones, and, most consistently across all three debates, users who post more negative or offensive content use more complex vocabulary. The authors read this as behavioral and ideological engagement producing richer lexical choices. A companion network analysis shows that influencers who share political stance and reliability ratings converge on shared vocabulary, forming distinct jargon communities.

What carries the argument

The central measure is Yule's K, defined for a text of $N$ tokens as $K = 10^4[-1/N + \sum_i V(i,N)(i/N)^2]$, where $V(i,N)$ counts how many distinct words occur exactly $i$ times; lower K indicates higher lexical richness, and the metric is treated by the authors as largely independent of text length. Two auxiliary measures, the gzip compression ratio of each user's concatenated tweets and the Flesch reading ease index, are used to corroborate the K-based findings. For the network analysis, the machinery is a weighted bipartite graph connecting influencers to the word types they use, binarized by the Revealed Comparative Advantage filter and projected onto the influencer layer with the Bipartite Configuration Model, a maximum-entropy null model that retains only statistically significant shared-vocabulary links; Louvain community detection on that projection yields clusters aligned with political stance and reliability.

What would settle it

Cut every user's tweets down to the same size — for example, 200 random tweets per account — recompute Yule's K, and check whether the drop in K across offensiveness and negative-sentiment quartiles survives; if the gradient vanishes or reverses, the 'involvement drives complexity' claim is an artifact of unequal text collections. Adding log tweet count as a covariate in a regression of K on offensiveness would settle the same question.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that engagement with an online debate predicts measurable linguistic complexity in the posts of influential users. The anchor result is a monotone decrease in Yule's K — lower values mean richer vocabulary — across rising offensiveness quartiles and rising negative-sentiment quartiles, statistically significant in all three datasets and corroborated by the gzip and Flesch metrics. The same direction appears for account type, political leaning, and reliability under at least two of the three measures, with the strongest and most consistent signal attached to offensiveness and negativity. The authors generalize this into the paper's title statement: users with higher involvement, whether political, ideological, or behavioral, generally show a more complex vocabulary.

Load-bearing premise

Complexity is measured on each user's full set of tweets pooled into one text, and groups of users post very different numbers of tweets; since the paper itself shows that vocabulary size grows as a power law with tweet count, the complexity differences between groups could reflect posting volume rather than a real difference in linguistic style.

Editorial extensions

If this is right

  • If the claim holds, lexical richness in at least three major contested debates is predictable from behavioral signals such as offensiveness and negative sentiment, with an account's complexity rising in step with its emotional engagement.
  • The validated word-sharing networks imply that political stance and reliability are visible in vocabulary alone, so like-minded influencer groups can be recovered from word co-occurrence without reading any content.
  • Because the pattern repeats across the three debates but is strongest in the science-adjacent ones (COVID-19 and COP26) and weakest in the geopolitical one, topic type appears to modulate how sharply linguistic camps form.
  • Since partisan and questionable-reliability accounts also score as more complex on standard readability metrics, the paper implies that complexity measures do not track credibility: more complex language is not better-sourced language.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because vocabulary size scales as a power law with tweet count (slope about 0.55) and complexity is computed on pooled per-user corpora, the group differences may be partly driven by posting volume; the paper does not run a size-matched or per-tweet control, so this remains an open test.
  • A directly testable extension the authors do not attempt: track individual influencers over time and check whether their lexical complexity rises during periods of intense engagement, which would support a causal reading of the title claim.
  • If the pattern generalizes, moderation and credibility pipelines that treat sophisticated phrasing as a quality signal could systematically favor hostile or low-reliability accounts — a consequence the paper leaves implicit.
  • The jargon-convergence result suggests lexical similarity networks could act as a stance-detection tool with no semantic annotation, a practical application the paper does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper analyzes Twitter/X data from three contested topics (COVID-19, COP26, and the Russia–Ukraine war) to examine whether linguistic complexity is associated with account type, political leaning, content reliability, sentiment, and offensiveness. Complexity is measured with Yule's K, gzip compression ratio, and Flesch Reading Ease, each computed on concatenated per-user text, and an influencer network is constructed from shared word types using an entropy-based bipartite null model. The main reported findings are that individuals use more complex language than organizations, partisan and questionable accounts show greater complexity than moderate and reliable accounts, and users who post more negative or offensive content use more complex vocabulary.

Significance. If the central association were established, the paper would be a useful contribution to the sociolinguistic analysis of digital discourse, connecting behavioral engagement to lexical complexity rather than formal expertise. The manuscript has clear strengths: it combines multiple complementary complexity metrics; it validates LLM-generated political and reliability labels against an external benchmark (MBFC, Cohen's kappa = 0.58 and 0.75); and the network analysis is based on an explicit maximum-entropy null model rather than ad hoc thresholds. However, the main empirical claim is currently not supported because the complexity metrics are computed on per-user concatenated corpora without controlling for corpus size, which the paper's own Figure 1 shows is strongly correlated with vocabulary size. With appropriate controls, the dataset and analyses could support a defensible association claim, but the present form overstates both the evidence and the causal interpretation.

major comments (3)
  1. [§2.5, §3.2, Figure 3, Table S1] The central claim that users in higher offensiveness and negativity quartiles use more complex language rests on Kruskal-Wallis comparisons of Yule's K computed on the concatenation of all of a user's tweets after stopword removal and stemming. Figure 1 shows that vocabulary size scales steeply with tweet count (log-log slope 0.55–0.59), and Yule's K is only asymptotically length-independent for fixed text populations; with finite samples more tweets sample more of the vocabulary tail and push K downward. Neither Table S1 nor the main text includes tweet count, total tokens, or total characters as a covariate, so the monotone decrease in K across offensiveness/negativity quartiles in Figure 3 could arise even if intrinsic stylistic complexity were identical across groups. The gzip and Flesch analyses in Figures S3 and S4 are subject to the same length sensitivity and do not remove the confound. Please re-run the group comparisons with activity-level controls, such as regression with corpus size, fixed-size random subsamples, or per-tweet metrics averaged after aggregation, before drawing conclusions about offensiveness, negativity, or account-type differences.
  2. [§3.2, Table S1, Abstract] The abstract's claim of significant differences across all four axes is not supported by the Yule's K results in Table S1: political leaning is non-significant for COP26 (p = 0.4948) and Ukraine (p = 0.5725), and reliability is non-significant for the same two datasets (p = 0.743 and p = 0.7568). The text in Section 3.2 acknowledges these null results for K, but the subsequent statement that Figures S3 and S4 show the patterns are persistent relies on gzip and Flesch, which, per Comment 1, are also length-sensitive and therefore do not independently confirm the pattern. The claims should be restricted to the specific comparisons that remain significant after controlling for corpus size.
  3. [Title, §4 Discussion] The title and the Discussion state that 'involvement drives complexity of language,' but the study is cross-sectional and the reported analyses are univariate group comparisons; no causal identification strategy or temporal ordering is presented. Even after the corpus-size confound is fixed, the appropriate conclusion would be an association between involvement-related attributes and measured complexity. Please replace the causal framing with association language or provide an explicit argument for why the direction of causality can be inferred from these data.
minor comments (5)
  1. [§2.5] The term 'readibility' should be 'readability,' and the metric name is written both as 'G-Zip' and 'gzip'; please standardize the terminology.
  2. [§2.4, §4] There are minor language errors, including 'excellence performance' (should be 'excellent performance'), 'deep more into' (should be 'delve deeper into'), and 'Ukranian' (should be 'Ukrainian').
  3. [Table S1] Reporting p-values only makes it difficult to assess the magnitude of group differences; please add group sizes and a standardized effect size for each Kruskal-Wallis comparison, such as epsilon-squared.
  4. [§2.6, §3.3] The network analysis measures lexical overlap rather than complexity; the text should make this explicit so that the network results are not read as direct evidence for the complexity conclusions.
  5. [§4] The limitations paragraph mentions platform and language scope but does not acknowledge the corpus-size confound identified above; this should be stated as a limitation and, ideally, addressed through supplementary controls.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: complexity metrics and LLM labels are independently computed, and no fitted parameter is recycled as a prediction.

full rationale

This paper does not exhibit circular reasoning in the specific sense of a claimed derivation reducing to its inputs. The complexity measures (Yule's K, gzip, Flesch) are computed from tweet text independently of the outcome groups; the offensiveness and negativity scores are obtained from pre-trained RoBERTa models, aggregated per user, and then compared with complexity via Kruskal-Wallis tests. The LLM political and reliability labels were validated against an external MBFC benchmark (Cohen's kappa = 0.58 and 0.75), so those labels are not self-referential. The dataset is inherited from Loru et al. (2024) with some author overlap, but that is a data-collection step rather than a proof step, and the complexity analyses are new computations on that data. No parameter is fitted and then renamed a prediction; no self-citation is used as a uniqueness theorem or to forbid alternatives. The clearest weakness—per-user corpora of very different sizes and no tweet-count covariate in the group comparisons—is a confounding and validity concern, not a circularity in the derivation chain. Under the stated criteria, no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fit to data: the complexity metrics, RCA threshold (1), and FDR level (0.05) are standard methodological constants. The main implicit assumptions are the validity of concatenated per-user corpora, the interchangeability of the three complexity metrics, the generalization of LLM labels from 56 matched users, and the latent involvement construct. No new entities are introduced.

assumptions (5)
  • domain assumption The concatenation of all of a user's tweets after stemming and stopword removal is a valid corpus for measuring that user's linguistic complexity.
    Invoked in Sections 2.2 and 2.5 before computing Yule's K and gzip ratio; this ignores tweet count as a confounder.
  • domain assumption Yule's K, gzip compression ratio, and Flesch index are complementary and comparable measures of a single latent construct 'textual complexity'.
    Section 2.5 states they 'provide a relatively orthogonal view' and treats them as interchangeable in the Discussion.
  • domain assumption Gemini-assigned political stance and reliability labels, validated on 56 MBFC-matched users, generalize to the full set of 3,016 influencers.
    Section 2.3 reports kappa of 0.58 and 0.75 on a small overlap and then uses the labels as ground truth for all users.
  • standard math The BiCM null model and FDR correction produce a meaningful influencer projection.
    Section 2.6 imports the maximum-entropy model from Saracco et al. and Vallarano et al.; this is a standard toolbox assumption.
  • ad hoc to paper Negative sentiment, offensiveness, and political partisanship are treated as proxies for a single latent 'involvement' that drives language complexity.
    Title and Discussion use 'involvement' as the cause, but no direct measure of involvement is provided; this is an interpretive assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Involvement drives complexity of language in online debates." pith.science (2026). https://pith.science/paper/ZVIYSTD2

@misc{pith2026250622098,
  author       = {Pith},
  title        = {Pith review of: Involvement drives complexity of language in online debates},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZVIYSTD2}},
  note         = {Machine review of arXiv:2506.22098}
}
read the original abstract

Language is a fundamental aspect of human societies, continuously evolving in response to various stimuli, including societal changes and intercultural interactions. Technological advancements have profoundly transformed communication, with social media emerging as a pivotal force that merges entertainment-driven content with complex social dynamics. As these platforms reshape public discourse, analyzing the linguistic features of user-generated content is essential to understanding their broader societal impact. In this paper, we examine the linguistic complexity of content produced by influential users on Twitter across three globally significant and contested topics: COVID-19, COP26, and the Russia-Ukraine war. By combining multiple measures of textual complexity, we assess how language use varies along four key dimensions: account type, political leaning, content reliability, and sentiment. Our analysis reveals significant differences across all four axes, including variations in language complexity between individuals and organizations, between profiles with sided versus moderate political views, and between those associated with higher versus lower reliability scores. Additionally, profiles producing more negative and offensive content tend to use more complex language, with users sharing similar political stances and reliability levels converging toward a common jargon. Our findings offer new insights into the sociolinguistic dynamics of digital platforms and contribute to a deeper understanding of how language reflects ideological and social structures in online spaces.

Figures

Figures reproduced from arXiv: 2506.22098 by the authors.

Figure 1
Figure 1. Bivariate distribution of the number of tweets and the vocabulary size for each influencer considered in the [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. K-complexity distributions according to a) the account type of the influencers (Individual or Organization), b) the political leaning of the influencer (Left, Center, Right), c) the reliability of the influencer (Questionable or Reliable). S2 and S3 summarize the original bipartite networks and their statistically validated projections (see Supplementary Information). Although the bipartite graphs are large, the BiC… view at source ↗
Figure 3
Figure 3. K-complexity distributions by (a) user offensiveness class and (b) negative sentiment class. The offensiveness class is defined based on how frequently a user posts offensive tweets, while the negative sentiment class is determined by the average negativity score of the user’s tweets. The results, shown in Panel b) and c) of [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: (a) Visualization of the influencers’ networks for the three datasets. In panels (b) and (c), Size represent the number of influencers in each community, colors represent the political leaning (b) or the reliability label (c) of the communities detected using the Louva…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

76 extracted references · 68 canonical work pages

  1. [1]

    The Cambridge encyclopedia of the English language

    David Crystal. The Cambridge encyclopedia of the English language. Cambridge university press, 2018

  2. [2]

    Language contact, creolization, and genetic linguistics

    Sarah Grey Thomason and Terrence Kaufman. Language contact, creolization, and genetic linguistics. Univ of California Press, 2023

  3. [3]

    Mediatization and sociolinguistic change, volume 36

    Jannis Androutsopoulos. Mediatization and sociolinguistic change, volume 36. Walter de Gruyter GmbH & Co KG, 2014

  4. [4]

    Evaluating the effect of viral posts on social media engagement

    Emanuele Sangiorgio, Niccolò Di Marco, Gabriele Etta, Matteo Cinelli, Roy Cerqueti, and Walter Quattrociocchi. Evaluating the effect of viral posts on social media engagement. Scientific Reports, 15(1):639, 2025

  5. [5]

    Filter bubbles, echo chambers, and online news consumption

    Seth Flaxman, Sharad Goel, and Justin M Rao. Filter bubbles, echo chambers, and online news consumption. Public opinion quarterly, 80(S1):298–320, 2016

  6. [6]

    The echo chamber effect on social media

    Matteo Cinelli, Gianmarco De Francisci Morales, Alessandro Galeazzi, Walter Quattrociocchi, and Michele Starnini. The echo chamber effect on social media. Proceedings of the National Academy of Sciences , 118(9):e2023301118, 2021

  7. [7]

    Like-minded sources on facebook are prevalent but not polarizing

    Brendan Nyhan, Jaime Settle, Emily Thorson, Magdalena Wojcieszak, Pablo Barberá, Annie Y Chen, Hunt Allcott, Taylor Brown, Adriana Crespo-Tenorio, Drew Dimmery, et al. Like-minded sources on facebook are prevalent but not polarizing. Nature, 620(7972):137–144, 2023

  8. [8]

    Social media, political polarization, and political disinformation: A review of the scientific literature

    Joshua A Tucker, Andrew Guess, Pablo Barberá, Cristian Vaccari, Alexandra Siegel, Sergey Sanovich, Denis Stukal, and Brendan Nyhan. Social media, political polarization, and political disinformation: A review of the scientific literature. Political polarization, and political disinformation: a review of the scientific literature (March 19, 2018), 2018

Show all 76 references
  1. [9]

    Growing polarization around climate change on social media

    Max Falkenberg, Alessandro Galeazzi, Maddalena Torricelli, Niccolò Di Marco, Francesca Larosa, Madalina Sas, Amin Mekacher, Warren Pearce, Fabiana Zollo, Walter Quattrociocchi, et al. Growing polarization around climate change on social media. Nature Climate Change, 12(12):111...

  2. [10]

    Se l’uomo non basta: Speranze e timori nell’uso della tecnologia contro Covid-19

    Paolo Benanti. Se l’uomo non basta: Speranze e timori nell’uso della tecnologia contro Covid-19. LIT EDIZIONI, 2020

  3. [11]

    Eugene Stanley, and Walter Quattrociocchi

    Michela Del Vicario, Alessandro Bessi, Fabiana Zollo, Fabio Petroni, Antonio Scala, Guido Caldarelli, H. Eugene Stanley, and Walter Quattrociocchi. The spreading of misinformation online.Proceedings of the National Academy of Sciences, 113(3):554–559, 2016

  4. [12]

    Sandra González-Bailón, David Lazer, Pablo Barberá, Meiqing Zhang, Hunt Allcott, Taylor Brown, Adriana Crespo-Tenorio, Deen Freelon, Matthew Gentzkow, Andrew M. Guess, Shanto Iyengar, Young Mie Kim, Neil Malhotra, Devra Moehler, Brendan Nyhan, Jennifer Pan, Carlos Velasco Rive...

  5. [13]

    Internet, social media and online hate speech

    Sergio Andrés Castaño-Pulgarín, Natalia Suárez-Betancur, Luz Magnolia Tilano Vega, and Harvey Mauricio Her- rera López. Internet, social media and online hate speech. systematic review. Aggression and Violent Behavior, 58:101608, 2021

  6. [14]

    Online hate speech

    Alexandra A Siegel. Online hate speech. Social media and democracy: The state of the field, prospects for reform, pages 56–88, 2020

  7. [15]

    A large-scale behavioural analysis of bots and humans on twitter

    Zafar Gilani, Reza Farahbakhsh, Gareth Tyson, and Jon Crowcroft. A large-scale behavioural analysis of bots and humans on twitter. ACM Trans. Web, 13(1), February 2019. 11 Complexity72h 23-27 J UNE 2025 - M ADRID

  8. [16]

    Persistent interaction patterns across social media platforms and over time

    Michele Avalle, Niccolò Di Marco, Gabriele Etta, Emanuele Sangiorgio, Shayan Alipour, Anita Bonetti, Lorenzo Alvisi, Antonio Scala, Andrea Baronchelli, Matteo Cinelli, et al. Persistent interaction patterns across social media platforms and over time. Nature, 628(8008):582–589, 2024

  9. [17]

    Ideo- logical fragmentation of the social media ecosystem: From echo chambers to echo platforms

    Edoardo Di Martino, Alessandro Galeazzi, Michele Starnini, Walter Quattrociocchi, and Matteo Cinelli. Ideo- logical fragmentation of the social media ecosystem: From echo chambers to echo platforms. arXiv preprint arXiv:2411.16826, 2024

  10. [18]

    Sentiment analysis with nlp on twitter data

    Md Rakibul Hasan, Maisha Maliha, and M Arifuzzaman. Sentiment analysis with nlp on twitter data. In 2019 international conference on computer, communication, chemical, materials and electronic engineering (IC4ME2), pages 1–4. IEEE, 2019

  11. [19]

    Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter

    Martin Müller, Marcel Salathé, and Per E Kummervold. Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter. Frontiers in artificial intelligence, 6:1023281, 2023

  12. [20]

    Inference of social media opinion trends in 2022 italian elections

    Simon Zollo, Matteo Cinelli, Gabriele Etta, Roy Cerqueti, and Walter Quattrociocchi. Inference of social media opinion trends in 2022 italian elections. Expert Systems with Applications, 269:126377, 2025

  13. [21]

    Using the president’s tweets to understand political diversion in the age of social media

    Stephan Lewandowsky, Michael Jetter, and Ullrich KH Ecker. Using the president’s tweets to understand political diversion in the age of social media. Nature communications, 11(1):5764, 2020

  14. [22]

    Twitter as arena for the authentic outsider: exploring the social media campaigns of trump and clinton in the 2016 us presidential election

    Gunn Enli. Twitter as arena for the authentic outsider: exploring the social media campaigns of trump and clinton in the 2016 us presidential election. European journal of communication, 32(1):50–61, 2017

  15. [23]

    Use of sentiment analysis for capturing patient experience from free-text comments posted online

    Felix Greaves, Daniel Ramirez-Cano, Christopher Millett, Ara Darzi, and Liam Donaldson. Use of sentiment analysis for capturing patient experience from free-text comments posted online. J Med Internet Res, 15(11):e239, Nov 2013

  16. [24]

    Sentiment analysis of comment texts based on bilstm

    Guixian Xu, Yueting Meng, Xiaoyu Qiu, Ziheng Yu, and Xu Wu. Sentiment analysis of comment texts based on bilstm. IEEE Access, 7:51522–51532, 2019

  17. [25]

    Sentiment analysis of comments in social media

    Abdulrahman Alrumaih, Ali Al-Sabbagh, Ruaa Alsabah, Harith Kharrufa, and James Baldwin. Sentiment analysis of comments in social media. International Journal of Electrical & Computer Engineering (2088-8708), 10(6), 2020

  18. [26]

    Inferring latent user properties from texts published in social media

    Svitlana V olkova, Yoram Bachrach, Michael Armstrong, and Vijay Sharma. Inferring latent user properties from texts published in social media. Proceedings of the AAAI Conference on Artificial Intelligence, 29(1), Mar. 2015

  19. [27]

    Evolution of the most common english words and phrases over the centuries

    Matjaž Perc. Evolution of the most common english words and phrases over the centuries. Journal of The Royal Society Interface, 9(77):3323–3328, 2012

  20. [28]

    Measurement of the size of general english vocabulary through the elementary grades and high school

    Mary K Smith. Measurement of the size of general english vocabulary through the elementary grades and high school. Genetic Psychology Monographs, 24:311–345, 1941

  21. [29]

    Breadth of vocabulary and advanced english study: An empirical investigation

    Erwin Tschirner. Breadth of vocabulary and advanced english study: An empirical investigation. Electronic Journal of Foreign Language Teaching, 1(1):27–39, 2004

  22. [30]

    V ocabulary size revisited: the link between vocabulary size and academic achievement

    James Milton and Jeanine Treffers-Daller. V ocabulary size revisited: the link between vocabulary size and academic achievement. Applied Linguistics Review, 4(1):151–172, 2013

  23. [31]

    Always on: Language in an online and mobile world

    Naomi S Baron. Always on: Language in an online and mobile world. Oxford University Press, 2010

  24. [32]

    Because internet: Understanding the new rules of language

    Gretchen McCulloch. Because internet: Understanding the new rules of language. Penguin, 2020

  25. [33]

    Discourse of Twitter and social media

    Michele Zappavigna. Discourse of Twitter and social media. Bloomsbury Publishing, 2012

  26. [34]

    Words onscreen: The fate of reading in a digital world

    Naomi S Baron. Words onscreen: The fate of reading in a digital world. Oxford University Press, 2015

  27. [35]

    Di Marco, Edoardo Loru, Anita Bonetti, Alessandra Olga Grazia Serra, Matteo Cinelli, and Walter Quattro- ciocchi

    N. Di Marco, Edoardo Loru, Anita Bonetti, Alessandra Olga Grazia Serra, Matteo Cinelli, and Walter Quattro- ciocchi. Patterns of linguistic simplification on social media platforms over time. Proceedings of the National Academy of Sciences, 121(50):e2412105121, 2024

  28. [36]

    Highly engaging events reveal semantic and temporal compression in online community discourse

    Antonio Desiderio, Anna Mancini, Giulio Cimini, and Riccardo Di Clemente. Highly engaging events reveal semantic and temporal compression in online community discourse. PNAS Nexus, 4(3):pgaf056, 02 2025

  29. [37]

    Song lyrics have become simpler and more repetitive over the last five decades

    Emilia Parada-Cabaleiro, Maximilian Mayerl, Stefan Brandl, Marcin Skowron, Markus Schedl, Elisabeth Lex, and Eva Zangerle. Song lyrics have become simpler and more repetitive over the last five decades. Scientific Reports, 14(1):5531, 2024

  30. [38]

    Decoding musical evolution through network science

    Niccolo’ Di Marco, Edoardo Loru, Alessandro Galeazzi, Matteo Cinelli, and Walter Quattrociocchi. Decoding musical evolution through network science. arXiv preprint arXiv:2501.07557, 2025

  31. [39]

    Who sets the agenda on social media? ideology and polarization in online debates

    Edoardo Loru, Alessandro Galeazzi, Anita Bonetti, Emanuele Sangiorgio, Niccolò Di Marco, Matteo Cinelli, Andrea Baronchelli, and Walter Quattrociocchi. Who sets the agenda on social media? ideology and polarization in online debates. arXiv preprint arXiv:2412.05176, 2024. 12 C...

  32. [40]

    Tweets in time of conflict: A public dataset tracking the twitter discourse on the war between Ukraine and Russia

    Emily Chen and Emilio Ferrara. Tweets in time of conflict: A public dataset tracking the twitter discourse on the war between Ukraine and Russia. In Proceedings of the International AAAI Conference on Web and Social Media, volume 17, pages 1006–1013, 2023

  33. [41]

    quanteda: An r package for the quantitative analysis of textual data

    Kenneth Benoit, Kohei Watanabe, Haiyan Wang, Paul Nulty, Adam Obeng, Stefan Müller, and Akitaka Matsuo. quanteda: An r package for the quantitative analysis of textual data. Journal of Open Source Software, 3(30):774, 2018

  34. [42]

    Republicans are flagged more often than democrats for sharing misinformation on x’s community notes

    Thomas Renault, Mohsen Mosleh, and David G Rand. Republicans are flagged more often than democrats for sharing misinformation on x’s community notes. Proceedings of the National Academy of Sciences , 122(25):e2502053122, 2025

  35. [43]

    Chatgpt outperforms crowd workers for text-annotation tasks

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

  36. [44]

    Predicting the topical stance and political leaning of media using tweets

    Peter Stefanov, Kareem Darwish, Atanas Atanasov, and Preslav Nakov. Predicting the topical stance and political leaning of media using tweets. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 527–537, 2020

  37. [45]

    Political polarization of news media and influencers on twitter in the 2016 and 2020 us presidential elections

    James Flamino, Alessandro Galeazzi, Stuart Feldman, Michael W Macy, Brendan Cross, Zhenkun Zhou, Matteo Serafino, Alexandre Bovet, Hernán A Makse, and Boleslaw K Szymanski. Political polarization of news media and influencers on twitter in the 2016 and 2020 us presidential ele...

  38. [46]

    Tweeteval: Unified benchmark and comparative evaluation for tweet classification

    Francesco Barbieri, Jose Camacho-Collados, Leonardo Neves, and Luis Espinosa-Anke. Tweeteval: Unified benchmark and comparative evaluation for tweet classification. arXiv preprint arXiv:2010.12421, 2020

  39. [47]

    Roberta: A robustly optimized bert pretraining approach

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019

  40. [48]

    Xlm-roberta based sentiment analysis of tweets on metaverse and 6g

    Akshat Gaurav, Brij B Gupta, Sachin Sharma, Ritika Bansal, and Kwok Tai Chui. Xlm-roberta based sentiment analysis of tweets on metaverse and 6g. Procedia Computer Science, 238:902–907, 2024

  41. [49]

    Sentiment analysis in tweets: an assessment study from classical to modern word representation models

    Sérgio Barreto, Ricardo Moura, Jonnathan Carvalho, Aline Paes, and Alexandre Plastino. Sentiment analysis in tweets: an assessment study from classical to modern word representation models. Data Mining and Knowledge Discovery, 37(1):318–380, 2023

  42. [50]

    Analyzing sentiments: A comprehensive study of roberta-based sentiment analysis on twitters

    A Krishnamoorthy, KA Sundhar, V Naveen Kumar, and V Karthik. Analyzing sentiments: A comprehensive study of roberta-based sentiment analysis on twitters. In 2024 4th International Conference on Advancement in Electronics & Communication Engineering (AECE), pages 626–630. IEEE, 2024

  43. [51]

    Transformer models for recognizing abusive language an investigation and review on tweeteval and solid dataset

    Fabeela Ali Rawther and Geevarghese Titus. Transformer models for recognizing abusive language an investigation and review on tweeteval and solid dataset. In 2023 Second International Conference on Electrical, Electronics, Information and Communication Technologies (ICEEICT), ...

  44. [52]

    Roberta-lstm: a hybrid model for sentiment analysis with transformer and recurrent neural network

    Kian Long Tan, Chin Poo Lee, Kalaiarasi Sonai Muthu Anbananthen, and Kian Ming Lim. Roberta-lstm: a hybrid model for sentiment analysis with transformer and recurrent neural network. IEEE Access, 10:21517–21525, 2022

  45. [53]

    Bertweet: A pre-trained language model for english tweets

    Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. Bertweet: A pre-trained language model for english tweets. arXiv preprint arXiv:2005.10200, 2020

  46. [54]

    TweetEval: Unified benchmark and comparative evaluation for tweet classification

    Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1644–1650, Online, November 2020...

  47. [55]

    A machine learning approach to identify toxic language in the online space

    Lisa Kaati, Amendra Shrestha, and Nazar Akrami. A machine learning approach to identify toxic language in the online space. In 2022 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), pages 396–402. IEEE, 2022

  48. [56]

    Uhh-lt at semeval-2020 task 12: Fine-tuning of pre-trained transformer networks for offensive language detection

    Gregor Wiedemann, Seid Muhie Yimam, and Chris Biemann. Uhh-lt at semeval-2020 task 12: Fine-tuning of pre-trained transformer networks for offensive language detection. arXiv preprint arXiv:2004.11493, 2020

  49. [57]

    TimeLMs: Diachronic language models from Twitter

    Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, and Jose Camacho-collados. TimeLMs: Diachronic language models from Twitter. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 251–26...

  50. [58]

    Bertopic: Neural topic modeling with a class-based tf-idf procedure

    Maarten Grootendorst. Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794, 2022. 13 Complexity72h 23-27 J UNE 2025 - M ADRID

  51. [59]

    Density-based clustering validation

    Davoud Moulavi, Pablo A Jaskowiak, Ricardo JGB Campello, Arthur Zimek, and Jörg Sander. Density-based clustering validation. In Proceedings of the 2014 SIAM international conference on data mining, pages 839–847. SIAM, 2014

  52. [60]

    Indicators of text complexity

    Kristian TH Jensen. Indicators of text complexity. Mees, IM; F . Alves & S. Göpferich (eds.), pages 61–80, 2009

  53. [61]

    The psycho-biology of language: An introduction to dynamic philology

    George Kingsley Zipf. The psycho-biology of language: An introduction to dynamic philology. Routledge, 2013

  54. [62]

    Computational constancy measures of texts—yule’s k and rényi’s entropy

    Kumiko Tanaka-Ishii and Shunsuke Aihara. Computational constancy measures of texts—yule’s k and rényi’s entropy. Computational Linguistics, 41(3):481–502, 2015

  55. [63]

    On measures of entropy and information

    Alfréd Rényi. On measures of entropy and information. In Proceedings of the fourth Berkeley symposium on mathematical statistics and probability, volume 1: contributions to the theory of statistics , volume 4, pages 547–562. University of California Press, 1961

  56. [64]

    Vocabulaire et stylistique, volume 8

    Daniel Dugast. Vocabulaire et stylistique, volume 8. Slatkine, 1979

  57. [65]

    How variable may a constant be? measures of lexical richness in perspective

    Fiona J Tweedie and R Harald Baayen. How variable may a constant be? measures of lexical richness in perspective. Computers and the Humanities, 32:323–352, 1998

  58. [66]

    The statistical study of literary vocabulary

    C Udny Yule. The statistical study of literary vocabulary. Cambridge University Press, 2014

  59. [67]

    Recurring patterns in online social media interactions during highly engaging events

    Antonio Desiderio, Anna Mancini, Giulio Cimini, and Riccardo Di Clemente. Recurring patterns in online social media interactions during highly engaging events. arXiv preprint arXiv:2306.14735, 2023

  60. [68]

    Trade liberalization and revealed comparative advantage

    Bela Balassa. Trade liberalization and revealed comparative advantage. The Manchester School of Economic and Social Studies, 33:99–123, 1965

  61. [69]

    Inferring monopartite projections of bipartite networks: an entropy-based approach

    Fabio Saracco, Mika J Straka, Riccardo Di Clemente, Andrea Gabrielli, Guido Caldarelli, and Tiziano Squartini. Inferring monopartite projections of bipartite networks: an entropy-based approach. New Journal of Physics, 19(5):053022, may 2017

  62. [70]

    The statistical physics of real-world networks

    Giulio Cimini, Tiziano Squartini, Fabio Saracco, Diego Garlaschelli, Andrea Gabrielli, and Guido Caldarelli. The statistical physics of real-world networks. Nature Reviews Physics, 1(1):58–71, Jan 2019

  63. [71]

    Controlling the false discovery rate: A practical and powerful approach to multiple testing

    Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society. Series B (Methodological), 57(1):289–300, 2023/06/13/

  64. [72]

    Fast and scalable likelihood maximization for exponential random graph models with local constraints

    Nicolò Vallarano, Matteo Bruno, Emiliano Marchese, Giuseppe Trapani, Fabio Saracco, Giulio Cimini, Mario Zanon, and Tiziano Squartini. Fast and scalable likelihood maximization for exponential random graph models with local constraints. Scientific Reports, 11(1):15227, 2021

  65. [73]

    Hagberg, Daniel A

    Aric A. Hagberg, Daniel A. Schult, and Pieter J. Swart. Exploring network structure, dynamics, and function using networkx. In Gaël Varoquaux, Travis Vaught, and Jarrod Millman, editors, Proceedings of the 7th Python in Science Conference, pages 11 – 15, Pasadena, CA USA, 2008

  66. [74]

    Gephi: An open source software for exploring and manipulating networks

    Mathieu Bastian, Sébastien Heymann, and Mathieu Jacomy. Gephi: An open source software for exploring and manipulating networks. In Proceedings of the International AAAI Conference on Weblogs and Social Media (ICWSM), 2009

  67. [75]

    L” (Left), combining “left

    Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, E...

  68. [1995]

    Full publication date: 1995

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.