{"id":"9fe816c4-f43c-4cc0-8844-e43db47f9e95","arxiv_id":"2504.19594","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A large-scale map of the Italian Telegram sphere shows ideological homophily, normalized toxicity across highly toxic communities, and consistent hate targets including Italians, Black people, Jews, and gay people.","lead":"This paper maps the Italian Telegram ecosystem by analyzing 186 million messages from over 13,000 public chats collected in 2023. It shows that toxic speech is spread broadly rather than confined to fringe groups, and that Italians, Black people, Jews, and gay people are the most common targets of hate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All reported Gini indices lie between 0.6 and 0.9, so Section 5.3's claim that toxicity is normalized across most chats in high-toxicity communities is not supported by the reported data.","rationale":"The paper's central contribution has three advertised findings: ideological homophily with a mixed far-left/far-right Warfare community, toxicity normalized in highly toxic communities, and stable hate targets (Black, Jewish, gay people). The reader's weakest assumption focused on the unvalidated ChatGPT-4o political labels. That is a legitimate measurement-validity concern, but the toxicity-normalization claim is more load-bearing because it is an inferential statement that can be checked against the paper's own reported numbers, and those numbers appear to contradict it. Figure 6 shows Gini indices between 0.6 and 0.9 for every community; even 0.6 denotes substantial inequality, and 0.9 denotes extreme concentration. Therefore 'low Gini' in this dataset is low only relative to other communities, not low in an absolute sense. The Pearson correlation (r = -0.87) shows that more toxic communities are less unequal, but it does not show that toxic messages are spread across most chats or that toxic behavior is normalized. The text in Section 5.3 explicitly says toxicity is 'normalized across most of their chats,' which goes beyond the evidence. Additionally, the unweighted chat-level aggregation gives equal weight to a 10-message chat and a 1-million-message chat, so the Gini cannot establish message-level or user-level normalization. A concrete re-analysis reporting zero-toxicity chat fractions and top-decile concentration would settle the point. Because this is a central claim repeated in the abstract and conclusions, the paper needs revision. The reader's CONDITIONAL verdict already requires revisions; this concern adds a specific, internally checkable reason for that condition rather than changing the verdict, so I leave the verdict as UNCHANGED. My agreement is partial because I agree with the reader's overall assessment of the paper but identify a different, and arguably more decisive, weak point than the political-label validation.","tokens_in":17701,"tokens_out":10425,"duration_ms":111203,"concrete_test":"Recompute the analysis behind Figure 6 and report, for each of the 15 communities: (i) the percentage of chats with zero toxic messages, and (ii) the share of the community's total toxic messages accounted for by its top 10% most toxic chats, weighting each chat by its message count. If, in the high-mean communities (e.g., General, Gaming, SportInfo, Adult), more than 50% of chats have zero toxic messages, or the top decile accounts for more than half of all toxic messages, then Section 5.3's claim that toxicity is 'evenly spread across many chats' and 'normalized across most of their chats' is not supported. As a robustness check, recompute the Gini-mean correlation after weighting each chat's toxicity by its message count; if the strong negative correlation disappears, the result is an artifact of unweighted chat-level aggregation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3 interprets the negative correlation between community mean chat toxicity and the Gini index (Pearson r = -0.872, R2 = 0.76, Fig. 6) as evidence that toxicity is 'widely normalized within highly toxic communities' and 'evenly spread across many chats.' This inference is not supported by the reported Gini values: all 15 communities fall in the 0.6-0.9 range, meaning every community has a highly unequal distribution of toxicity across chats. A Gini of about 0.7 is compatible with roughly 70% of chats having zero toxic messages and the remaining 30% carrying all toxicity; a Gini of 0.9 is compatible with about 90% zero-toxicity chats. Thus the negative correlation is a relative statement that high-mean communities are less concentrated than low-mean communities; it does not establish that toxicity is shared across most chats in any absolute sense. The text goes further, claiming toxicity is 'normalized across most of their chats,' which is an overclaim. In addition, the analysis uses unweighted chat-level toxicity percentages, so a 10-message chat and a 1-million-message chat count equally. Consequently, even if many small chats have nonzero toxicity, a few large chats could account for most toxic messages; the Gini cannot rule this out. The reader's concern about unvalidated ChatGPT-4o political labels remains valid, but the toxicity-normalization claim is internally checkable from the paper's own reported quantities and is one of the three headline findings in the abstract and conclusions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a large-scale descriptive mapping of the Italian public Telegram ecosystem. The authors collected 186 million messages from 15,378 chats in 2023, built a directed weighted forwarding network, detected 15 communities with Louvain, labeled chat topics with Mixtral:8x7B, labeled political leaning with ChatGPT-4o, scored toxicity and identity attacks with the Perspective API, and identified hate-speech targets with Llama-3.1-Nemotron-70B. The main empirical claims are: (i) strong thematic and ideological homophily, including a mixed far-left/far-right geopolitical community (Warfare); (ii) toxicity is widely spread across chats in highly toxic communities rather than concentrated in a few outliers; and (iii) hate speech consistently targets Black, Jewish, and gay individuals across all communities, together with intra-national hostility against Italians.","tokens_in":18064,"tokens_out":2500,"duration_ms":26993,"significance":"If the findings hold, this would be a valuable first comprehensive map of a national Telegram ecosystem, providing a rare bird's-eye view that connects community structure, political orientation, toxicity, and hate-speech targets. The paper's strengths include the size and scope of the dataset, the explicit research questions, transparently reported prompts, and the commitment to release an anonymized network. The descriptive network construction and community detection are straightforward and largely reproducible. However, the three headline conclusions rest on two unvalidated LLM-based labeling pipelines and a single toxicity API with no Italian-language validation, and one of the main toxicity-distribution claims is internally contradicted by the reported Gini values. These issues make the current evidence insufficient to support the central claims as stated.","major_comments":[{"comment":"The political-orientation labels that drive RQ1 and the ideological-homophily and mixed-ideology findings are produced by ChatGPT-4o with no human validation, no agreement metrics, and no robustness checks. A community is classified as political if at least 50% of its chats receive a political label, but the accuracy of those chat-level labels for Italian-language chats is never established. Since the far-left/far-right coexistence in Warfare and the classification of AltNews as far-right and Activism as far-left are headline results, the absence of any validation of the labeling pipeline is load-bearing. The authors should report a human-annotated validation sample with inter-annotator agreement, or at minimum a sensitivity analysis over the 50% threshold and the labeling prompt.","section":"Section 4.3, Table A2"},{"comment":"The claim that toxicity is 'widely normalized within highly toxic communities' and 'evenly spread across many chats' is not supported by the reported Gini indices, which all lie between 0.6 and 0.9. A Gini index in this range indicates substantial concentration: a Gini near 0.7 is compatible with roughly 70% of chats carrying zero toxicity and 30% carrying all toxicity, and a Gini near 0.9 with even stronger concentration. The negative correlation (Pearson r = -0.872) shows only that high-toxicity communities are relatively less concentrated in chat-level toxicity than low-toxicity communities; it does not establish that toxicity is shared by most chats in any absolute sense. Moreover, the analysis uses unweighted chat-level toxicity percentages, so a 10-message chat and a million-message chat contribute equally, and the Gini cannot rule out that a few large chats account for most toxic messages. The authors should report the share of chats with at least one toxic message, the message-weighted concentration, and the full distribution, not just the Gini mean correlation.","section":"Section 5.3, Figure 6"},{"comment":"The toxicity and hate-speech findings depend on Perspective API scores with a threshold of 0.7, but no validation of this threshold or of the API's accuracy on Italian-language Telegram messages is provided. Likewise, the identity-target extraction by Llama-3.1-Nemotron-70B and the subsequent ChatGPT-4o label standardization have no human evaluation or agreement metrics. Given that the paper's cross-community claims about Black, Jewish, and gay targets are potentially sensitive and are used to draw conclusions about Italian online discourse, the authors should provide precision/recall or agreement numbers against a human-annotated Italian sample, and should report how sensitive the results are to the 0.7 threshold.","section":"Sections 4.4 and 4.5, Figures 7-11"},{"comment":"The Mann-Whitney U tests in Figure 5 compare each community's chat-toxicity distribution against the distribution across all chats in the network, but the community's own chats are included in the comparison distribution. This creates a non-independence problem that can bias the reported p-values, especially for large communities such as General, which contains about a third of the chats. The authors should either use a hold-out baseline excluding the community being tested or report effect sizes and confidence intervals instead of relying on these significance stars.","section":"Section 5.3, Figure 5"}],"minor_comments":[{"comment":"The number of chats is inconsistent across the paper: the abstract says 13,151 chats, Section 3.5 says 15,378 chats, and Section 4.1 reports 13,144 nodes after removing disconnected chats and 11,305 after community filtering. Please clarify which number corresponds to which stage of the pipeline.","section":"Abstract and Sections 3.5, 4.1"},{"comment":"The caption of Figure 5b says 'statistical significance levels indicate communities with notably higher or lower toxicity than average,' but the panel shows identity attack percentages, not toxicity. Please correct the wording.","section":"Section 5.3, Figure 5 caption"},{"comment":"There are several typographical and grammatical errors, including 'A the end of this process' (Section 3.5), 'corresponfing' (Section 5.3), 'rethoric' (Section 5.2), and 'hl(e.g.,' (Section 5.4). A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The topic-labeling prompt allows up to three categories per chat, but the community-level analysis appears to use only the majority topic. Please clarify how ties and multi-label outputs are handled when aggregating to the community level.","section":"Section 4.2, Table A1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a largely descriptive empirical study whose central claims are plausible but currently rest on unvalidated LLM labels and an over-interpreted Gini analysis. The weaknesses are fixable within the manuscript's scope: adding a focused validation study, correcting the toxicity-normalization interpretation, and fixing the statistical test procedure would substantially strengthen the contribution. I do not see grounds for rejection, but I would not accept the paper in its present form. The relationship to the authors' earlier arXiv paper [4] should be clarified in the revision, particularly how much of the data-collection and network-construction methodology is reused and what is genuinely new here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what to know: this is a real empirical contribution. It is the first whole-ecosystem map of a national Telegram sphere at this scale, and the descriptive results are worth having. Second, the paper's most prominent interpretive claim—that toxicity is widely normalized in highly toxic communities—does not survive contact with its own Figure 6. Read this for the map and the descriptive statistics, not for the normalization conclusion.\n\nThe data collection is the strong part. The snowball forwarding method is inherited from the authors' earlier paper, but that is fine; method reuse is not a flaw. The result is 186M messages across 15,378 public chats, with clear community labels and a sensible distinction between channels, groups, and linked chats. The finding that political discourse is concentrated in three communities, including a mixed far-left/far-right Warfare cluster, is interesting and plausible. The cross-community hate-target results (Italians, Black people, Jewish people, gay people) are descriptive but useful for future work. The limitations section honestly acknowledges the public-only scope. The promised anonymized release is good, though currently conditional on acceptance.\n\nThe soft spots are real. Political labels come entirely from ChatGPT-4o with no human validation, no agreement metrics, and a hand-set rule that 50% of chats must be political. The hate-target pipeline stacks Perspective's identity-attack score, Nemotron classification, and ChatGPT-4o label mapping, again with no Italian-language validation. Perspective is known to be biased across languages, and no threshold sensitivity is reported. The 15 Mann-Whitney tests against the global mean are not corrected for multiple comparisons. These are fixable, but they are load-bearing for RQ1 and RQ3.\n\nOn RQ2, the stress-test note is correct and important. Every Gini index in Figure 6 is between 0.6 and 0.9. A Gini of 0.7 already means most chats carry little or none of the toxicity; 0.9 is extreme concentration. The negative correlation only says that high-toxicity communities are less unequal than low-toxicity ones—not that toxicity is evenly shared across most chats. The abstract and conclusion go further than the data. In addition, the chat-level toxicity percentages are unweighted by message count, so a 10-message chat and a million-message chat count equally; the Gini cannot rule out a few large chats driving the toxicity. This is an internal inconsistency, not just a missing validation.\n\nWho is this for? Computational social scientists studying Telegram, moderation, Italian political communication, or online hate. It deserves a serious referee. I would send it to peer review, but with a request for a validation appendix, corrected multiple-comparison statistics, and a rewritten RQ2. I would cite it for the map; I would not cite the normalization claim as established.","headline":"A genuinely new national Telegram map with useful descriptive findings, but the headline toxicity-normalization claim is not supported by the paper's own Gini values.","tokens_in":18558,"tokens_out":2498,"would_cite":true,"duration_ms":30185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Italian Telegram communities split by ideology, share hate targets.","keywords":["Telegram","Italian-language social media","community detection","ideological homophily","toxicity normalization","hate speech targets","forwarding network","online political polarization"],"falsifier":"Recompute the community-level regression of mean toxicity on Gini index after removing the Adult community, and separately after removing the largest General community; if the negative correlation (Pearson = -0.8721) loses significance in either removal, the 'toxicity is normalized' conclusion depends on a single community rather than a general mechanism.","tokens_in":17521,"feed_emoji":"💬","tokens_out":7141,"duration_ms":70576,"temperature":0.7,"pith_summary":"The paper sets out to map an entire national Telegram sphere rather than a single fringe phenomenon. Using 186 million Italian-language messages from 13,151 public chats collected across 2023, it builds a forwarding network, detects 15 communities, and labels each chat's topic, political leaning, toxicity, and hate-speech targets. It aims to show that political communities split along far-left and far-right lines except for a geopolitical Warfare community where both extremes converge; that toxicity in the most toxic communities is spread evenly across many chats rather than concentrated in a few; and that Black, Jewish, and gay people are the most consistent targets of hate. These findings matter because they locate online harm not only in extremist bubbles but also in mainstream entertainment and sports spaces, and because they tie Italian online hostility to long-standing internal regional divisions.","feed_headline":"Italian Telegram communities split by ideology, share hate targets","feed_subtitle":"War news mixes far-left and far-right; toxicity is spread, not concentrated; Black, Jewish, gay users are top hate targets.","key_machinery":"The central object is a directed, weighted forwarding network G = (N, E), where each node is a Telegram chat and each edge n → m carries the number of messages forwarded from n to m; communities are detected with a weighted directed Louvain algorithm. The quantitative engine for the toxicity claim is the pairing of each community's mean toxicity percentage with the Gini index of toxicity across its chats: the negative correlation between the two is what elevates the observation 'toxic chats exist' to the structural claim 'toxicity is normalized.' Toxicity is measured with the Perspective API's Toxicity attribute at a score threshold of 0.7. Topic labels come from Mixtral, political labels from ChatGPT-4o with a 50-percent chat threshold for calling a community political, and hate targets from an instruction-tuned LLM whose free-form outputs are mapped to standardized identity categories.","core_discovery":"The paper's central claim is that the Italian public Telegram ecosystem, sampled through message forwarding, is structured by thematic and ideological homophily: chats that forward to each other cluster into 15 communities, and the three political communities separate into a far-right alternative-news cluster (AltNews), a far-left activism cluster (Activism), and a geopolitical Warfare community in which far-left and far-right rhetoric coexist, especially around Ukraine and Israel. A second claim is structural: comparing each community's mean toxicity with the Gini index of toxicity across its chats yields a strong negative correlation (Pearson = -0.8721, $R^{2}$ = 0.7605, p < 0.0001), so the most toxic communities are toxic because many chats are moderately toxic rather than because a few chats are extremely toxic. Third, hate-speech targets are stable across communities: Black and African American people, Jewish people, and gay men (the term often standing in for the whole LGBTQ+ community) are attacked everywhere, while nationality-based hate is context-dependent and includes a striking pattern of Italians attacking other Italians along regional lines.","pith_inferences":["If the political labels are validated by human annotation, the Warfare mixing result suggests a testable mechanism: attention to the Ukraine and Israel conflicts may override domestic ideological divides, so the same mixed community should appear in other national Telegram ecosystems during the same period.","The Gini-toxicity relationship is framed cross-sectionally; a longitudinal extension could test whether communities become toxic by a few chats first (high Gini) and then spread (low Gini), which would give the 'normalization' claim a temporal direction the current data cannot support.","The 'gay' target dominance may partly be an artifact of Italian hate vocabulary using 'gay' generically; a lexicon study could separate attacks aimed at gay men specifically from slurs aimed at the whole LGBTQ+ community.","The intra-Italian hostility result suggests that for Telegram studies in other countries, nationality-based hate should be broken into subnational and regional categories rather than a single 'compatriot' target."],"forward_implications":["Moderation in highly toxic Italian Telegram communities should be community-wide cultural intervention, since toxicity is spread across many chats; targeted bans on a few outlier chats would miss most of the harm.","Entertainment, adult, and sports communities carry toxicity comparable to or above political spaces, so safety research and policy cannot concentrate only on extremist political channels.","Because Black, Jewish, and gay people are attacked consistently across all communities, hate-speech detection on Italian Telegram should treat these groups as default high-priority targets independent of topic.","The coexistence of far-left and far-right rhetoric in the Warfare community means geopolitical crises can create cross-spectrum convergence, so studies of polarization should measure mixed-ideology communities, not only single-leaning clusters.","The attack pattern of Italians on other Italians indicates that national identity is not a protective category in this ecosystem; regional and intra-national frames belong in hate-speech taxonomies for Italy."],"supporting_citations":[{"why":"supplies the snowball-sampling method that follows forwarded messages to discover chats and uses forwarding as a homophily proxy","marker":"[4]"},{"why":"provides the directed weighted Louvain algorithm used to detect the 15 communities","marker":"[33]"},{"why":"labels each chat's topic from name, description, and sampled messages","marker":"[34]"},{"why":"assigns the seven-point political leaning label to each chat and later standardizes hate-target terms","marker":"[37]"},{"why":"scores message toxicity and identity attacks through the multilingual Perspective API","marker":"[42]"},{"why":"supplies the concept that toxicity becomes normalized within communities, used to interpret the Gini result","marker":"[52]"},{"why":"references the Nemotron instruction-tuned model used to extract the identity groups targeted by hate messages","marker":"[47, 48]"}],"fun_headline_variants":["Italian Telegram: ideological clusters, shared hate targets","Toxicity is widespread, not extreme, in Italian Telegram","Far-left and far-right unite on geopolitics in Italian Telegram","Italians attack Italians: regional hostility on Telegram","186M Italian Telegram messages map hate and homophily"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's ideological mapping (the far-left/far-right Warfare mix and the far-right/far-left labels) assumes the ChatGPT-4o political labels, applied to chat name, description, and a random 5,000-character message sample with no human agreement check, are accurate enough that labeling a community 'political' when at least half of its chats show a leaning does not distort the result.","fun_headline_variants_meta":{"raw":{"variants":["Italian Telegram: ideological clusters, shared hate targets","Toxicity is widespread, not extreme, in Italian Telegram","Far-left and far-right unite on geopolitics in Italian Telegram","Italians attack Italians: regional hostility on Telegram","186M Italian Telegram messages map hate and homophily"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1320,"prompt_tokens":1037,"completion_tokens":283,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":203}},"tokens_in":653,"tokens_out":283,"duration_ms":3362,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:48:06.950763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the community-level regression of mean toxicity on Gini index after removing the Adult community, and separately after removing the largest General community; if the negative correlation (Pearson = -0.8721) loses significance in either removal, the 'toxicity is normalized' conclusion depends on a single community rather than a general mechanism.","supporting_citations":[{"cited_title":"Unraveling the Italian and English Telegram Conspiracy Spheres through Message Forwarding","cited_arxiv_id":"2404.18602","evidence_quote":"supplies the snowball-sampling method that follows forwarded messages to discover chats and uses forwarding as a homophily proxy"},{"cited_title":"PhD thesis, Universit´e d’Orl´eans (2015)","cited_arxiv_id":null,"evidence_quote":"provides the directed weighted Louvain algorithm used to detect the 15 communities"},{"cited_title":"Accessed: 2024-02-10 (2024)","cited_arxiv_id":null,"evidence_quote":"assigns the seven-point political leaning label to each chat and later standardizes hate-target terms"},{"cited_title":"In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp","cited_arxiv_id":null,"evidence_quote":"scores message toxicity and identity attacks through the multilingual Perspective API"},{"cited_title":"In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp","cited_arxiv_id":null,"evidence_quote":"supplies the concept that toxicity becomes normalized within communities, used to interpret the Gini result"}],"review_version":1}