{"id":"879756ed-6e7d-4a9b-8e60-714acfac1752","arxiv_id":"2507.13398","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A descriptive study of German conspiracy Telegram chats finds event-driven activity spikes, top-heavy forwarding concentrated in 10% of chats, top-down flow into regional chats, and 42.7% of links to untrustworthy domains.","lead":"Researchers analyzed 57 million German-language Telegram messages from the Schwurbelarchiv, mapping how conspiracy-related activity, geography, and message forwarding behaved during the COVID-19 pandemic. They found that a small set of channels dominates content spread, that national and transnational chats feed regional groups more than the reverse, and that 42.7% of shared links point to domains NewsGuard scores as untrustworthy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline '42.7% untrustworthy links, far exceeding other platforms and contexts' rests on an unmatched comparison to political-elite Twitter data; the correct within-Telegram control is available but not used.","rationale":"The paper is a solid descriptive study of a new large Telegram corpus. The network findings (top 10% of chats accounting for 94% of forwards, disassortative forwarding structure, slower content decay than Twitter) are clearly presented and do not depend on the contested cross-platform comparison. The authors also explicitly acknowledge key limitations: about 50% estimated coverage, exclusion of private chats and deleted messages, and likely overestimation of regional chats from name matching. Those acknowledgements support a fair reading. The specific weakness I identify is narrower than the reader's general representativeness concern: the 'far exceeding other platforms and contexts' statement is not supported by the cited comparison because it compares different actor types, platforms, time periods, and possibly denominators. This is a correctness risk for the abstract's strongest claim, but it is addressable by re-running the link-quality analysis on the Mohr dataset, which the authors already use for other comparisons. If the Mohr-based control shows a similar or higher untrustworthy-link share, the claim would need to be weakened; if it shows a much lower share, the original claim is strengthened. The reader's conditional verdict remains appropriate: the paper should be accepted only after this matched comparison is added or the cross-platform language is qualified. I therefore keep the verdict unchanged rather than moving to rejection, because the descriptive in-corpus statistics and the network analyses are valuable and largely independent of the comparison flaw.","tokens_in":21455,"tokens_out":4526,"duration_ms":60220,"concrete_test":"Compute the same NewsGuard untrustworthy-link proportion on the Mohr (2023) German-language Telegram dataset that the authors already use for comparison in Section 4.2, using identical link extraction, URL normalization, NewsGuard thresholds, and the same observation window (September 2015-August 2022). Report the proportion for (i) all Mohr German chats, (ii) Mohr chats that also appear in the Schwurbelarchiv, and (iii) Schwurbelarchiv-only chats. If (iii) is close to (i), the 42.7% reflects conspiracy Telegram discourse rather than corpus selection; if (iii) substantially exceeds (i), the headline should be reworded as a property of the sampled corpus. Additionally, report the share of links in the denominator that point to NewsGuard-rated domains; if that share is below roughly 70%, recompute the headline proportion restricted to rated domains and compare with Lasser et al.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing defect is not the raw 42.7% figure for the sampled corpus, but the abstract's claim that this proportion is 'far exceeding their share on other platforms and in other discourse contexts' (Abstract; Section 4.4). The comparison is to Lasser et al. (2022), which measured untrustworthy-link shares among links posted by politicians on Twitter over 2016-2022. The populations differ (political elites vs. conspiracy Telegram chats), the platforms differ, the time windows differ, and the denominator rules may differ (all links vs. links to rated news domains). Any of these differences could produce the gap, so the comparison does not identify Telegram or conspiracy discourse as the causal explanation. The paper already has the right control in hand: in Section 4.2 it compares its regional chat counts against Mohr's 23,000-chat German Telegram dataset, but it never computes link-quality statistics on that dataset. Without such a matched comparison, the headline's cross-platform and cross-context claim is unsupported. The snowball-sampling concern raised by the reader compounds this: if the human-in-the-loop selection over-represented chats known to share low-quality links, the 42.7% in-corpus statistic itself could be inflated. These are fixable with additional analysis, not fatal flaws, so the conditional verdict stands.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a descriptive, observational analysis of the Schwurbelarchiv, a corpus of German-language Telegram messages from conspiracy-related public chats, spanning roughly 2015 to mid-2022. It addresses four research questions: temporal activity dynamics during COVID-19; the regional, national, and transnational distribution of chats and message flows; the concentration of forwarding activity and its implications for influence; and the prevalence of links to untrustworthy sources as rated by NewsGuard. The main quantitative findings are that message activity peaks coincide with major socio-political events, that 94% of forwarded content originates from the top 10% of spreader chats, that information flows predominantly from transnational and national chats into regional chats, and that 42.7% of the 2,308,880 links in the corpus point to NewsGuard-rated untrustworthy domains. The authors claim this share 'far exceeds' the share seen on other platforms and in other discourse contexts, citing a comparison with links shared by political elites on Twitter.","tokens_in":21629,"tokens_out":3834,"duration_ms":45552,"significance":"If the central descriptive findings are accepted, this would be a valuable large-scale contribution to the empirical literature on conspiracy-related Telegram discourse. The paper uses a novel corpus that is substantially larger than most existing studies, reports descriptive statistics directly from scraped messages without fitted parameters, makes the forwarding network publicly available, and explicitly discusses data coverage and limitations. The regional-to-national flow analysis and the concentration of forwarding activity are concrete, falsifiable descriptions that could inform future work on platform governance and misinformation. The significance of the headline misinformation claim, however, depends on two things the paper does not currently establish: that the snowball-sampled corpus represents German conspiracy Telegram discourse, and that the cross-platform comparison is valid. These issues are fixable, but they are load-bearing for the abstract's strongest claim.","major_comments":[{"comment":"The claim that the 42.7% untrustworthy-link share is 'far exceeding their share on other platforms and in other discourse contexts' is not supported by the comparison offered. The comparison population in Lasser et al. (2022) is political elites on Twitter over 2016-2022, whereas the present corpus is conspiracy-focused Telegram chats over a different time window; the platform, the population, and the possible denominator rules all differ. Any of these differences could account for the gap, so the abstract's cross-platform and cross-context claim overreaches. The authors should either restrict the claim to the sampled corpus or provide a matched comparison, for example by computing the same link-quality statistics on the Mohr (2023) dataset already used for comparison in Section 4.2.","section":"Abstract; Section 4.4"},{"comment":"The denominator for the 42.7% figure is ambiguous. The text states that 'among the 2,308,880 links posted in the Telegram chats, 42.7% point to domains classified as not trustworthy by NewsGuard,' but NewsGuard rates news and information domains, not all URLs. If the denominator is all links, the authors should explain how non-news links (e.g., YouTube, Twitter, shopping links) were classified; if the denominator is limited to links to NewsGuard-rated domains, that denominator should be reported explicitly and used consistently when comparing with Lasser et al. (2022), whose denominator rules may differ.","section":"Section 4.4"},{"comment":"The robustness of the headline 42.7% statistic to sampling selection is not assessed. The paper acknowledges that the Schwurbelarchiv likely contains about 50% of the relevant discourse and that chats were selected via snowball sampling with a human-in-the-loop component focused on conspiracy-related content. If that selection over-represented chats that disproportionately share low-quality links, the in-corpus statistic itself could be inflated. Because RQ4 and the abstract's central claim depend on this number, the authors should provide a sensitivity analysis, for example by computing link-quality statistics on the Mohr (2023) dataset or on clearly defined subsets of the corpus, and should explicitly discuss how selection could affect the estimate.","section":"Section 3.1; Section 5.0.1"}],"minor_comments":[{"comment":"The abstract reports '43%' of links while Section 4.4 reports 42.7%; these should be made numerically consistent.","section":"Abstract; Section 4.4"},{"comment":"There is a typo in the first sentence of the abstract: 'the structure of of conspiracy-related.'","section":"Abstract"},{"comment":"The caption describes 'active authors' while the y-axis label reads 'Active Groups'; please correct the mismatch.","section":"Figure 2 caption"},{"comment":"The table reports forwarding probabilities but does not specify the denominator; the authors should state explicitly whether probabilities are conditional on all forwarded messages or on messages originating from the source chat type.","section":"Table 1"},{"comment":"The text refers to a 'Section Ethics and Data Protection' that does not appear in the manuscript; this cross-reference should be removed or the section added.","section":"Section 5.0.1"},{"comment":"Several city names in the appendix lists are concatenated without spaces (e.g., 'BadKreuznach', 'VillingenSchwenningen'); please clarify whether these are intentional matching patterns and how they interact with the space/punctuation boundary rule described in Section 3.3.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong descriptive contribution for a computational social science venue, and the corpus release and public forwarding network are assets. My main concern, reflected in the major comments, is that the abstract's strongest claim about cross-platform comparison is not supported by the analysis as presented. I would be comfortable with acceptance after the authors either add the matched analysis using the Mohr data they already have or substantially weaken the cross-platform framing. The checklist-style appendix appears to be a template artifact and should be removed from the final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a solid descriptive map of a large German-language Telegram conspiracy corpus during COVID-19. What's actually new: linking geographical scope classification (regional/national/transnational) to forwarding asymmetries, and applying NewsGuard link-quality ratings to a corpus of this size. The finding that regional chats receive more national-to-regional flow than the reverse (0.2% vs 0.04%) is interesting, and the 94% of forwards from top 10% of chats quantifies concentration cleanly. The temporal spike analysis is plausible, though manually matched to events.\n\nThe authors do several things right. They benchmark their corpus against Mohr and Zehring datasets, state the ~50% coverage estimate, and explicitly acknowledge most limitations in the discussion. The regional chat identification via names is honestly flagged as likely overestimating regional prevalence. The message longevity analysis is straightforward and the comparison to Pfeffer's Twitter half-life is properly hedged in the text itself.\n\nThe main soft spot is the headline 42.7% untrustworthy-link share. The raw figure for this corpus is what it is, but the abstract's claim that this 'far exceeds' other platforms and contexts rests on a comparison to Lasser et al.'s political-elite Twitter data. That population differs in platform, time window, and denominator rules, so the gap doesn't license the causal read. The paper has the right control in hand—Mohr's German Telegram dataset—and computes nothing on it beyond chat counts. A matched link-quality analysis there would make the cross-context claim credible. Second, the snowball sample with human-in-the-loop selection could overrepresent chats known for sharing low-quality links; the authors estimate 50% coverage but don't test how selection affects link quality. These are fixable with additional analysis, not fatal flaws. Reproducibility is limited: no code, and data URLs deleted for anonymity.\n\nOverall, the descriptive findings are internally consistent and clearly presented. The paper deserves a serious referee; it's a useful contribution to the within-subfield literature. I'd bring it to reading group and would cite the forwarding asymmetry and concentration numbers, though not the cross-platform comparison.","headline":"Solid descriptive map of a German Telegram conspiracy corpus, but the headline misinformation claim leans on an unmatched Twitter comparison.","tokens_in":22216,"tokens_out":1842,"would_cite":true,"duration_ms":18416,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"German conspiracy-related Telegram chats during the pandemic are dominated by a small number of spreader channels and carry an unusually high share of untrustworthy links: 42.7% of shared links point to domains that NewsGuard scores below…","keywords":["conspiracy theories","Telegram","misinformation","COVID-19","network analysis","NewsGuard","Germany","information flow"],"falsifier":"Collect an independent census of German-language public Telegram chats from the same period—for example, starting from Telegram's own search or from the larger Mohr corpus—and measure the share of links to NewsGuard-rated untrustworthy domains and the forwarding concentration. If the share drops well below 42.7% or the top-decile concentration falls far below 94% once the snowball seed is removed, the paper's headline claims would be artifacts of sampling rather than properties of the discourse.","tokens_in":21212,"feed_emoji":"🔗","tokens_out":6396,"duration_ms":60768,"temperature":0.7,"pith_summary":"This paper tries to establish a structural portrait of German-language conspiracy discourse on Telegram during the COVID-19 pandemic, using a large scraped corpus of public chats. It argues that the discourse is event-driven, peaks during major political and health events, and is concentrated: the top 10% of chats account for 94% of all forwarded messages. It also claims that information flows predominantly from national and transnational chats down to regional groups, rather than the reverse. Its most consequential quantitative claim is that 42.7% of the 2.3 million links shared point to domains that NewsGuard rates as untrustworthy, a share far above the single-digit percentages seen among political elites on Twitter. If these claims hold, Telegram's lightly moderated public channels function as a major vector for misinformation in German-speaking countries.","feed_headline":"43% of links in German conspiracy Telegram chats are untrustworthy","feed_subtitle":"A large new dataset shows a few spreader chats dominate and content flows down to regional groups.","key_machinery":"The analysis rests on three instruments. The Schwurbelarchiv, a snowball-sampled corpus of roughly 6,000 public German-language Telegram chats (an estimated 50% of the relevant discourse), provides messages, forwards, authors, and timestamps from September 2015 to August 2022. A forwarded-message network, built by matching forwarded messages to their originals through author, text, and timestamp, carries the structural analysis. Trustworthiness is operationalized through NewsGuard domain ratings, with scores below 60 counted as untrustworthy, and the same operationalization is used to compare against Twitter data from the cited literature.","core_discovery":"The central discovery, as the authors frame it, is that conspiracy-related German Telegram discourse during the pandemic is not a decentralized many-to-many conversation but a concentrated structure: a small number of broadcast channels produce most forwarded content, while the many groups mostly receive it. The dataset reveals that chats with high out-degree (spreader chats) do not forward much to each other and instead send content to low-traffic groups, while chats that receive many forwards tend to cluster together. At the same time, activity spikes align with real-world events such as the storming of the U.S. Capitol and the inauguration of Joe Biden, and forwarding longevity on Telegram outlasts that on Twitter. The paper's headline figure is that 42.7% of 2,308,880 links in the corpus point to domains that NewsGuard scores below 60 ('not trustworthy'), classifying the ecosystem as a dense misinformation vector.","pith_inferences":["The 42.7% figure depends on NewsGuard's binary threshold; applying other domain-quality ratings (which the paper cites) could shift the share significantly, so a multi-rater robustness check would settle how stable the number is.","Because the corpus covers only about 50% of the discourse and excludes private chats, the untrustworthy-link share could be systematically higher or lower; the paper's own remark that deleted messages likely skew untrustworthy suggests the true share may be understated.","Re-running the same forwarding-concentration and link-quality analysis on the independent Mohr dataset, which the authors already use for completeness benchmarking, would test whether the snowball-selection procedure biases the headline results.","The regional flow asymmetry implies a plausible mechanism for narrative spread—national channels broadcast, regional groups adopt—that could be tested by tracking specific conspiracy narratives from their first appearance in a national channel to their echo in local groups."],"forward_implications":["If 42.7% of shared links are untrustworthy, then public Telegram chats in German-speaking countries were a high-density misinformation vector during the pandemic, and the same method could quantify similar ecosystems in other languages.","Because the top 10% of chats produce 94% of forwarded content, interventions that reach or counter those few channels could plausibly reduce most content diffusion.","The finding that national and transnational messages are more likely to reach regional chats than the reverse implies that local conspiracy communities are partly seeded by super-regional narratives.","The slower decay of forwards on Telegram relative to Twitter suggests misinformation lingers longer where algorithmic recommendation is absent.","The chat-typing heuristic (usernames containing the chat name) and the geographic matching by chat names offer a transferable method for other unmoderated messenger corpora."],"supporting_citations":[{"why":"Companion paper describing the Schwurbelarchiv dataset, its collection via snowball sampling with human-in-the-loop, and the ~50% coverage estimate.","marker":"Angermaier et al. [2025]"},{"why":"Provides the domain trustworthiness ratings that operationalize misinformation; scores below 60 define 'not trustworthy' links.","marker":"NewsGuard [2020]"},{"why":"Supply the Twitter baseline for the share of links to untrustworthy sources, used to argue Telegram's proportion is far higher.","marker":"Lasser et al. [2022]"},{"why":"Provide the Twitter half-life benchmark for content decay, used to argue Telegram forwards last longer.","marker":"Pfeffer et al. [2023]"},{"why":"Independent larger Telegram corpus used to benchmark the Schwurbelarchiv's regional coverage and completeness.","marker":"Mohr [2023]"},{"why":"Closest prior network analysis of German conspiracy-related Telegram chats; its dataset is compared for regional coverage.","marker":"Zehring and Domahidi [2023]"},{"why":"Documents that deleted messages are common on Telegram, supporting the limitation that the corpus misses such content.","marker":"Buehling [2024]"}],"fun_headline_variants":["43% of links in German conspiracy chats are untrustworthy","German conspiracy Telegram: few spreader chats dominate","Top 10% of conspiracy chats fuel 94% of forwards","Telegram conspiracy activity spikes with real-world events","Study: conspiracy Telegram ecosystem is a misinformation vector"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the Schwurbelarchiv's snowball collection, with a human selecting conspiracy-related public chats, captures a representative slice of German-language conspiracy Telegram discourse, so that aggregate statistics like the 42.7% untrustworthy-link share and the 94% forwarding concentration generalize beyond the sampled chats.","fun_headline_variants_meta":{"raw":{"variants":["43% of links in German conspiracy chats are untrustworthy","German conspiracy Telegram: few spreader chats dominate","Top 10% of conspiracy chats fuel 94% of forwards","Telegram conspiracy activity spikes with real-world events","Study: conspiracy Telegram ecosystem is a misinformation vector"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000569,"raw_usage":{"total_tokens":2708,"prompt_tokens":977,"completion_tokens":1731,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":1653}},"tokens_in":593,"tokens_out":1731,"duration_ms":14912,"temperature":1.0,"reasoning_tokens":1653,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:51:28.224612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect an independent census of German-language public Telegram chats from the same period—for example, starting from Telegram's own search or from the larger Mohr corpus—and measure the share of links to NewsGuard-rated untrustworthy domains and the forwarding concentration. If the share drops well below 42.7% or the top-decile concentration falls far below 94% once the snowball seed is removed, the paper's headline claims would be artifacts of sampling rather than properties of the discourse.","supporting_citations":[{"cited_title":"Newsguard rating process and criteria (internet archive), 2020","cited_arxiv_id":null,"evidence_quote":"Provides the domain trustworthiness ratings that operationalize misinformation; scores below 60 define 'not trustworthy' links."},{"cited_title":"Inference of modular structures in dynamical systems and the application to telegram data","cited_arxiv_id":null,"evidence_quote":"Independent larger Telegram corpus used to benchmark the Schwurbelarchiv's regional coverage and completeness."},{"cited_title":"German corona protest mobilizers on telegram and their relations to the far right: A network and topic analysis","cited_arxiv_id":null,"evidence_quote":"Closest prior network analysis of German conspiracy-related Telegram chats; its dataset is compared for regional coverage."}],"review_version":1}