REVIEW 3 major objections 4 minor 30 references
TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper introduces TeleScope, a public dataset of about 120 million Telegram messages from 71,000 public channels, with reconstructed forwarding paths that let researchers trace how content spreads.
desk verdict Useful and genuinely novel Telegram dataset with forwarding flows, but the 'largest' claim is not supported by the paper's own numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is Telegram's own message-forwarding feature, used twice: once as a discovery engine and once as a tracing tool. The crawler starts from 251 seed channels—selected from a public registry's top-100 lists ranked by subscribers, citations, and reach—and whenever it sees a forwarded message, it adds the source channel and recurses. This snowball sampling both expands the channel list to over 1.2 million discovered channels and produces the backward traces needed to construct message forwarding flows and the channel-to-channel graph. The graph is the load-bearing derived object: it turns a flat collection of messages into a network along which information spread can be measured.
What would settle it
A reader could draw an independent random sample of Telegram public channels—say, through Telegram's own search or an unrelated directory—and check whether TeleScope's language distribution and forwarding graph match that sample; a large mismatch, such as most non-Russian channels being absent, would falsify the representativeness implied by the snowball convergence.
Extended reading notes
Core claim
The central claim is that a large, longitudinal, multilingual Telegram corpus can be assembled from public data and enriched so that message propagation is visible instead of opaque. The authors present TeleScope as the largest publicly available collection of its kind: 534,137 discovered channels, 71,048 fully downloaded public channels, and 120,024,020 messages, spanning February 1 to October 29, 2024. The novel part is the derived layer—message forwarding flows that trace each forwarded message back to its source channel, a directed channel-to-channel graph with 261,171 nodes and 2,733,720 edges, and aggregated user interaction counts computed across all crawled copies of a message. On top of this, TeleScope adds language labels, per-channel hourly activity, and extracted Telegram entities such as hashtags, mentions, URLs, and formatting spans. The authors argue that these pieces together make Telegram research possible at the level of detail previously limited to Twitter-style platforms.
Load-bearing premise
The entire dataset starts from 251 channels picked from the top-100 lists of a public registry, and the paper assumes that snowballing from these seeds eventually gives a representative view of Telegram's public channels without comparing against a full census of Telegram.
Editorial extensions
If this is right
- Researchers can study message propagation at scale by using the forwarding flows instead of scraping each channel separately.
- Retweet-style analyses—virality prediction, information diffusion, and community detection—can be replicated on Telegram using the channel graph and aggregated interaction scores.
- The 47 detected languages, including low-resource ones, open a path for multilingual NLP and cross-community comparisons not feasible with earlier Telegram datasets.
- Public channel metadata for 534K channels, including creation dates and flag statuses, supports longitudinal studies of platform growth and content moderation.
- Planned annual snapshots would let later work track how channels, forwarding networks, and communities evolve over time.
Reading between the lines
- The convergence of the three seed criteria shown in the paper is evidence that snowballing saturates around the seed region, not that it covers all Telegram; a random-sample validation would strengthen any cross-platform generalization.
- Because the language distribution is 82% Russian, cross-lingual studies built on this dataset will need reweighting or stratified subsampling to avoid conflating 'Telegram' with 'Russian-speaking Telegram'.
- The message forwarding flows could be used to estimate cascade lengths and branching factors, but any such estimate is a lower bound, since forwards that leave the 71K crawled channels are invisible.
- A natural next step, not demonstrated in the paper, is to join the channel-to-channel graph with extracted URLs and hashtags to study topic-level diffusion across communities.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces TeleScope, a Telegram dataset suite containing metadata for 534,137 channels, downloaded message metadata for 71,048 public channels (120,024,020 messages), and derived artefacts including a channel-to-channel forwarding graph, message forwarding flows, and aggregated user interaction statistics. The collection procedure starts from 251 seed channels selected from TGStat's top-100 lists by subscribers, citations, and reach, and expands through snowballing via Telegram's message-forwarding metadata. The dataset is released through GESIS with a DOI, and the authors provide enrichment scripts and crawler information on GitHub. The paper claims that TeleScope is the largest publicly available, multilingual, longitudinal, and continuous Telegram collection to date, and it presents exploratory statistics on language distribution, channel properties, active periods, hashtags, and user engagement.
Significance. If validated, TeleScope would be a substantial reusable resource for social media research, particularly for studies of information propagation, multilingual discourse, and low-resource language communities. The paper's strengths include a publicly shared seed list for provenance, a clearly described collection pipeline, reproducible enrichment code, FAIR-aligned hosting with a DOI, and the release of derived propagation and interaction data that are genuinely useful additions over raw Telegram metadata. However, the two most prominent claims of the paper are not yet established: the 'largest of its kind' statement is not supported by the comparators cited in the paper itself, and the snowball-sampling coverage has only been checked internally, not against an external census or independent channel registry. These issues affect the framing and generalizability of the dataset more than its raw utility.
major comments (3)
- [Abstract and §1, compared with §2] The headline claim that TeleScope is 'the largest publicly available, multilingual, longitudinal, and continuous collection from Telegram' is not supported by the paper's own cited comparators. Pushshift Telegram (Baumgartner et al. 2020) is reported as containing 317M messages, about 2.6 times TeleScope's 120M messages, and TGDataset (La Morgia et al. 2023) is reported as containing 120K channels, about 1.7 times TeleScope's 71K fully downloaded channels. The paper never states the metric on which 'largest' is judged. I ask the authors to define the comparison metric (e.g., 534K channel metadata entries, or the 261K-node forwarding graph), add a comparison table with existing Telegram datasets, and soften the claim if no dimension is actually the largest.
- [§3.1, Figure 1, and §6.1] The convergence among the three seed criteria in Figure 1 is an internal property of the snowball process; it does not establish that the resulting channel set is representative of Telegram's public channel space. The language distribution in Table 3 (82.29% Russian) and the dominance of Russia/Ukraine-related hashtags in Figure 8 indicate strong selection effects from the TGStat seed. If the paper retains the 'comprehensive data coverage' framing in §1 and the general-purpose applicability claims in §7, please add an external validation of coverage—for example, a comparison of the discovered channels against an independent channel registry, or against the channel sets of Pushshift and TGDataset, and an estimate of what fraction of active public channels the 71,048 downloaded channels represent.
- [§4.3, Table 2] The propagation statistics are not well-defined. 'Total number of messages' in Table 2 is 31,227,109, which does not match the 19.6% forwarded share of the 120M messages (about 23.5M forwarded messages), and 'Number of unique messages' (308,147) is not explained. Please define the unit of a forwarding flow, what counts as a message in the flow table, and how unique messages are identified, and report the number of forwarding events and the distribution of flow lengths so that the derived interaction data can be evaluated.
minor comments (4)
- [§8 and §5] There are a few typos: 'dataset suit' in §8 should be 'dataset suite', and 'isFindable' in §5 should be 'is Findable'.
- [§4.1] The phrase 'metadata for all the channels fully downloaded in our dataset (534,137)' is confusing because only 71,048 channels have downloaded messages; please rephrase to distinguish channel-metadata-only entries from full message downloads.
- [§6.3 and Figure 7] The active-period analysis aggregates all channels in UTC without per-channel local-time normalization; because the sample is 82% Russian, the observed morning peak may reflect a single timezone. Either normalize to channel-local time or explicitly state this as a limitation.
- [Paper Checklist, item 5(g)] The authors state that no Datasheet for the Dataset was created and that one will be provided if the paper is accepted; for a dataset contribution, a datasheet is standard practice and should be added to the final release to document provenance, biases, and intended uses.
Circularity Check
No significant circularity: TeleScope's dataset construction and enrichments are computed directly from observed Telegram API data, not from fitted parameters or self-referential definitions.
full rationale
This is a data-collection and dataset-release paper, not a derivation of a model or theory from fitted inputs. The channel seed list is a stated snowball sample from TGStat rankings; the message metadata, forwarding flows, channel-to-channel graph, and aggregated interaction statistics are computed directly from Telegram's API-observed message fields (e.g., the source of a forwarded message is backward-traced from the destination message's metadata). No equation equates an output to an input by construction, and no fitted parameter is later renamed as a prediction. The only self-citations (Fafalios et al. 2018; Gangopadhyay et al. 2024) are used to illustrate GESIS's hosting commitments for other datasets and are not load-bearing for TeleScope's central claims. The paper's 'largest of its kind' superlative is an empirical comparison claim that may be under-supported against the cited Pushshift and TGDataset numbers, but that ambiguity is a correctness or evidence concern, not a circularity: the claim is not defined in terms of the dataset's own output. The acknowledged limitation that the snowball sample may over-represent Russian-speaking channels (Section 6.1) is a coverage caveat, not a circular step. Overall, the dataset's content is self-contained external evidence, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The Telethon API returns complete and accurate message metadata for public channels.
- domain assumption The TGStat top-100 lists per subscriber, citation, and reach provide a representative seed set for snowball sampling.
- domain assumption Backward tracing of forwarded messages reconstructs actual propagation paths within the dataset.
- domain assumption The Langdetect library, applied to concatenated channel messages, estimates channel language with sufficient accuracy.
Cite this review
Pith. "Pith review of TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram." pith.science (2026). https://pith.science/paper/QYUD7242
@misc{pith2026250419536,
author = {Pith},
title = {Pith review of: TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram},
year = {2026},
howpublished = {\url{https://pith.science/paper/QYUD7242}},
note = {Machine review of arXiv:2504.19536}
}
read the original abstract
Telegram is a globally popular instant messaging platform known for its strong emphasis on security, privacy, and unique social networking features. It has recently emerged as the host for various cross-domain analysis and research works, such as social media influence, propaganda studies, and extremism. This paper introduces TeleScope, an extensive dataset suite that, to our knowledge, is the largest of its kind. It comprises metadata for about 500K Telegram channels and downloaded message metadata for about 71K public channels, accounting for around 120M crawled messages. We also release channel connections and user interaction data built using Telegram's message-forwarding feature to study multiple use cases, such as information spread and message forwarding patterns. In addition, we provide data enrichments, such as language detection, active message posting periods for each channel, and Telegram entities extracted from messages, that enable online discourse analysis beyond what is possible with the original data alone. The dataset is designed for diverse applications, independent of specific research objectives, and sufficiently versatile to facilitate the replication of social media studies comparable to those conducted on platforms like X (formerly Twitter)
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Alvisi, L.; Tardelli, S.; and Tesconi, M. 2024. Unraveling the Italian and English Telegram Conspiracy Spheres through Message Forwarding. arXiv preprint arXiv:2404.18602
arXiv 2024
-
[2]
Baumgartner, J.; Zannettou, S.; Squire, M.; and Blackburn, J. 2020. The pushshift telegram dataset. In Proceedings of the international AAAI conference on web and social media, volume 14, 840--847
work page 2020
-
[3]
Drolsbach, C. P.; and Pr\" o llochs, N. 2023. Diffusion of Community Fact-Checked Misinformation on Twitter. Proc. ACM Hum.-Comput. Interact., 7(CSCW2)
work page 2023
-
[4]
Ellis, C. H.; Moore, J. B.; Ho, P.; et al. 2024. Social network and linguistic analysis of the\# nutrition discourse on the social network platform X, formerly known as Twitter. Social Network Analysis and Mining, 14(1): 238
work page 2024
-
[5]
Elmas, T.; Stephane, S.; and Houssiaux, C. 2023. Measuring and Detecting Virality on Social Media: The Case of Twitter’s Viral Tweets Topic. In Companion Proceedings of the ACM Web Conference 2023, WWW '23 Companion, 314–317. New York, NY, USA. ISBN 9781450394192
work page 2023
-
[6]
Fafalios, P.; Iosifidis, V.; Ntoutsi, E.; and Dietze, S. 2018. Tweetskb: A public and large-scale rdf corpus of annotated tweets. In European Semantic Web Conference. Springer
work page 2018
-
[7]
Gangopadhyay, S.; Schellhammer, S.; Hafid, S.; et al. 2024. Investigating Characteristics, Biases and Evolution of Fact-Checked Claims on the Web. In Proceedings of the 35th ACM Conference on Hypertext and Social Media, 246--258
work page 2024
-
[8]
Ghaffary, S.; and Heilweil, R. 2024. How Trump’s internet built and broadcast the Capitol insurrectionHow Trump’s internet built and broadcast the Capitol insurrection. Accessed: 2024-06-06
work page 2024
Show all 30 references
-
[9]
W.; and Durumeric, Z
Hanley, H. W.; and Durumeric, Z. 2024. Partial mobilization: Tracking multilingual information flows amongst russian media outlets and telegram. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 528--541
2024
-
[10]
o hn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A new dataset for bias research from a “dark
H \"o hn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A new dataset for bias research from a “dark” platform. In Proceedings of the international AAAI Conference on Web and social media, volume 16, 1268--1274
2022
-
[11]
Hoseini, M.; de Freitas Melo, P.; Benevenuto, F.; Feldmann, A.; and Zannettou, S. 2024. Characterizing Information Propagation in Fringe Communities on Telegram. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 583--595
2024
-
[12]
Höhn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A New Dataset for Bias Research from a “Dark” Platform. Proceedings of the International AAAI Conference on Web and Social Media, 16: 1268--1274
2022
-
[13]
J.; and Carley, K
Kloo, I.; Cruickshank, I. J.; and Carley, K. M. 2024. A Cross-Platform Topic Analysis of the Nazi Narrative on Twitter and Telegram During the 2022 Russian Invasion of Ukraine. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 839--850
2024
-
[14]
La Morgia, M.; Mei, A.; Mongardini, A.; and Wu, J. 2023. It’s a Trap! Detection and Analysis of Fake Channels on Telegram. 97--104
2023
-
[15]
La Morgia, M.; Mei, A.; and Mongardini, A. M. 2023. Tgdataset: a collection of over one hundred thousand telegram channels. arXiv preprint arXiv:2303.05345
2023 arXiv
-
[16]
H.; Lee, H.; Cesare, N.; Shojaie, A.; and Spiro, E
McCormick, T. H.; Lee, H.; Cesare, N.; Shojaie, A.; and Spiro, E. S. 2017. Using Twitter for demographic and social science research: tools for data collection and processing. Sociological methods & research, 46(3): 390--421
2017
-
[17]
Ovadia, S. 2009. Exploring the potential of Twitter as a research tool. Behavioral & Social Sciences Librarian, 28(4): 202--205
2009
-
[18]
Qudar, M. M. A.; and Mago, V. 2020. Tweetbert: a pretrained language representation model for twitter text analysis. arXiv preprint arXiv:2010.11091
2020 arXiv
-
[19]
Schulze, H.; Hohner, J.; et al. 2022. Far-right conspiracy groups on fringe platforms: A longitudinal analysis of radicalization dynamics on Telegram. Convergence: The International Journal of Research into New Media Technologies
2022
-
[20]
Solopova, V.; Scheffler, T.; and Popa-Wyatt, M. 2021. A Telegram corpus for hate speech, offensive language, and online harm. Journal of Open Humanities Data, 7
2021
-
[21]
Sosa, J.; and Sharoff, S. 2022. Multimodal Pipeline for Collection of Misinformation Data from Telegram. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, 1480--1489
2022
-
[22]
Tikhomirova, K.; and Makarov, I. 2021. Community detection based on the nodes role in a network: The telegram platform case. In Analysis of Images, Social Networks and Texts: 9th International Conference, AIST 2020, 294--302
2021
-
[23]
R.; Ferreira, C
Ven \^a ncio, O. R.; Ferreira, C. H.; Almeida, J. M.; et al. 2024. Unraveling User Coordination on Telegram: A Comprehensive Analysis of Political Mobilization during the 2022 Brazilian Presidential Election. In Proceedings of the International AAAI Conference on Web and Socia...
2024
-
[24]
Walther, S.; and McCoy, A. 2021. US extremism on Telegram. Perspectives on Terrorism, 15(2): 100--124
2021
-
[25]
D.; Dumontier, M.; Aalbersberg, I
Wilkinson, M. D.; Dumontier, M.; Aalbersberg, I. J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L. B.; Bourne, P. E.; et al. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3(1): 1--9
2016
-
[26]
H.; Porter, M
Xiao, Z.; Zhu, J.; Wang, Y.; Zhou, P.; Lam, W. H.; Porter, M. A.; and Sun, Y. 2023. Detecting political biases of named entities and hashtags on Twitter. EPJ Data Science, 20
2023
-
[27]
S.; and Speckhard, A
Yayla, A. S.; and Speckhard, A. 2017. Telegram: The mighty application that ISIS loves. International Center for the Study of Violent Extremism, 9
2017
-
[28]
R.; Herbrich, R.; Van Gael, J.; and Stern, D
Zaman, T. R.; Herbrich, R.; Van Gael, J.; and Stern, D. 2010. Predicting information spreading in twitter. In Workshop on computational social science and the wisdom of crowds, nips, volume 104, 17599--601. Citeseer
2010
-
[29]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[30]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.