Pith. sign in

REVIEW 3 major objections 4 minor 30 references

TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper introduces TeleScope, a public dataset of about 120 million Telegram messages from 71,000 public channels, with reconstructed forwarding paths that let researchers trace how content spreads.

desk verdict Useful and genuinely novel Telegram dataset with forwarding flows, but the 'largest' claim is not supported by the paper's own numbers. read the letter →

arxiv 2504.19536 v1 pith:QYUD7242 submitted 2025-04-28 cs.SI

classification cs.SI
keywords Telegramdatasetmessageforwardingchannel-to-channelgraphinformationpropagationlongitudinaldatamultilingualcorpussocialmediaanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TeleScope is a public dataset suite that the authors say is the largest of its kind for Telegram: metadata for about 534,000 channels, complete message metadata for 71,048 public channels, and roughly 120 million messages crawled between February and October 2024. The dataset's distinctive feature is that it does not stop at raw messages: it reconstructs forwarding flows, builds a channel-to-channel graph, and aggregates views, forwards, and reactions per message across the crawled network. The authors' aim is to give social scientists the kind of fine-grained material for Telegram that Twitter once provided for studying information spread, communities, virality, and misinformation. If the resource works as advertised, it removes a major access barrier to one of the world's largest and least transparent messaging platforms.

What carries the argument

The mechanism that carries the argument is Telegram's own message-forwarding feature, used twice: once as a discovery engine and once as a tracing tool. The crawler starts from 251 seed channels—selected from a public registry's top-100 lists ranked by subscribers, citations, and reach—and whenever it sees a forwarded message, it adds the source channel and recurses. This snowball sampling both expands the channel list to over 1.2 million discovered channels and produces the backward traces needed to construct message forwarding flows and the channel-to-channel graph. The graph is the load-bearing derived object: it turns a flat collection of messages into a network along which information spread can be measured.

What would settle it

A reader could draw an independent random sample of Telegram public channels—say, through Telegram's own search or an unrelated directory—and check whether TeleScope's language distribution and forwarding graph match that sample; a large mismatch, such as most non-Russian channels being absent, would falsify the representativeness implied by the snowball convergence.

Watch

Extended reading notes

Core claim

The central claim is that a large, longitudinal, multilingual Telegram corpus can be assembled from public data and enriched so that message propagation is visible instead of opaque. The authors present TeleScope as the largest publicly available collection of its kind: 534,137 discovered channels, 71,048 fully downloaded public channels, and 120,024,020 messages, spanning February 1 to October 29, 2024. The novel part is the derived layer—message forwarding flows that trace each forwarded message back to its source channel, a directed channel-to-channel graph with 261,171 nodes and 2,733,720 edges, and aggregated user interaction counts computed across all crawled copies of a message. On top of this, TeleScope adds language labels, per-channel hourly activity, and extracted Telegram entities such as hashtags, mentions, URLs, and formatting spans. The authors argue that these pieces together make Telegram research possible at the level of detail previously limited to Twitter-style platforms.

Load-bearing premise

The entire dataset starts from 251 channels picked from the top-100 lists of a public registry, and the paper assumes that snowballing from these seeds eventually gives a representative view of Telegram's public channels without comparing against a full census of Telegram.

Editorial extensions

If this is right

  • Researchers can study message propagation at scale by using the forwarding flows instead of scraping each channel separately.
  • Retweet-style analyses—virality prediction, information diffusion, and community detection—can be replicated on Telegram using the channel graph and aggregated interaction scores.
  • The 47 detected languages, including low-resource ones, open a path for multilingual NLP and cross-community comparisons not feasible with earlier Telegram datasets.
  • Public channel metadata for 534K channels, including creation dates and flag statuses, supports longitudinal studies of platform growth and content moderation.
  • Planned annual snapshots would let later work track how channels, forwarding networks, and communities evolve over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The convergence of the three seed criteria shown in the paper is evidence that snowballing saturates around the seed region, not that it covers all Telegram; a random-sample validation would strengthen any cross-platform generalization.
  • Because the language distribution is 82% Russian, cross-lingual studies built on this dataset will need reweighting or stratified subsampling to avoid conflating 'Telegram' with 'Russian-speaking Telegram'.
  • The message forwarding flows could be used to estimate cascade lengths and branching factors, but any such estimate is a lower bound, since forwards that leave the 71K crawled channels are invisible.
  • A natural next step, not demonstrated in the paper, is to join the channel-to-channel graph with extracted URLs and hashtags to study topic-level diffusion across communities.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces TeleScope, a Telegram dataset suite containing metadata for 534,137 channels, downloaded message metadata for 71,048 public channels (120,024,020 messages), and derived artefacts including a channel-to-channel forwarding graph, message forwarding flows, and aggregated user interaction statistics. The collection procedure starts from 251 seed channels selected from TGStat's top-100 lists by subscribers, citations, and reach, and expands through snowballing via Telegram's message-forwarding metadata. The dataset is released through GESIS with a DOI, and the authors provide enrichment scripts and crawler information on GitHub. The paper claims that TeleScope is the largest publicly available, multilingual, longitudinal, and continuous Telegram collection to date, and it presents exploratory statistics on language distribution, channel properties, active periods, hashtags, and user engagement.

Significance. If validated, TeleScope would be a substantial reusable resource for social media research, particularly for studies of information propagation, multilingual discourse, and low-resource language communities. The paper's strengths include a publicly shared seed list for provenance, a clearly described collection pipeline, reproducible enrichment code, FAIR-aligned hosting with a DOI, and the release of derived propagation and interaction data that are genuinely useful additions over raw Telegram metadata. However, the two most prominent claims of the paper are not yet established: the 'largest of its kind' statement is not supported by the comparators cited in the paper itself, and the snowball-sampling coverage has only been checked internally, not against an external census or independent channel registry. These issues affect the framing and generalizability of the dataset more than its raw utility.

major comments (3)
  1. [Abstract and §1, compared with §2] The headline claim that TeleScope is 'the largest publicly available, multilingual, longitudinal, and continuous collection from Telegram' is not supported by the paper's own cited comparators. Pushshift Telegram (Baumgartner et al. 2020) is reported as containing 317M messages, about 2.6 times TeleScope's 120M messages, and TGDataset (La Morgia et al. 2023) is reported as containing 120K channels, about 1.7 times TeleScope's 71K fully downloaded channels. The paper never states the metric on which 'largest' is judged. I ask the authors to define the comparison metric (e.g., 534K channel metadata entries, or the 261K-node forwarding graph), add a comparison table with existing Telegram datasets, and soften the claim if no dimension is actually the largest.
  2. [§3.1, Figure 1, and §6.1] The convergence among the three seed criteria in Figure 1 is an internal property of the snowball process; it does not establish that the resulting channel set is representative of Telegram's public channel space. The language distribution in Table 3 (82.29% Russian) and the dominance of Russia/Ukraine-related hashtags in Figure 8 indicate strong selection effects from the TGStat seed. If the paper retains the 'comprehensive data coverage' framing in §1 and the general-purpose applicability claims in §7, please add an external validation of coverage—for example, a comparison of the discovered channels against an independent channel registry, or against the channel sets of Pushshift and TGDataset, and an estimate of what fraction of active public channels the 71,048 downloaded channels represent.
  3. [§4.3, Table 2] The propagation statistics are not well-defined. 'Total number of messages' in Table 2 is 31,227,109, which does not match the 19.6% forwarded share of the 120M messages (about 23.5M forwarded messages), and 'Number of unique messages' (308,147) is not explained. Please define the unit of a forwarding flow, what counts as a message in the flow table, and how unique messages are identified, and report the number of forwarding events and the distribution of flow lengths so that the derived interaction data can be evaluated.
minor comments (4)
  1. [§8 and §5] There are a few typos: 'dataset suit' in §8 should be 'dataset suite', and 'isFindable' in §5 should be 'is Findable'.
  2. [§4.1] The phrase 'metadata for all the channels fully downloaded in our dataset (534,137)' is confusing because only 71,048 channels have downloaded messages; please rephrase to distinguish channel-metadata-only entries from full message downloads.
  3. [§6.3 and Figure 7] The active-period analysis aggregates all channels in UTC without per-channel local-time normalization; because the sample is 82% Russian, the observed morning peak may reflect a single timezone. Either normalize to channel-local time or explicitly state this as a limitation.
  4. [Paper Checklist, item 5(g)] The authors state that no Datasheet for the Dataset was created and that one will be provided if the paper is accepted; for a dataset contribution, a datasheet is standard practice and should be added to the final release to document provenance, biases, and intended uses.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: TeleScope's dataset construction and enrichments are computed directly from observed Telegram API data, not from fitted parameters or self-referential definitions.

full rationale

This is a data-collection and dataset-release paper, not a derivation of a model or theory from fitted inputs. The channel seed list is a stated snowball sample from TGStat rankings; the message metadata, forwarding flows, channel-to-channel graph, and aggregated interaction statistics are computed directly from Telegram's API-observed message fields (e.g., the source of a forwarded message is backward-traced from the destination message's metadata). No equation equates an output to an input by construction, and no fitted parameter is later renamed as a prediction. The only self-citations (Fafalios et al. 2018; Gangopadhyay et al. 2024) are used to illustrate GESIS's hosting commitments for other datasets and are not load-bearing for TeleScope's central claims. The paper's 'largest of its kind' superlative is an empirical comparison claim that may be under-supported against the cited Pushshift and TGDataset numbers, but that ambiguity is a correctness or evidence concern, not a circularity: the claim is not defined in terms of the dataset's own output. The acknowledged limitation that the snowball sample may over-represent Russian-speaking channels (Section 6.1) is a coverage caveat, not a circular step. Overall, the dataset's content is self-contained external evidence, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim (provision of a large, versatile Telegram dataset) depends on the completeness and representativeness of the collection process, not on fitted parameters or invented entities.

assumptions (4)
  • domain assumption The Telethon API returns complete and accurate message metadata for public channels.
    Used in Section 3.2 for all data collection; no validation against ground truth (e.g., sample spot-checks) is reported.
  • domain assumption The TGStat top-100 lists per subscriber, citation, and reach provide a representative seed set for snowball sampling.
    Introduced in Section 3.1; the convergence shown in Figure 1 is internal and does not demonstrate coverage of the full public channel universe.
  • domain assumption Backward tracing of forwarded messages reconstructs actual propagation paths within the dataset.
    Described in Section 4.3; the paper notes that source channels may be absent, so flows can be truncated, which limits completeness but does not invalidate the approach.
  • domain assumption The Langdetect library, applied to concatenated channel messages, estimates channel language with sufficient accuracy.
    Used in Section 4.1; the cited accuracy (99.77%) comes from news article benchmarks, not Telegram text, which may contain code-switching, slang, or short posts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram." pith.science (2026). https://pith.science/paper/QYUD7242

@misc{pith2026250419536,
  author       = {Pith},
  title        = {Pith review of: TeleScope: A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on Telegram},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QYUD7242}},
  note         = {Machine review of arXiv:2504.19536}
}
read the original abstract

Telegram is a globally popular instant messaging platform known for its strong emphasis on security, privacy, and unique social networking features. It has recently emerged as the host for various cross-domain analysis and research works, such as social media influence, propaganda studies, and extremism. This paper introduces TeleScope, an extensive dataset suite that, to our knowledge, is the largest of its kind. It comprises metadata for about 500K Telegram channels and downloaded message metadata for about 71K public channels, accounting for around 120M crawled messages. We also release channel connections and user interaction data built using Telegram's message-forwarding feature to study multiple use cases, such as information spread and message forwarding patterns. In addition, we provide data enrichments, such as language detection, active message posting periods for each channel, and Telegram entities extracted from messages, that enable online discourse analysis beyond what is possible with the original data alone. The dataset is designed for diverse applications, independent of specific research objectives, and sufficiently versatile to facilitate the replication of social media studies comparable to those conducted on platforms like X (formerly Twitter)

Figures

Figures reproduced from arXiv: 2504.19536 by the authors.

Figure 1
Figure 1. Overlap of discovered channels across different [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Depicts the data collection pipeline. Channels collected using three criteria are merged into a seedlist, enriched via [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Channels discovered via snowball sampling over [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: An example of message forwarding flows in TeleScope. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Example of channel-to-channel graph. nel to each edge [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Channels created per year within TeleScope data. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Active periods reporting the average number of [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: The top 10 hashtags in the messages. these 71K channels. According to our analysis, we notice a distinct pattern in messaging activity, with the lowest ac￾tivity occurring during the early morning hours (Hour 00 to Hour 04). The activity begins to rise sharply from Hou…
Figure 9
Figure 9. Figure 9: The correlation between (a) views and forwards on the left, (b) views and reactions on the right. [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 24 canonical work pages

  1. [1]

    Alvisi, L.; Tardelli, S.; and Tesconi, M. 2024. Unraveling the Italian and English Telegram Conspiracy Spheres through Message Forwarding. arXiv preprint arXiv:2404.18602

  2. [2]

    Baumgartner, J.; Zannettou, S.; Squire, M.; and Blackburn, J. 2020. The pushshift telegram dataset. In Proceedings of the international AAAI conference on web and social media, volume 14, 840--847

  3. [3]

    P.; and Pr\" o llochs, N

    Drolsbach, C. P.; and Pr\" o llochs, N. 2023. Diffusion of Community Fact-Checked Misinformation on Twitter. Proc. ACM Hum.-Comput. Interact., 7(CSCW2)

  4. [4]

    H.; Moore, J

    Ellis, C. H.; Moore, J. B.; Ho, P.; et al. 2024. Social network and linguistic analysis of the\# nutrition discourse on the social network platform X, formerly known as Twitter. Social Network Analysis and Mining, 14(1): 238

  5. [5]

    Elmas, T.; Stephane, S.; and Houssiaux, C. 2023. Measuring and Detecting Virality on Social Media: The Case of Twitter’s Viral Tweets Topic. In Companion Proceedings of the ACM Web Conference 2023, WWW '23 Companion, 314–317. New York, NY, USA. ISBN 9781450394192

  6. [6]

    Fafalios, P.; Iosifidis, V.; Ntoutsi, E.; and Dietze, S. 2018. Tweetskb: A public and large-scale rdf corpus of annotated tweets. In European Semantic Web Conference. Springer

  7. [7]

    Gangopadhyay, S.; Schellhammer, S.; Hafid, S.; et al. 2024. Investigating Characteristics, Biases and Evolution of Fact-Checked Claims on the Web. In Proceedings of the 35th ACM Conference on Hypertext and Social Media, 246--258

  8. [8]

    Ghaffary, S.; and Heilweil, R. 2024. How Trump’s internet built and broadcast the Capitol insurrectionHow Trump’s internet built and broadcast the Capitol insurrection. Accessed: 2024-06-06

Show all 30 references
  1. [9]

    W.; and Durumeric, Z

    Hanley, H. W.; and Durumeric, Z. 2024. Partial mobilization: Tracking multilingual information flows amongst russian media outlets and telegram. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 528--541

  2. [10]

    o hn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A new dataset for bias research from a “dark

    H \"o hn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A new dataset for bias research from a “dark” platform. In Proceedings of the international AAAI Conference on Web and social media, volume 16, 1268--1274

  3. [11]

    Hoseini, M.; de Freitas Melo, P.; Benevenuto, F.; Feldmann, A.; and Zannettou, S. 2024. Characterizing Information Propagation in Fringe Communities on Telegram. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 583--595

  4. [12]

    Höhn, S.; Mauw, S.; and Asher, N. 2022. BelElect: A New Dataset for Bias Research from a “Dark” Platform. Proceedings of the International AAAI Conference on Web and Social Media, 16: 1268--1274

  5. [13]

    J.; and Carley, K

    Kloo, I.; Cruickshank, I. J.; and Carley, K. M. 2024. A Cross-Platform Topic Analysis of the Nazi Narrative on Twitter and Telegram During the 2022 Russian Invasion of Ukraine. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 839--850

  6. [14]

    La Morgia, M.; Mei, A.; Mongardini, A.; and Wu, J. 2023. It’s a Trap! Detection and Analysis of Fake Channels on Telegram. 97--104

  7. [15]

    La Morgia, M.; Mei, A.; and Mongardini, A. M. 2023. Tgdataset: a collection of over one hundred thousand telegram channels. arXiv preprint arXiv:2303.05345

  8. [16]

    H.; Lee, H.; Cesare, N.; Shojaie, A.; and Spiro, E

    McCormick, T. H.; Lee, H.; Cesare, N.; Shojaie, A.; and Spiro, E. S. 2017. Using Twitter for demographic and social science research: tools for data collection and processing. Sociological methods & research, 46(3): 390--421

  9. [17]

    Ovadia, S. 2009. Exploring the potential of Twitter as a research tool. Behavioral & Social Sciences Librarian, 28(4): 202--205

  10. [18]

    Qudar, M. M. A.; and Mago, V. 2020. Tweetbert: a pretrained language representation model for twitter text analysis. arXiv preprint arXiv:2010.11091

  11. [19]

    Schulze, H.; Hohner, J.; et al. 2022. Far-right conspiracy groups on fringe platforms: A longitudinal analysis of radicalization dynamics on Telegram. Convergence: The International Journal of Research into New Media Technologies

  12. [20]

    Solopova, V.; Scheffler, T.; and Popa-Wyatt, M. 2021. A Telegram corpus for hate speech, offensive language, and online harm. Journal of Open Humanities Data, 7

  13. [21]

    Sosa, J.; and Sharoff, S. 2022. Multimodal Pipeline for Collection of Misinformation Data from Telegram. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, 1480--1489

  14. [22]

    Tikhomirova, K.; and Makarov, I. 2021. Community detection based on the nodes role in a network: The telegram platform case. In Analysis of Images, Social Networks and Texts: 9th International Conference, AIST 2020, 294--302

  15. [23]

    R.; Ferreira, C

    Ven \^a ncio, O. R.; Ferreira, C. H.; Almeida, J. M.; et al. 2024. Unraveling User Coordination on Telegram: A Comprehensive Analysis of Political Mobilization during the 2022 Brazilian Presidential Election. In Proceedings of the International AAAI Conference on Web and Socia...

  16. [24]

    Walther, S.; and McCoy, A. 2021. US extremism on Telegram. Perspectives on Terrorism, 15(2): 100--124

  17. [25]

    D.; Dumontier, M.; Aalbersberg, I

    Wilkinson, M. D.; Dumontier, M.; Aalbersberg, I. J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L. B.; Bourne, P. E.; et al. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, 3(1): 1--9

  18. [26]

    H.; Porter, M

    Xiao, Z.; Zhu, J.; Wang, Y.; Zhou, P.; Lam, W. H.; Porter, M. A.; and Sun, Y. 2023. Detecting political biases of named entities and hashtags on Twitter. EPJ Data Science, 20

  19. [27]

    S.; and Speckhard, A

    Yayla, A. S.; and Speckhard, A. 2017. Telegram: The mighty application that ISIS loves. International Center for the Study of Violent Extremism, 9

  20. [28]

    R.; Herbrich, R.; Van Gael, J.; and Stern, D

    Zaman, T. R.; Herbrich, R.; Van Gael, J.; and Stern, D. 2010. Predicting information spreading in twitter. In Workshop on computational social science and the wisdom of crowds, nips, volume 104, 17599--601. Citeseer

  21. [29]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  22. [30]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.