REVIEW 5 major objections 5 minor 19 references
PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany
T0 review · 5 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read German political TikTok posts are deleted at 14–19 times TikTok's reported platform-wide rate, according to a new multimodal dataset.
desk verdict A potentially valuable Saxony TikTok dataset, but the abstract and body disagree on scale and the body's deletion percentage is wrong; the headline numbers are currently unsubstantiated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the four-stage collection pipeline: (I) keyword queries to TikTok's research API with a 96-hour delay to account for indexing lag; (II) direct web scraping to pull video, audio, images, and metadata; (III) a single re-scrape ten days after the collection window to mark posts as deleted; and (IV) human annotation of a sample for politics, Saxony-relevance, intolerance, and hedonic/eudaimonic entertainment. The deletion snapshot (Stage III) is what turns an ordinary scraped corpus into a dataset about platform removal, and the annotation codebook is what makes the deleted content interpretable.
What would settle it
Replicate the collection with daily re-scrapes over at least 30 days instead of a single check at day 10, and cross-check the keyword-API query against a broader crawl of German election TikTok content. If the resulting platform-deletion rate drops below, say, 5%—or close to TikTok's reported 1%—the paper's headline rate of 13% is an artifact of the single-snapshot measurement; if it stays above 10%, the central claim holds.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that political TikTok content is being removed from the platform at rates far exceeding both TikTok's official 1% deletion figure and its claim that 94% of removals happen within 24 hours, and that this deleted content is substantively important rather than noise. Using keyword queries via TikTok's research API plus web scraping to capture full media, the authors assembled a corpus of German election posts and then re-scraped the platform once, ten days after the collection window, to mark which posts had vanished. They report that 17.3% of the Saxony corpus (18,842 of 195,373 posts) was deleted; in the extended federal-election collection, the
Load-bearing premise
The paper's deletion statistics rest on a single availability check ten days after the collection window and on keyword-API queries that may miss posts deleted within the 96-hour indexing delay or outside the keyword set; if those samples are unrepresentative, the computed deletion rates do not describe the true platform behavior.
Editorial extensions
If this is right
- Researchers can study what kinds of political speech TikTok removes (or creators withdraw) by comparing available and deleted posts on content, user, and engagement attributes.
- The multimodal format (video, audio, images, text) allows studying intolerance and humor as they interact across modalities, going beyond text-only or meme-image analyses.
- The high deletion rate of election-related posts suggests that analyses relying only on currently available TikTok content may systematically underrepresent far-right and otherwise problematic speech.
- The public release of post IDs with hydration code enables replication and extension to other elections or platforms.
- The case study's annotated subset provides a benchmark for multimodal intolerance detection, though inter-annotator agreement is below conventional thresholds.
Reading between the lines
- If the deletion-rate finding generalizes beyond Germany, single-snapshot studies of TikTok politics may be biased in a predictable direction: toward mainstream content, since extremist or intolerant posts are removed first.
- The 96-hour lag the authors used suggests TikTok's research API may miss posts that are posted and deleted within that window; the true deletion rate could be even higher than reported.
- A natural next step is to track the same post IDs over a longer horizon (months) to distinguish quick platform removals from slow creator withdrawals, and to check whether deletion rates spike around elections.
- The low inter-annotator agreement on 'intolerance' and 'eudaimonic entertainment' hints that multimodal implicit intolerance may resist reliable human labeling, which is a challenge for training automated detectors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents PoliTok-DE, a multimodal TikTok dataset for the 2024 Saxony state election. The body describes a four-stage pipeline (research API, scraping, re-scraping to identify deletions, annotations), a corpus of 195,373 posts with 18,842 deletions, keyword distributions, and a case study of intolerance and entertainment in a subsample. The abstract supplied with the manuscript, however, advertises a two-election corpus of over 930,000 posts and over 330,000 deleted posts, with a 13.0% platform-deletion rate. These headline statistics do not appear in the full text; the body's own deletion percentage is also arithmetically inconsistent (18,842/195,373 ≈ 9.6%, not 17.3%). The paper also contains annotation-reliability problems (Krippendorff's α ≤ 0.55 for key constructs) and a codebook copy-paste error. If the Saxony-only dataset and pipeline are taken as the actual contribution, the paper is a useful data note, but the current abstract and text cannot both be right.
Significance. A well-documented multimodal dataset of political TikTok posts with deletion snapshots would be a valuable resource for computational social science. The release of post IDs and hydration code, the explicit treatment of API limitations in §2, and the detailed codebook in Appendix B (aside from the error) are strengths. The deletion snapshot is a useful feature, and the keyword prevalence findings (e.g., AfD overrepresentation in deleted content, Figure 2) are interesting. However, the current significance is undermined by the unsupported federal-election statistics and the arithmetic error. The annotation reliability is too low to support the case-study prevalence claims. The dataset's methodological contribution is potentially real, but the published claims must be corrected and substantiated before the paper can be evaluated.
major comments (5)
- [Abstract vs. §2–§3 and §8] The abstract claims a two-election corpus of over 930,000 posts, over 330,000 deletions, 18.7%/39.7% per-election deletion rates, and a 13.0% platform-deletion rate. Neither the federal-election collection nor any of the 930k/330k/13.0% figures appear in the full text; §2 and §3 describe only a Saxony corpus of 195,373 posts, and §8 concludes with the same Saxony-only numbers. The 13.0% figure and the 14–19x comparison are absent from the body. The central advertised contribution is therefore unsupported. The revised manuscript must either document the federal-election collection with full methodology, search queries, and statistics, or withdraw these claims and reframe the paper around the Saxony dataset.
- [§3 Dataset Statistics] The deletion count is arithmetically inconsistent: 18,842/195,373 = 9.6%, not 17.3% as stated in §3 ('17.3% of the dataset'), in §8, and in the narrative of Stage III (§2). If the intended proportion is 17.3%, the deleted count should be about 33,800, not 18,842. This error affects every deletion-related statistic and comparison, including the claim that the deletion rate is 'surprising' relative to TikTok's reported rate. The numerator, denominator, or percentage must be corrected, and all dependent statements revised accordingly.
- [§5.2 Results, Table 1] The inter-annotator agreement for the key constructs is below commonly accepted thresholds: Krippendorff's α = 0.48 for intolerance, 0.55 for hedonic entertainment, and 0.38 for eudaimonic entertainment. The authors acknowledge in the text that these do not meet the recommended α ≥ 0.66 (Krippendorff, 2018). Despite this, the results are presented as point estimates (20.5% intolerance, 62.9% hedonic) and are echoed in the abstract ('about one in five posts conveyed intolerance and a majority conveyed humor'). With these α values, the prevalence estimates are not reliable enough to support the case-study claims; the claims should be explicitly downgraded to exploratory, or the annotation instrument should be refined and the data re-coded before any substantive interpretation.
- [Appendix B.3] The codebook for Intolerance reproduces the Politics section verbatim: the question reads 'Does the video refer to political topics (policy, politics or polity)?' with the same explanation and examples as B.1. If this is the codebook actually used by annotators, the intolerance annotations are not defined by the specified protocol; if it is a clerical error, the appendix must be corrected before the dataset is used or cited. As written, this undermines the reproducibility of the annotation procedure and the credibility of the intolerance results.
- [§2 Stage I and §6 Limitations] The deletion rate is estimated from a single re-scrape performed 10 days after the end of the collection window (§2 Stage III, §6). Additionally, daily API queries are run 96 hours after publication (§2 Stage I), and the paper acknowledges that content deleted before the API query 'never makes it to the research API.' Consequently, posts deleted within those 96 hours are systematically missing from the corpus, making the deletion count among collected posts a lower bound rather than an unbiased rate. The manuscript does not quantify this censoring or its impact on the headline deletion percentages. A sensitivity analysis or at least an explicit bounding statement is needed.
minor comments (5)
- [General] The headline abstract and the abstract in the full text describe different corpora (930k two-election corpus vs. 195k Saxony-only corpus). Even after addressing the major data mismatch, the two versions must be harmonized.
- [Figure 2] The two panels ('Full' and 'Deleted') would benefit from explicit axis labels and a caption clarifying that percentages are computed within each subset. The current presentation is visually dense and hard to read.
- [§5.2] The text reports that 'about 13% of the posts were annotated as conveying both hedonic and eudaimonic entertainment,' but Table 1 reports only marginal distributions. The joint distribution should be reported to support this statement.
- [§2 Stage I] The phrase 'we collected the data for each day 96 hours (4 days) later' is ambiguous. Specify whether each day's query was run exactly 96 hours after publication date, after the end of the day, or at some other offset.
- [References] The citation 'TikTok [2025]' in §3 refers to an online transparency report; the reference list gives a URL but the citation format is inconsistent with the author-date style used elsewhere. Minor formatting issue.
Circularity Check
No circularity: the corpus, deletion rates, and annotation results are all observational measurements, not predictions derived from fitted inputs or self-citations.
full rationale
The paper's derivation chain is a measurement pipeline, not a derivation: posts are collected via the TikTok research API and scraping (§2, Stages I–II), deletion status is determined by a re-scrape (§2, Stage III), and the 18,842/195,373 deletion count is a direct count, not a fitted parameter (§3). The comparison with TikTok's platform-wide reported rate is an external benchmark, not an input used to construct the dataset. The case-study annotations are human labels (Table 1) with stated inter-annotator agreement; low α for intolerance/entertainment is a reliability limitation, not a circular step. The paper explicitly acknowledges the coarse single re-scrape (§6) and keyword-matching limits (§6), which are measurement caveats. The input's top abstract contains federal-election figures (930k posts, 330k deletions) that do not appear in the body, whose abstract and §3 report 195,373 posts and 18,842 deletions; moreover 18,842/195,373 = 9.6%, not the stated 17.3%. These are internal-consistency/arithmetic problems, not circularity: no claim is defined in terms of itself, no fitted input is relabeled as a prediction, and no load-bearing argument rests on a self-citation. Hence circularity score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The TikTok research API returns a representative and complete set of posts matching the election keywords.
- domain assumption A single re-scrape 10 days after the collection window correctly determines which posts are deleted.
- ad hoc to paper Annotation agreement with Krippendorff's alpha as low as 0.38 is sufficient to support the reported prevalence of intolerance and entertainment.
Cite this review
Pith. "Pith review of PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany." pith.science (2026). https://pith.science/paper/XD4UP2A2
@misc{pith2026250915860,
author = {Pith},
title = {Pith review of: PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany},
year = {2026},
howpublished = {\url{https://pith.science/paper/XD4UP2A2}},
note = {Machine review of arXiv:2509.15860}
}
read the original abstract
We present PoliTok-DE, a large-scale multimodal dataset (video, audio, images, text) of TikTok posts from two German elections: the 2024 Saxony state election and the 2025 German federal election. The corpus contains over 930,000 posts, of which over 330,000 were later deleted from the platform (18.7% of Saxony posts, 39.7% of federal posts). In the federal-election collection, about two thirds of the deletions were creator withdrawals, and the platform-deletion rate we computed was 13.0% of all posts, more than an order of magnitude (14-19x) above the platform-wide rate TikTok reported. Posts were identified via the TikTok research API and complemented with web scraping to retrieve full multimodal media and metadata. PoliTok-DE supports social science research across substantive and methodological agendas: substantive work on intolerance and political communication, and methodological work on platform policies around deleted content and qualitative-quantitative multimodal research. To illustrate, we report a case study on intolerance and entertainment in an annotated subset of deleted posts: about one in five posts conveyed intolerance and a majority conveyed humor.
Figures
Reference graph
Works this paper leans on
-
[1]
Shah, Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim
Aashish Bhandari, Siddhant B. Shah, Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim. Crisishatemm: Multimodal analysis of directed and undirected hate speech in text-embedded images from russia-ukraine conflict. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1994--2003, 2023. doi:10.1109/CVPRW59228.2023.00193
arXiv 2023
-
[2]
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. Latent hatred: A benchmark for understanding implicit hate speech. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processin...
-
[3]
Tiktok's research api: Problems without explanations, 2025
Carlos Entrena-Serrano, Martin Degeling, Salvatore Romano, and Raziye Buse Çetin. Tiktok's research api: Problems without explanations, 2025. URL https://arxiv.org/abs/2506.09746
arXiv 2025
-
[4]
S em E val-2022 task 5: Multimedia automatic misogyny identification
Elisabetta Fersini, Francesca Gasparini, Giulia Rizzi, Aurora Saibene, Berta Chulvi, Paolo Rosso, Alyssa Lees, and Jeffrey Sorensen. S em E val-2022 task 5: Multimedia automatic misogyny identification. In Guy Emerson, Natalie Schluter, Gabriel Stanovsky, Ritesh Kumar, Alexis Palmer, Nathan Schneider, Siddharth Singh, and Shyam Ratan, editors, Proceedings...
2022
-
[5]
Social-media-partei afd? Digitale Landtagswahlk \"a mpfe im Vergleich
Maik Fielitz, Harald Sick, Michael Schmidt, and Christian Donner. Social-media-partei afd? Digitale Landtagswahlk \"a mpfe im Vergleich. Hg. v. Otto Brenner Stiftung. Frankfurt aM (OBS-Arbeitspapier, 73) , 2024
2024
-
[6]
Rechtsextremismus im social web
Lara Franke and Daniel Hajok. Rechtsextremismus im social web. mit neuen propagandastrategien nun auch bei tiktok? JMS Jugend Medien Schutz-Report, 45 0 (3): 0 2--4, 2022
2022
-
[7]
Landtagswahlen 2024
Freistaat Sachsen . Landtagswahlen 2024 . https://www.wahlen.sachsen.de/landtagswahlen-2024.html, 2024. Accessed: August 21, 2025
2024
-
[8]
Analyzing radical visuals at scale: How far-right groups mobilize on tiktok
Julian Hohner, Azade Kakavand, and Sophia Rothut. Analyzing radical visuals at scale: How far-right groups mobilize on tiktok. Journal of Digital Social Research, 6 0 (1): 0 10--30, 2024
2024
Show all 19 references
-
[9]
Content analysis: An introduction to its methodology
Klaus Krippendorff. Content analysis: An introduction to its methodology. Sage publications, 2018
2018
-
[10]
News on tiktok: An annotated dataset of tiktok videos from german-speaking news outlets in 2023
Anna-Theresa Mayer, Lion Wedel, Jan Batzner, Jonathan Hendrickx, Emma Bremer, Alexander Iwan, Volker Stocker, and Jakob Ohme. News on tiktok: An annotated dataset of tiktok videos from german-speaking news outlets in 2023. In Proceedings of the International AAAI Conference on...
2023
-
[11]
Afd auf tiktok: So abgehängt sind die anderen parteien, 2024
Nils Metzger. Afd auf tiktok: So abgehängt sind die anderen parteien, 2024. URL https://www.zdfheute.de/politik/deutschland/afd-tiktok-erfolg-strategie-jugendliche-100.html
2024
-
[12]
Entertainment as pleasurable and meaningful: Identifying hedonic and eudaimonic motivations for entertainment consumption
Mary Beth Oliver and Arthur A Raney. Entertainment as pleasurable and meaningful: Identifying hedonic and eudaimonic motivations for entertainment consumption. Journal of Communication, 61 0 (5): 0 984--1004, 2011
2011
-
[13]
Beyond the margin of error: a systematic and replicable audit of the tiktok research api
George DH Pearson, Nathan A Silver, Jessica Y Robinson, Mona Azadi, Barbara A Schillo, and Jennifer M Kreslake. Beyond the margin of error: a systematic and replicable audit of the tiktok research api. Information, Communication & Society, 28 0 (3): 0 452--470, 2025
2025
-
[14]
Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty
Shraman Pramanick, Dimitar Dimitrov, Rituparna Mukherjee, Shivam Sharma, Md. Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty. Detecting harmful memes and their targets. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Findings of the Association for Comp...
2021 doi
-
[15]
Beyond incivility: Understanding patterns of uncivil and intolerant discourse in online political talk
Patr \' cia Rossini. Beyond incivility: Understanding patterns of uncivil and intolerant discourse in online political talk. Communication Research, 49 0 (3): 0 399--425, 2022
2022
-
[16]
Just another hour on tiktok: Reverse-engineering unique identifiers to obtain a complete slice of tiktok, 2025
Benjamin Steel, Miriam Schirmer, Derek Ruths, and Juergen Pfeffer. Just another hour on tiktok: Reverse-engineering unique identifiers to obtain a complete slice of tiktok, 2025. URL https://arxiv.org/abs/2504.13279
2025 arXiv
-
[17]
Multimodal meme dataset ( M ulti OFF ) for identifying offensive content in image and text
Shardul Suryawanshi, Bharathi Raja Chakravarthi, Mihael Arcan, and Paul Buitelaar. Multimodal meme dataset ( M ulti OFF ) for identifying offensive content in image and text. In Ritesh Kumar, Atul Kr. Ojha, Bornini Lahiri, Marcos Zampieri, Shervin Malmasi, Vanessa Murdock, and...
2020
-
[18]
Community guidelines enforcement - 2025-1
TikTok. Community guidelines enforcement - 2025-1. https://www.tiktok.com/transparency/en/community-guidelines-enforcement-2025-1, 2025. Accessed: August 21, 2025
2025
-
[19]
Tiktok regional election monitor 2024: Methods report
Johannes Wolfgram, Jasper Tjaden, Aaron Philipp, Sarah Wei mann, Ulrich Kohler, Licia Bobzien, and Roland Verwiebe. Tiktok regional election monitor 2024: Methods report. Release version, 1, 2024
2024
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.