Pith. sign in

REVIEW 3 major objections 5 minor 34 references

Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Age splits toxic German comments, 33k-comment dataset shows

desk verdict A useful German multi-platform toxicity dataset whose headline age findings rest on an acknowledged ecological inference; the resource is worth having, but the age patterns should be treated as exploratory. read the letter →

arxiv 2508.21084 v1 pith:D5BLQ2QN submitted 2025-08-26 cs.CL cs.CY

classification cs.CLcs.CY
keywords toxicspeechagedemographicsGermandatasetLLMannotationcontentmoderationsocialmediaplatformsdisinformationpublicbroadcasting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds a German-language dataset of 33,048 social media comments (3,024 human-annotated, 30,024 annotated by a fine-tuned LLM) from the public broadcaster funk's Instagram, TikTok, and YouTube channels, pairing toxicity labels with platform-provided audience-age estimates. It claims that toxicity patterns differ by age: users under 30 use more expressive, sarcastic, emoji-heavy language; the 31-35 group produces the most insults and disinformation; users over 35 engage more in devaluation and disinformation. It also reports platform differences, with Instagram showing the most toxic content. The dataset is presented as the first large-scale German resource combining toxic speech with any demographic marker, intended to support age-aware moderation and linguistic analysis.

What carries the argument

The dataset itself is the load-bearing artifact: two corpora sharing one 18-label annotation scheme with Target and Type classes, one human-annotated (3,024 comments, three annotators, majority vote, Fleiss kappa 0.64) and one LLM-annotated (30,024 comments, fine-tuned GPT-4o-mini). The age mechanism is channel-level age inference: platforms provide the age distribution of a channel's followers (Instagram) or viewers (YouTube), funk converts it into an average age per account, and every comment from that channel inherits that average; TikTok supplies no age data and is excluded from age analysis. Comments were pre-filtered with a toxic-keyword list, so the sample is enriched for problematic

What would settle it

Collect a sample of actual commenter ages, e.g., via an opt-in survey linked to anonymized comment IDs, and compare them to the channel-average age assignments; if the assignments are wrong for a substantial share of comments, or if a within-channel regression using actual commenter ages does not reproduce the age-label patterns, the age findings collapse. Comparing channels where audiences and commenters are known to diverge would provide a direct test.

Watch

Extended reading notes

Core claim

The central claim is that online toxicity is measurably age-structured in German comment sections: a chi-square test on the human-annotated data (chi-squared = 61.66, df = 36, p = 0.0049) supports a relationship between age group and label distribution, and the pattern repeats in the LLM-extrapolated set. Younger users (0-30) show the highest share of non-toxic comments and prefer emojis and sarcastic terms in their insults; the 31-35 group leads in insults, disinformation, and discrimination; the 35+ group leads in devaluation, disinformation, and religion- or ethnicity-based hate. Platform-level analysis finds Instagram highest in toxic content (22.53%), followed by YouTube (13.03%) and Ti

Load-bearing premise

The entire age analysis assumes that people who comment on a channel have the same age distribution as the channel's overall audience, because each comment is assigned the channel's average follower or viewer age rather than the commenter's actual age.

Editorial extensions

If this is right

  • If the age patterns hold, moderation systems could weight categories by age: disinformation detection would be most relevant for the 31-35 and 35+ groups, while emoji and sarcasm detection would matter more for users under 30.
  • The platform differences imply that moderation strategies may need to be platform-specific, with Instagram requiring the most attention for insults, devaluation, and disinformation.
  • The fine-tuned LLM pipeline shows a practical path to scaling human toxicity labels to tens of thousands of comments for about four dollars, enabling broader demographic toxicity monitoring.
  • The 31-35 group appears as a distinct toxicity peak rather than a monotonic increase with age, pointing to a middle-age pattern worth investigating in other languages and platforms.
  • The dataset provides a benchmark for German toxicity annotation with demographic context, supporting training and evaluation of age-aware models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The channel-averaged age method likely compresses within-channel age variation; if commenters skew younger or older than followers/viewers, the reported age effects could be artifacts. A small opt-in survey of actual commenter ages compared with the channel-average assignments would test this.
  • Because comments were pre-filtered by toxic keywords, the reported proportions (e.g., 16.7% problematic) are conditional on that filter and cannot be read as base rates of toxicity across all comments on these channels.
  • The 31-35 peak in disinformation and devaluation could partly reflect topic effects, since different age groups may comment on different videos; the paper acknowledges this confound, but future within-topic analysis would be needed to separate age from subject matter.
  • The platform-level discrepancy between the human and LLM annotations (where YouTube's toxicity ranking shifts in the larger set) suggests that LLM-annotated demographic patterns, while broadly reproducing age trends, should be treated as exploratory for platform comparisons.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a new German-language toxic-speech dataset collected from funk's Instagram, TikTok, and YouTube channels, comprising 3,024 human-annotated comments and 30,024 LLM-annotated comments. It compares four LLMs for annotation, fine-tunes GPT-4o-mini, and reports analyses of label distributions by age group, platform, and keywords. The authors claim age-based differences, with younger users using more expressive language and older users more disinformation/devaluation. However, age is not observed per comment; it is inferred from channel-level audience age distributions, which is acknowledged in Section 10. The dataset and annotation materials are intended for public release with restricted access.

Significance. The resource addresses a real gap: no existing German multi-platform toxicity dataset includes demographic estimates, and the collaboration with a public broadcaster provides a unique avenue for platform-provided audience demographics. The paper's transparent data statement, detailed annotation protocol, evaluation of multiple LLMs, and release of scripts/guidelines are strengths. The LLM annotation comparison is useful, and the keyword-filtered sampling is documented. If the age-related claims are substantially reframed or supported by appropriate statistical modeling, the dataset could enable meaningful demographic studies of online toxicity. At present, however, the headline age finding is not adequately supported.

major comments (3)
  1. [§7.1 and §3] The central age finding rests on an ecological inference. Each comment is assigned the average audience age of its channel (141 accounts; §3), so the 3,024 comments are not independent age observations. The chi-square test (χ²=61.66, df=36, p=0.0049) ignores clustering and also treats multi-label counts as independent per-label entries; both violate test assumptions. The effective sample for age is at account level, not comment level, and channel age profile is confounded with topic/moderation (§10). The abstract and §9 state age-based differences as findings, but the current analysis cannot support that claim. A channel-level or mixed-effects analysis is required.
  2. [§8] The LLM-extrapolated set reproduces some age-group orderings, but this cannot validate the age finding because the same channel-average age assignment is used. The failure to reproduce platform-level trends (§H) further indicates that LLM annotations are not reliable for cross-demographic claims. Please discuss what conclusions can legitimately be drawn from the LLM set and do not present replication as independent confirmation.
  3. [§7.3 and Abstract] The claim that younger users favor expressive language is supported only by emoji counts among the ten most frequent words for two labels (15 vs. 10 vs. 7). No hypothesis test or effect size is reported, and the analysis is based on tiny subsets. This descriptive observation does not support a general age-group finding; the abstract should be reworded or the analysis strengthened.
minor comments (5)
  1. [§7.1 / Tables 13–17] Age group definitions use '31–35' and '35+', overlapping at 35; '0–30' includes an implausible age of 0. Clarify boundaries and justify the choice of bins.
  2. [Terminology] Appendix tables use 'no_hate' while §7 uses 'non-toxic'; main text inconsistently capitalizes 'Funk' vs. 'funk'. Standardize.
  3. [Figures and tables] Figure 1 appears as caption only, and Table 15 contains rendering artifacts (e.g., 'u200d', missing emoji glyphs). Ensure all figures are embedded and Unicode is handled.
  4. [References / placeholders] Several references to 'Annonymous' and 'Anonymous' need resolving; the word list and scripts should be linked with accessible identifiers.
  5. [Dataset access] Since the dataset is access-restricted, include a sample of annotated examples in the appendix or a data card to allow review of label quality.

Circularity Check

0 steps flagged · score 1.0 of 10

No circularity: the age findings rest on an acknowledged ecological proxy, not on a fitted/self-referential construction; self-citations are incidental.

full rationale

The paper's central empirical claim—age-group differences in toxic labels—is a statistical association computed from human/LLM labels and channel-level age proxies. It is not derived from its inputs by construction: the age proxy could in principle yield null or opposite patterns, and the toxicity labels come from human annotation and from a fine-tuned LLM evaluated on a held-out set. The self-citation to Fillies et al. (2025) is used only as taxonomic inspiration for the annotation scheme, not as a load-bearing proof or prediction; Fillies et al. (2023) is cited as related work. Section 10 candidly states that the age inference assumes 'users engaging in the comment section of a channel fit the overall audience age profile' and that confounding variables may drive the observed patterns. That is a validity limitation (ecological inference, clustered data, uncontrolled channel-level confounds), not circularity. No equation-level reduction, no fitted parameter renamed as a prediction, and no uniqueness claim imported from self-citation occur. Therefore the analysis is self-contained; the minor self-citations do not raise the circularity score beyond 1.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central age findings depend on two unvalidated domain assumptions: that channel audience demographics equal commenter demographics, and that uniform distribution within age buckets is accurate. The keyword filter is an independent selection mechanism that shapes the dataset. No new theoretical entities are introduced.

free parameters (3)
  • Age binning for analysis = 0-30, 31-35, 35+
    The age analysis groups users into three arbitrary bins chosen based on funk's target audience. The choice is ad hoc and affects the reported age-difference patterns.
  • Keyword filter list = Not fully enumerated in paper
    Comments were pre-selected using a predefined toxic keyword list, which directly determines the 16.7% problematic rate and biases prevalence estimates.
  • Uniform age distribution assumption = Uniform within platform age buckets
    The paper assumes followers/views are evenly distributed within each age group to compute an average age per account. This is a modeling choice that injects uncertainty into every age assignment.
assumptions (4)
  • domain assumption Commenters on a channel match the channel's overall audience age profile
    Invoked in Section 3 and explicitly acknowledged in Section 10; without this, age labels for comments are invalid.
  • domain assumption Platform-provided age groups are based on user-reported birth dates and are accurate enough
    Mentioned in Section 3: 'the age groups provided by the platform rely on user-reported birth dates, which may be inaccurate'.
  • domain assumption The annotation schema faithfully captures toxicity and is appropriate for all platforms
    The schema is based on funk's content policies and may not generalize; the paper does not validate it externally.
  • domain assumption LLM annotations can reliably extrapolate human annotations to unseen data
    The fine-tuned GPT-4o-mini is used to label 30,024 comments despite the paper's own evidence of platform-level mismatches and acknowledged LLM biases (Section 10).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting." pith.science (2026). https://pith.science/paper/D5BLQ2QN

@misc{pith2026250821084,
  author       = {Pith},
  title        = {Pith review of: Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D5BLQ2QN}},
  note         = {Machine review of arXiv:2508.21084}
}
read the original abstract

A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public service content network, this research introduces the first large-scale German dataset annotated for toxicity and enriched with platform-provided age estimates. The dataset includes 3,024 human-annotated and 30,024 LLM-annotated anonymized comments from Instagram, TikTok, and YouTube. To ensure relevance, comments were consolidated using predefined toxic keywords, resulting in 16.7\% labeled as problematic. The annotation pipeline combined human expertise with state-of-the-art language models, identifying key categories such as insults, disinformation, and criticism of broadcasting fees. The dataset reveals age-based differences in toxic speech patterns, with younger users favoring expressive language and older users more often engaging in disinformation and devaluation. This resource provides new opportunities for studying linguistic variation across demographics and supports the development of more equitable and age-aware content moderation systems.

Figures

Figures reproduced from arXiv: 2508.21084 by the authors.

Figure 1
Figure 1. A figure displaying the age distribution in the [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Heatmaps labels per platform 13 [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Heatmaps labels per platform Annotator Pair Agreement (%) Annotator 1 vs Annotator 2 87.7 Annotator 1 vs Annotator 3 89.6 Annotator 2 vs Annotator 3 88.2 [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 19 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Riehle, and Heike Trautmann

    Dennis Assenmacher, Marco Niemann, Kilian M \"u ller, Moritz Vinzent Seiler, Dennis M. Riehle, and Heike Trautmann. 2021. https://api.semanticscholar.org/CorpusID:244906720 Rp-mod&rp-crowd: Moderator- and crowd-annotated german news comment datasets . In NeurIPS Datasets and Benchmarks

  4. [4]

    Uwe Bretschneider and Ralf Peters. 2017. Detecting offensive statements towards foreigners in social media. In Hawaii International Conference on System Sciences

  5. [5]

    Davide Chicco and Giuseppe Jurman. 2020. The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation. BMC genomics, 21:1--13

  6. [6]

    Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini. 2019. https://doi.org/10.18653/v1/P19-1271 CONAN - CO unter NA rratives through nichesourcing: a multilingual dataset of responses to fight online hate speech . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2819--2829, Florence,...

  7. [7]

    Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1):37--46

  8. [8]

    Amit Das, Mostafa Rahgouy, Dongji Feng, Zheng Zhang, Tathagata Bhattacharya, Nilanjana Raychawdhary, Fatemeh Jamshidi, Vinija Jain, Aman Chadha, Mary J Sandage, et al. 2024 a . Offensivelang: A community based implicit offensive language dataset. IEEE Access

Show all 34 references
  1. [9]

    Amit Das, Zheng Zhang, Najib Hasan, Souvika Sarkar, Fatemeh Jamshidi, Tathagata Bhattacharya, Mostafa Rahgouy, Nilanjana Raychawdhary, Dongji Feng, Vinija Jain, Aman Chadha, Mary Sandage, Lauramarie Pope, Gerry Dozier, and Cheryl Seals. 2024 b . https://arxiv.org/abs/2406.1110...

  2. [10]

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. https://arxiv.org/abs/1703.04009 Automated hate speech detection and the problem of offensive language . Preprint, arXiv:1703.04009

  3. [12]

    Ona de Gibert, Naiara Perez, Aitor Garc \'i a-Pablos, and Montse Cuadros. 2018 b . https://doi.org/10.18653/v1/W18-5102 Hate speech dataset from a white supremacy forum . In Proceedings of the 2nd Workshop on Abusive Language Online ( ALW 2) , pages 11--20, Brussels, Belgium. ...

  4. [13]

    Christoph Demus, Jonas Pitz, Mina Sch \"u tz, Nadine Probol, Melanie Siegel, and Dirk Labudde. 2022. https://doi.org/10.18653/v1/2022.woah-1.14 Detox: A comprehensive dataset for G erman offensive language and conversation analysis . In Proceedings of the Sixth Workshop on Onl...

  5. [14]

    Penelope Eckert. 2012. Three waves of variation study: The emergence of meaning in the study of sociolinguistic variation. Annual review of Anthropology, 41(1):87--100

  6. [15]

    Mai ElSherief, Vivek Kulkarni, Dana Nguyen, William Yang Wang, and Elizabeth Belding. 2018. Hate lingo: A target-based linguistic analysis of hate speech in social media. In Proceedings of the international AAAI conference on web and social media, volume 12

  7. [16]

    Jan Fillies, Silvio Peikert, and Adrian Paschke. 2023. Hateful messages: A conversational data set of hate speech produced by adolescents on discord. In International Data Science Conference, pages 37--44. Springer

  8. [17]

    Jan Fillies, Esther Theisen, Michael Hoffmann, Robert Jung, Elena Jung, Nele Fischer, and Adrian Paschke. 2025. A novel german tiktok hate speech dataset: far-right comments against politicians, women, and others. Discover Data, 3(1):4

  9. [18]

    Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378

  10. [19]

    Antigoni Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018. Large scale crowdsourcing and characterization of twitter abusive behavior. In Proceedings of th...

  11. [20]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daum\' e III, and Kate Crawford. 2021. https://doi.org/10.1145/3458723 Datasheets for datasets . Commun. ACM, 64(12):86–92

  12. [21]

    Janis Goldzycher, Paul Röttger, and Gerold Schneider. 2024. https://arxiv.org/abs/2403.19559 Improving adversarial data collection by supporting annotators: Lessons from gahd, a german hate speech dataset . Preprint, arXiv:2403.19559

  13. [22]

    Saba Gulzar. 2023. The role of media in shaping public opinion and social discourse. Contemporary Journal of Social Science Review, 1(1):30--40

  14. [23]

    Homa Hosseinmardi, Sabrina Arredondo Mattson, Rahat Ibn Rafiq, Richard Han, Qin Lv, and Shivakant Mishra. 2015. Analyzing labeled cyberbullying incidents on the instagram social network. In Social Informatics, pages 49--66, Cham. Springer International Publishing

  15. [24]

    Max-Emanuel Keller, Maximilian Auch, Alexander D \"o schl, Fabian Vlk, Julian Quernheim, Mike Hartmann, Peter Mandl, Alexander Kaul, and Markus Franz. 2025. Hocon34k: A corpus of hate speech in online comments from german newspapers. In International Conference on Information ...

  16. [25]

    Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Mandlia, and Aditya Patel. 2019. https://doi.org/10.1145/3368567.3368584 Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages . In...

  17. [26]

    Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. https://doi.org/10.1609/aaai.v35i17.17745 Hatexplain: A benchmark dataset for explainable hate speech detection . Proceedings of the AAAI Conference on Artificial Intelligen...

  18. [27]

    Tharindu Ranasinghe and Marcos Zampieri. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.470 Multilingual offensive language identification with cross-lingual embeddings . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), page...

  19. [28]

    Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.407 Probing LLM s for hate speech detection: strengths and vulnerabilities . In Findings of the Association for Computational Linguistics: EMNLP 2023, ...

  20. [29]

    Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Viviana Patti, and Marco Stranisci. 2018. https://aclanthology.org/L18-1443/ An I talian T witter corpus of hate speech against immigrants . In Proceedings of the Eleventh International Conference on Language Resources and Ev...

  21. [30]

    Andrew Schwartz, Johannes C

    H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Stephanie M. Ramones, Megha Agrawal, Achal Shah, Michal Kosinski, David Stillwell, Martin E.P. Seligman, and Lyle H. Ungar. 2013. https://doi.org/10.1371/journal.pone.0073791 Personality, gender,...

  22. [31]

    Rachele Sprugnoli, Stefano Menini, Sara Tonelli, Filippo Oncini, and Enrico Piras. 2018. https://doi.org/10.18653/v1/W18-5107 Creating a W hats A pp dataset to study pre-teen cyberbullying . In Proceedings of the 2nd Workshop on Abusive Language Online ( ALW 2) , pages 51--59,...

  23. [32]

    Bertie Vidgen and Leon Derczynski. 2021. https://doi.org/10.1371/journal.pone.0243300 Directions in abusive language training data, a systematic review: Garbage in, garbage out . PLOS ONE, 15(12):1--32

  24. [33]

    Zeerak Waseem and Dirk Hovy. 2016. https://doi.org/10.18653/v1/N16-2013 Hateful symbols or hateful people? predictive features for hate speech detection on T witter . In Proceedings of the NAACL Student Research Workshop , pages 88--93, San Diego, California. Association for C...

  25. [34]

    Xin Wen, Yulan Wang, Kai Wang, and Ran Sui. 2022. https://doi.org/10.1109/BigDataSecurityHPSCIDS54978.2022.00018 A russian hate speech corpus for cybersecurity applications . In 2022 IEEE 8th Intl Conference on Big Data Security on Cloud (BigDataSecurity), IEEE Intl Conference...

  26. [35]

    Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer. 2019. https://nbn-resolving.org/urn:nbn:de:bsz:mh39-84935 Overview of the germeval 2018 shared task on the identification of offensive language . Proceedings of GermEval 2018, 14th Conference on Natural Language Processin...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.