{"id":"c8b412b3-6b5f-4c07-ac0d-10c46b58ef38","arxiv_id":"2508.21084","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A new German toxic-comment dataset with channel-level age estimates reveals age differences in toxic speech, but age is inferred from channel audience profiles, not individual users.","lead":"This paper presents a new German-language dataset of social media comments with toxicity labels and age estimates supplied by the platforms. The authors report age-related patterns in toxic language and compare four AI models for annotation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Age findings in §7.1 rely on assigning each comment its channel's average audience age; with only 141 channel-level age values, the chi-square test overstates evidence and channel-level confounding is uncontrolled.","rationale":"The reader's weakest assumption—that channel audience age equals commenter age—is exactly the load-bearing concern. The paper's own limitation section admits this, but the abstract and conclusions still assert age-based differences. My proposed test would determine whether the reported age patterns survive when the analysis respects the actual level of variation (account level) and controls for channel confounding. This concern does not invalidate the dataset as a resource, but it does mean the age-related findings should be presented as exploratory and hypothesis-generating, not as established demographic differences. Since the reader already recommended a conditional acceptance with a reframing of the age conclusions, my assessment leaves the verdict unchanged. I agree with the reader's reading and see no additional load-bearing issue beyond this one; other problems (LLM platform mismatch, restricted availability) are secondary and do not change the conditional verdict.","tokens_in":15714,"tokens_out":3128,"duration_ms":37711,"concrete_test":"Re-run the §7.1 age analysis at the account level. For each of the 141 accounts, compute the proportion of toxic comments (or the proportion of each label) and the account's average age. Then perform a weighted linear regression of toxic proportion on three age groups (0–30, 31–35, 35+), weighting by the number of comments per account, and also a chi-square test on the account-level counts (aggregating comments within accounts). If the p-value exceeds 0.05 or the effect sizes shrink by more than half, the age-based finding is an artifact of treating comments as independent samples. As a second check, fit a mixed-effects logistic regression predicting toxic/non-toxic with age group as a fixed effect and account as a random intercept; if the age coefficient is no longer significant, the effect is confounded with channel identity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that the dataset reveals age-based differences in toxic speech — is built on an ecological inference. As stated in §3 and §10, the age assigned to each comment is the average age of the channel's audience (followers/viewers), not the age of the commenter. This produces at most 141 distinct age values (one per account) for the 3,024 human-annotated comments, yet §7.1 treats each comment as an independent observation in a chi-square test (χ²=61.66, df=36, p=0.0049). That test ignores the clustering of comments within accounts, so the effective sample size for age is far smaller than 3,024, and the p-value is likely anticonservative. Worse, the channel-level average age is a proxy for the channel itself: channels differ in topic, content style, moderation policy, and audience composition. The age-group patterns reported (e.g., 31–35 group having more disinformation) may therefore reflect differences between channels rather than between individual commenters of different ages. The paper explicitly acknowledges this limitation in §10, but the abstract and §9 nevertheless state the age-based differences as a finding. The LLM-extrapolated dataset (§8) does reproduce the age-group ordering for some labels, but since the same channel-average age assignment is used there, that consistency does not validate the ecological assumption. The lack of TikTok in the age analysis further narrows the claim. This is the load-bearing concern because if the age effect does not survive account-level analysis, the paper's headline contribution—demographic mapping of toxicity—loses its empirical support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a new German-language toxic-speech dataset collected from funk's Instagram, TikTok, and YouTube channels, comprising 3,024 human-annotated comments and 30,024 LLM-annotated comments. It compares four LLMs for annotation, fine-tunes GPT-4o-mini, and reports analyses of label distributions by age group, platform, and keywords. The authors claim age-based differences, with younger users using more expressive language and older users more disinformation/devaluation. However, age is not observed per comment; it is inferred from channel-level audience age distributions, which is acknowledged in Section 10. The dataset and annotation materials are intended for public release with restricted access.","tokens_in":16030,"tokens_out":5249,"duration_ms":55726,"significance":"The resource addresses a real gap: no existing German multi-platform toxicity dataset includes demographic estimates, and the collaboration with a public broadcaster provides a unique avenue for platform-provided audience demographics. The paper's transparent data statement, detailed annotation protocol, evaluation of multiple LLMs, and release of scripts/guidelines are strengths. The LLM annotation comparison is useful, and the keyword-filtered sampling is documented. If the age-related claims are substantially reframed or supported by appropriate statistical modeling, the dataset could enable meaningful demographic studies of online toxicity. At present, however, the headline age finding is not adequately supported.","major_comments":[{"comment":"The central age finding rests on an ecological inference. Each comment is assigned the average audience age of its channel (141 accounts; §3), so the 3,024 comments are not independent age observations. The chi-square test (χ²=61.66, df=36, p=0.0049) ignores clustering and also treats multi-label counts as independent per-label entries; both violate test assumptions. The effective sample for age is at account level, not comment level, and channel age profile is confounded with topic/moderation (§10). The abstract and §9 state age-based differences as findings, but the current analysis cannot support that claim. A channel-level or mixed-effects analysis is required.","section":"§7.1 and §3"},{"comment":"The LLM-extrapolated set reproduces some age-group orderings, but this cannot validate the age finding because the same channel-average age assignment is used. The failure to reproduce platform-level trends (§H) further indicates that LLM annotations are not reliable for cross-demographic claims. Please discuss what conclusions can legitimately be drawn from the LLM set and do not present replication as independent confirmation.","section":"§8"},{"comment":"The claim that younger users favor expressive language is supported only by emoji counts among the ten most frequent words for two labels (15 vs. 10 vs. 7). No hypothesis test or effect size is reported, and the analysis is based on tiny subsets. This descriptive observation does not support a general age-group finding; the abstract should be reworded or the analysis strengthened.","section":"§7.3 and Abstract"}],"minor_comments":[{"comment":"Age group definitions use '31–35' and '35+', overlapping at 35; '0–30' includes an implausible age of 0. Clarify boundaries and justify the choice of bins.","section":"§7.1 / Tables 13–17"},{"comment":"Appendix tables use 'no_hate' while §7 uses 'non-toxic'; main text inconsistently capitalizes 'Funk' vs. 'funk'. Standardize.","section":"Terminology"},{"comment":"Figure 1 appears as caption only, and Table 15 contains rendering artifacts (e.g., 'u200d', missing emoji glyphs). Ensure all figures are embedded and Unicode is handled.","section":"Figures and tables"},{"comment":"Several references to 'Annonymous' and 'Anonymous' need resolving; the word list and scripts should be linked with accessible identifiers.","section":"References / placeholders"},{"comment":"Since the dataset is access-restricted, include a sample of annotated examples in the appendix or a data card to allow review of label quality.","section":"Dataset access"}],"recommendation":"major_revision","confidential_remarks":"The paper would be stronger if the age analysis were presented as an exploratory, hypothesis-generating analysis at the channel-audience level, not as a confirmed finding. The restricted dataset access makes it hard to verify the exact age assignments and labels; please ensure the review version contains enough examples for evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a genuine resource: the first German multi-platform toxic-comment dataset that includes platform-provided age estimates. That combination is new, and the authors document the construction carefully—sampling, anonymization, annotation scheme, inter-annotator agreement, and a comparison of four LLMs with fine-tuning. The decision to publish scripts and guidelines while keeping the comment data access-restricted is defensible, and the data statement is a nice touch. I'd trust the label distributions and the platform-level comparisons more than anything else here.\n\nThe soft spot is exactly what the reader and the stress test flag: the age analysis assigns each comment the average audience age of its channel (141 channel-level values for 3,024 comments), then runs a chi-square test treating comments as independent. That test overstates the evidence, and the channel-level confound (topic, moderation, content style) is real. The authors do acknowledge this in the Limitations section, which is honest, but the abstract and conclusions still present age-based differences as a finding, and that is overselling. The LLM-extrapolated set reproducing the same age ordering doesn't help, because the same channel-average assignment is baked into both sets. Also, the absence of TikTok from the age analysis narrows the claim further.\n\nThe paper's own Section 10 is the most accurate summary of what we can conclude from the age results, and the authors deserve credit for writing it. The platform analysis (§7.2) is on firmer ground, and the LLM annotation comparison (§5–6) is useful for people building similar pipelines.\n\nWho is this for? Researchers working on German toxicity, content moderation, or sociolinguistic variation will find the dataset and annotation materials useful, even if they reanalyze the age relationships more carefully. The methodological lesson—about ecological inference in demographic social-media data—is also worth keeping in mind.\n\nI'd send this to peer review. It's not a clean acceptance; the authors should be asked to reframe the age findings as exploratory, add account-clustered tests or per-channel confidence intervals, and fix the abstract's overstatement. But the dataset itself is a legitimate contribution, and the transparency about limitations is a reason to engage rather than reject.","headline":"A useful German multi-platform toxicity dataset whose headline age findings rest on an acknowledged ecological inference; the resource is worth having, but the age patterns should be treated as exploratory.","tokens_in":16553,"tokens_out":1378,"would_cite":true,"duration_ms":16371,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Age splits toxic German comments, 33k-comment dataset shows","keywords":["toxic speech","age demographics","German dataset","LLM annotation","content moderation","social media platforms","disinformation","public broadcasting"],"falsifier":"Collect a sample of actual commenter ages, e.g., via an opt-in survey linked to anonymized comment IDs, and compare them to the channel-average age assignments; if the assignments are wrong for a substantial share of comments, or if a within-channel regression using actual commenter ages does not reproduce the age-label patterns, the age findings collapse. Comparing channels where audiences and commenters are known to diverge would provide a direct test.","tokens_in":15606,"feed_emoji":"💬","tokens_out":4579,"duration_ms":45579,"temperature":0.7,"pith_summary":"This paper builds a German-language dataset of 33,048 social media comments (3,024 human-annotated, 30,024 annotated by a fine-tuned LLM) from the public broadcaster funk's Instagram, TikTok, and YouTube channels, pairing toxicity labels with platform-provided audience-age estimates. It claims that toxicity patterns differ by age: users under 30 use more expressive, sarcastic, emoji-heavy language; the 31-35 group produces the most insults and disinformation; users over 35 engage more in devaluation and disinformation. It also reports platform differences, with Instagram showing the most toxic content. The dataset is presented as the first large-scale German resource combining toxic speech with any demographic marker, intended to support age-aware moderation and linguistic analysis.","feed_headline":"Age splits toxic German comments, 33k-comment dataset shows","feed_subtitle":"Younger users favor sarcasm and emojis; the 31-35 group leads in insults and disinformation.","key_machinery":"The dataset itself is the load-bearing artifact: two corpora sharing one 18-label annotation scheme with Target and Type classes, one human-annotated (3,024 comments, three annotators, majority vote, Fleiss kappa 0.64) and one LLM-annotated (30,024 comments, fine-tuned GPT-4o-mini). The age mechanism is channel-level age inference: platforms provide the age distribution of a channel's followers (Instagram) or viewers (YouTube), funk converts it into an average age per account, and every comment from that channel inherits that average; TikTok supplies no age data and is excluded from age analysis. Comments were pre-filtered with a toxic-keyword list, so the sample is enriched for problematic","core_discovery":"The central claim is that online toxicity is measurably age-structured in German comment sections: a chi-square test on the human-annotated data (chi-squared = 61.66, df = 36, p = 0.0049) supports a relationship between age group and label distribution, and the pattern repeats in the LLM-extrapolated set. Younger users (0-30) show the highest share of non-toxic comments and prefer emojis and sarcastic terms in their insults; the 31-35 group leads in insults, disinformation, and discrimination; the 35+ group leads in devaluation, disinformation, and religion- or ethnicity-based hate. Platform-level analysis finds Instagram highest in toxic content (22.53%), followed by YouTube (13.03%) and Ti","pith_inferences":["The channel-averaged age method likely compresses within-channel age variation; if commenters skew younger or older than followers/viewers, the reported age effects could be artifacts. A small opt-in survey of actual commenter ages compared with the channel-average assignments would test this.","Because comments were pre-filtered by toxic keywords, the reported proportions (e.g., 16.7% problematic) are conditional on that filter and cannot be read as base rates of toxicity across all comments on these channels.","The 31-35 peak in disinformation and devaluation could partly reflect topic effects, since different age groups may comment on different videos; the paper acknowledges this confound, but future within-topic analysis would be needed to separate age from subject matter.","The platform-level discrepancy between the human and LLM annotations (where YouTube's toxicity ranking shifts in the larger set) suggests that LLM-annotated demographic patterns, while broadly reproducing age trends, should be treated as exploratory for platform comparisons."],"forward_implications":["If the age patterns hold, moderation systems could weight categories by age: disinformation detection would be most relevant for the 31-35 and 35+ groups, while emoji and sarcasm detection would matter more for users under 30.","The platform differences imply that moderation strategies may need to be platform-specific, with Instagram requiring the most attention for insults, devaluation, and disinformation.","The fine-tuned LLM pipeline shows a practical path to scaling human toxicity labels to tens of thousands of comments for about four dollars, enabling broader demographic toxicity monitoring.","The 31-35 group appears as a distinct toxicity peak rather than a monotonic increase with age, pointing to a middle-age pattern worth investigating in other languages and platforms.","The dataset provides a benchmark for German toxicity annotation with demographic context, supporting training and evaluation of age-aware models."],"supporting_citations":[{"why":"Provides the keyword-based relevance filtering method used to consolidate comments before selection.","marker":"(Waseem and Hovy, 2016)"},{"why":"Supplies the foundational taxonomy framework that the annotation scheme's categories were inspired by.","marker":"Fillies et al. (2025)"},{"why":"Designed the best-performing prompt that this paper extends from binary to multi-class toxicity classification.","marker":"Das et al. (2024a)"},{"why":"Guides the iterative annotation process and biweekly annotator discussions used to maintain annotation quality.","marker":"(Vidgen and Derczynski, 2021)"},{"why":"Provides the datasheet framework used for the dataset's transparency statement.","marker":"Gebru et al. (2021)"},{"why":"Cited as prior evidence that language use varies with age, supporting the interpretation of age-linked toxicity differences as generational communication styles.","marker":"(Schwartz et al., 2013)"},{"why":"Establishes that age influences language use and topic preferences, motivating the study's focus on age as a demographic dimension.","marker":"(Eckert, 2012)"}],"fun_headline_variants":["Age splits toxic German comments: 33k dataset","Young mock, old misinform: German toxicity by age","Instagram tops toxic share in German public comments","Chi-square confirms age links to German hate speech","Age-based toxicity patterns in German online comments"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire age analysis assumes that people who comment on a channel have the same age distribution as the channel's overall audience, because each comment is assigned the channel's average follower or viewer age rather than the commenter's actual age.","fun_headline_variants_meta":{"raw":{"variants":["Age splits toxic German comments: 33k dataset","Young mock, old misinform: German toxicity by age","Instagram tops toxic share in German public comments","Chi-square confirms age links to German hate speech","Age-based toxicity patterns in German online comments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1106,"prompt_tokens":719,"completion_tokens":387,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":325}},"tokens_in":463,"tokens_out":387,"duration_ms":5424,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:52:28.070539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a sample of actual commenter ages, e.g., via an opt-in survey linked to anonymized comment IDs, and compare them to the channel-average age assignments; if the assignments are wrong for a substantial share of comments, or if a within-channel regression using actual commenter ages does not reproduce the age-label patterns, the age findings collapse. Comparing channels where audiences and commenters are known to diverge would provide a direct test.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the foundational taxonomy framework that the annotation scheme's categories were inspired by."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that age influences language use and topic preferences, motivating the study's focus on age as a demographic dimension."}],"review_version":1}