{"id":"a260dbb4-3af7-4eb5-8486-ff0a5ea7a1d0","arxiv_id":"2411.15462","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HateDay provides a representative global sample of one day of Twitter and shows that hate speech detection models achieve much lower average precision on real-world data than on academic datasets.","lead":"Researchers released HateDay, a dataset of 240,000 tweets randomly sampled from one day on Twitter across eight languages and four countries, annotated for hate speech. It shows that hate speech is rare but that current detection models perform far worse on this real-world sample than on academic benchmarks, especially outside Europe.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Raw AP comparisons across HateDay (0.7% prevalence), enriched academic datasets, and balanced functional tests conflate base rates with model skill; the central overestimation claim needs a prevalence-normalized metric.","rationale":"The reader's weakest_assumption identifies exactly this issue: AP values are compared across datasets with very different hate prevalence, and no random baseline or prevalence-normalized metric is provided in the main comparison. I agree that this is the most load-bearing concern. The paper makes a valuable empirical contribution with a new representative dataset, careful annotation, and descriptive findings on prevalence and targets, and the moderation analysis is thoughtfully derived from precision and recall. However, the headline claim that academic evaluations overestimate real-world detection performance rests on raw AP comparisons across evaluation sets with wildly different base rates. Because a random classifier's AP equals the base rate, the 40% versus 9.4% gap may reflect dataset composition rather than model failure. The paper's own limitations section acknowledges the small number of positives and wide confidence intervals on HateDay, which compounds the problem but is distinct from the base-rate confound. The proposed concrete test would settle the issue directly using existing scores and labels. I therefore keep the reader's conditional verdict: the paper should be accepted only after the authors re-analyze the performance comparison in a prevalence-normalized way and report whether the overestimation conclusion survives.","tokens_in":23667,"tokens_out":7126,"duration_ms":72196,"concrete_test":"Using the already-computed model scores behind Tables 2 and 3, compute for every language/country and each evaluation set: (i) the hate prevalence b from the labels actually used in evaluation, (ii) the random-baseline AP (equal to b), (iii) normalized AP = AP / b (lift over random), and (iv) precision at fixed recalls such as 20% and 50%. Then compare AD versus HD on the normalized metrics, with bootstrap confidence intervals as in Tables 5 and 6. If the normalized AD-HD gap is small, changes sign, or becomes statistically insignificant, the paper's overestimation claim is an artifact of base-rate differences; if the normalized gap remains large and in the same direction, the concern is resolved. This requires no new data collection, only the label prevalence in the 10% academic samples and existing model scores.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in Section 4.1 is that average precision is 9.4% on HateDay versus 40% on academic datasets and 87.2% on functional tests, implying academic evaluations greatly overestimate real-world performance. This comparison is not interpretable as a statement about model skill because average precision is prevalence-dependent: a random classifier achieves AP equal to the positive-class base rate. HateDay's hate prevalence is about 0.7%, so random AP there is roughly 0.7%; academic hate-speech supersets are constructed from deliberately enriched datasets with far higher positive rates, and HateCheck is balanced by design. Thus, a large part of the 40%-versus-9.4% gap is a mechanical consequence of label prevalence rather than of worse ranking in real-world data. In fact, a model with AP 9.4% on HateDay is roughly 13x the random baseline, whereas a model with AP 40% on a 50%-prevalence academic set is no better than chance. The paper never reports a prevalence-normalized comparison (AP lift over random, ROC-AUC, or precision at fixed recall) in its main evaluation, so Tables 2 and 3 cannot support the overestimation claim as stated. This is not a minor caveat: the abstract and conclusion rest on this comparison, and Appendix E.3 shows the authors are aware of base-rate effects in the moderation analysis. The same concern applies to the functional-test comparison, whose random baseline is near 50%. Without normalization, the headline result could even reverse.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces HATEDAY, a dataset of 240,000 tweets sampled uniformly at random from all tweets posted on September 21, 2022, covering eight languages and four English-majority countries, with 20,000 tweets per stratum. Three annotators labeled each tweet for hateful, offensive, or neutral content following a prescriptive guideline, with majority voting and high agreement. The authors estimate hate speech prevalence (average 0.7%), analyze target composition, and evaluate public hate speech detection models on HATEDAY, academic supersets, and HateCheck functional tests using average precision. They report that average precision is 9.4% on HATEDAY versus 40% on academic datasets and 87.2% on HateCheck, and draw conclusions about performance overestimation, cross-lingual gaps, and the infeasibility of fully automated moderation.","tokens_in":23946,"tokens_out":8469,"duration_ms":76066,"significance":"The dataset is a valuable resource: it is publicly released, uses a rigorous random sampling design, documents annotator demographics and guidelines, and provides bootstrapped confidence intervals. The cross-geographic scope (eight languages and four countries from one platform day) is substantially broader than prior representative-sample evaluations such as NaijaHate. If the main result survives reanalysis with a prevalence-appropriate metric, it would be an important corrective to benchmark-based evaluations of hate speech detection and would strengthen the case for human-in-the-loop moderation. The paper also contributes a concrete analysis of target-alignment and offensive false positives. However, the headline quantitative claim is currently under-supported because the chosen metric is not comparable across datasets with very different class prevalences.","major_comments":[{"comment":"The central comparison of average precision across HATEDAY, academic supersets, and HateCheck is not interpretable as a statement about model skill, because AP is prevalence-dependent. For a random scoring model, the expected AP equals the positive-class base rate. HATEDAY has roughly 0.7% hate prevalence, so a random model has AP near 0.7%; the academic supersets are enriched and HateCheck is approximately balanced, so random AP there is much higher. Consequently, the report that AP is 9.4% on HATEDAY versus 40% on academic datasets and 87.2% on HateCheck conflates base rates with ranking quality. In fact, an AP of 9.4% at 0.7% prevalence is about 13 times the random baseline, while an AP of 40% on a 50%-prevalence benchmark is no better than chance. The paper should report a prevalence-normalized metric (for example, AP lift over the random baseline, ROC-AUC, or precision at a fixed recall level) and include random baselines in Tables 2 and 3. Without this, the abstract's claim that 'evaluations on academic datasets greatly overestimate real-world detection performance' is not supported as stated.","section":"Section 4.1, Tables 2 and 3; abstract and Section 5"},{"comment":"The aggregate figures 9.4%, 40%, and 87.2% are not defined in the text, and the 'Best OS' rows select different models per language/country based on their HATEDAY AP. This makes the cross-dataset comparison sensitive to model selection. For example, if the best model on HATEDAY happens to be weaker on academic data than the best model on academic data, the gap is inflated. Please specify exactly how the aggregates are computed (which models, which strata, weighting), and verify that the main overestimation conclusion holds when using a fixed model per language or paired per-model comparisons across evaluation sets.","section":"Section 4.1, Tables 2 and 3"},{"comment":"The number of hate positives in several strata is very small (e.g., 31 in Kenya), and the paper itself notes the resulting uncertainty in Tables 5 and 6. This is particularly consequential for the moderation feasibility curves in Figure 4, where recall targets of 80-90% are extrapolated from very few positives, and for the cross-country ranking (e.g., Kenya's AP of 9.1% with a 95% interval of ±6.1). The paper should provide bootstrap or other uncertainty estimates for the moderation curves and state which cross-language/country performance differences are statistically significant, rather than relying on the noisy point estimates in the headline.","section":"Section 2.3, Section 4.3, Limitations"}],"minor_comments":[{"comment":"The main text says 36 annotators (three per 12 strata), but Appendix B.1 opens with 'a team of 30 annotators'; please correct the inconsistency in B.1.","section":"Section 2.2 and Appendix B.1"},{"comment":"The caption contains 'T witter' with an unintended space; it should be 'Twitter'.","section":"Figure 1 caption"},{"comment":"The name 'Perpective API' is misspelled twice in the model-type paragraph; it should be 'Perspective API'.","section":"Section 4.1"},{"comment":"The abstract describes HATEDAY as 'representative of social media settings', but the data reflects one platform and one 24-hour period. The Limitations section already qualifies this; the abstract and conclusion should include the same qualification or use wording such as 'representative of a day on Twitter'.","section":"Abstract and Section 5"},{"comment":"For the HateCheck columns, the paper should remind readers that HateCheck is a curated challenge suite rather than a random sample, so its AP values are not directly interpretable as real-world performance estimates; the 'overestimation' interpretation applies primarily to the academic dataset comparison.","section":"Tables 2 and 5"}],"recommendation":"major_revision","confidential_remarks":"The base-rate confound is the main threat to the paper's headline. A reanalysis with normalized metrics and random baselines is feasible within the paper's scope and will likely preserve a qualitative overestimation finding, though the magnitude may shrink. The dataset and annotation effort are strong, and the paper fits the journal's scope. I recommend major revision rather than reject, because the issue is methodological and fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this is a genuinely useful dataset and the authors mostly deserve credit for careful collection and annotation. But the headline quantitative claim—that academic evaluations \"greatly overestimate\" real-world detection performance—rests on an average-precision comparison that mixes model skill with label prevalence. The paper should be reviewed, but the central comparison needs to be redone.\n\nWhat's new: HateDay is the first random-sample, one-day, multi-context hate speech resource: 240k tweets, 8 languages plus 4 English-speaking countries, consistent annotation guidelines, 3 annotators per item, majority vote, bootstrapped confidence intervals. That is a real contribution. The prevalence and composition results (hate around 0.7%, offensive far more common, target mix varying by context) are credible and useful. The moderation cost-recall analysis is a solid practical addition, and the PR-curve evidence that precision and recall are both low in deployment supports the authors' skepticism about fully automatic moderation. The limitations section is unusually candid, and the NaijaHate precedent is cited and discussed rather than hidden.\n\nThe soft spot is Section 4.1. AP for a random ranker equals the positive-class base rate. HateDay has roughly 0.7% hate; academic supersets are enriched, and HateCheck is balanced. So 9.4% AP on HateDay and 40% AP on academic sets are not comparable as measures of ranking quality. In fact, 9.4% on a 0.7% base rate is about 13x random, while 40% on a 50%-prevalence set is below random. The paper never reports a prevalence-normalized metric in the main comparison, so the \"greatly overestimate\" sentence in the abstract is not supported as written. This is load-bearing for the framing; it is not fatal to the moderation recommendation, which the PR curves carry. But the headline numbers need reanalysis and reframing.\n\nMinor issues: the abstract says \"all tweets\" while retweets were dropped; the annotator count is 36 in Section 2.2 but 30 in Appendix B.1; and there is no released evaluation script, which would make the numbers easier to check.\n\nThis paper is for anyone building or evaluating hate speech detection, especially for non-English and non-US contexts. It deserves a serious referee after the AP comparison is fixed. I would send it to review with a request for revision, not desk-reject.","headline":"HateDay is a genuinely useful representative dataset, but its headline claim that academic evaluations overestimate real-world performance rests on an average-precision comparison that conflates base rates with model skill.","tokens_in":24463,"tokens_out":4090,"would_cite":true,"duration_ms":39296,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HATEDAY, a random sample of 240,000 tweets from one day on Twitter, shows public hate speech detectors reach only 9.4% average precision in real-world data versus 40% on academic datasets.","keywords":["hate speech detection","representative dataset","Twitter","average precision","cross-lingual evaluation","content moderation","human-in-the-loop","offensive language"],"falsifier":"Evaluate the same models on a HATEDAY subsample where hateful tweets are oversampled to match the prevalence of academic datasets; if average precision climbs to near 40%, the reported overestimation is largely a prevalence artifact, and if it stays near 9.4%, the paper's conclusion stands.","tokens_in":23499,"feed_emoji":"🤬","tokens_out":5743,"duration_ms":46881,"temperature":0.7,"pith_summary":"HATEDAY is a dataset of 240,000 tweets randomly sampled from everything posted on Twitter on September 21, 2022, covering eight languages and four English-speaking countries. The paper's central claim is that this sample represents what a real moderation system would encounter, and that on such data publicly available hate speech detectors achieve only 9.4% average precision, versus 40% on academic datasets and 87.2% on functional tests. If correct, the field's standard evaluations substantially overstate real-world detection ability, especially for Arabic, Indonesian, and Turkish. The paper also argues that fully automatic moderation is therefore unsafe, while human-in-the-loop review can work only with substantial human effort.","feed_headline":"Hate speech detectors hit 9.4% on a real day of Twitter","feed_subtitle":"A random sample of 240,000 tweets shows academic benchmarks overstate model performance by about four times.","key_machinery":"The load-bearing mechanism is HATEDAY itself: 240,000 tweets randomly sampled from a full 24-hour corpus of Twitter posts, with 20,000 tweets per language or country, labeled by three annotators using a prescriptive hateful, offensive, or neutral taxonomy with majority-vote labels. This random sampling is what makes the evaluation representative of real-world conditions, and it is paired with average precision as the evaluation metric because it handles severe class imbalance. The second mechanism is the target-alignment measure, a cosine similarity between the target focus of academic datasets and the targets observed in HATEDAY, which the paper uses to explain cross-language and cross-country performance differences.","core_discovery":"The central discovery is that real-world hate speech detection performance is far below what academic benchmarks report, and that the gap is systematic rather than random. Using HATEDAY, the paper measures the prevalence of hate across languages and countries at about 0.7% of posts on average, then evaluates publicly available supervised and zero-shot models, finding average precision of 9.4% on HATEDAY versus 40% on academic datasets and 87.2% on HateCheck functional tests. The paper traces the gap to two main drivers: offensive tweets crowding the top of the hate score distribution, and a mismatch between targets emphasized in academic datasets such as religion and race and the political and gender hate that dominates real-world data; target alignment correlates with performance (Pearson's r = 0.76) while dataset size does not. The paper concludes that fully automatic moderation is not viable and that human-in-the-loop moderation requires reviewing at least 10% of daily tweets to catch more than 80% of hate.","pith_inferences":["Our inference: a prevalence-matched control, such as evaluating the same models on a HATEDAY subsample oversampled to academic base rates, would isolate how much of the 40 to 9.4 percentage-point drop is a pure class-imbalance artifact; part of the gap would likely persist but shrink.","Our inference: the target-alignment correlation suggests a concrete fix, namely rebalancing training data toward political hate speech, especially for English and US contexts, and that this fix should transfer to other harm-detection tasks.","Our inference: the same random-day sampling design could be applied to misinformation or toxicity detection, where non-representative benchmarks are also the norm and where real-world prevalence may similarly be much lower than benchmark prevalence."],"forward_implications":["Removing retweets and sampling uniformly from a full day's output makes HATEDAY the first evaluation set whose class balance and topic mix match what a deployed moderator would actually see.","Because average precision on HATEDAY is 9.4% versus 40% on academic sets, models that look strong in the lab would flood a real moderation queue with false positives and still miss most hate.","Human-in-the-loop moderation can catch 70 to 90% of hate in most languages, but only if human reviewers examine at least 10% of all daily tweets.","Performance gaps across languages track how well academic datasets match real-world hate targets, not how much annotated data exists for a language.","Public models are currently ill-suited for fully automatic hate speech moderation, so deployments should assume a substantial human review component."],"supporting_citations":[{"why":"Supplies TWITTER DAY, the complete one-day corpus of all Twitter posts from which HATEDAY is randomly sampled.","marker":"Pfeffer et al., 2023"},{"why":"Supplies the hateful, offensive, and neutral classification taxonomy and the distinction that drives the false-positive analysis.","marker":"Davidson et al., 2017"},{"why":"Supplies HateCheck functional tests, the second comparison evaluation setting used to show overestimation.","marker":"Röttger et al., 2021"},{"why":"Supplies the language-level academic dataset supersets used as the academic evaluation condition.","marker":"Tonneau et al., 2024a"},{"why":"Supplies the Nigerian representative-data evaluation and classifier that HATEDAY extends to more languages and countries.","marker":"Tonneau et al., 2024b"},{"why":"Supplies target-focus shares of academic hate speech datasets used for the target-alignment and correlation analysis.","marker":"Yu et al., 2024"},{"why":"Supplies Perspective API, the widely used toxicity detection system evaluated as a supervised baseline.","marker":"Lees et al., 2022"},{"why":"Supplies Aya 23, one of the zero-shot large language model baselines.","marker":"Aryabumi et al., 2024"},{"why":"Supplies Llama 3.1, the other zero-shot large language model baseline.","marker":"Dubey et al., 2024"}],"fun_headline_variants":["Hate speech detectors hit just 9.4% precision on real Twitter","Real Twitter data shows hate detection precision of only 9.4%","Academic hate speech benchmarks overstate real detection by 4x","To catch 80% of hate on Twitter, humans must review 10% of posts","HateDay: real-world detection 4x worse than benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gap in average precision, the metric that rewards ranking hateful tweets above non-hateful ones, measures model skill rather than the rarity of hate in real-world data, since a random classifier's average precision equals the base rate.","fun_headline_variants_meta":{"raw":{"variants":["Hate speech detectors hit just 9.4% precision on real Twitter","Real Twitter data shows hate detection precision of only 9.4%","Academic hate speech benchmarks overstate real detection by 4x","To catch 80% of hate on Twitter, humans must review 10% of posts","HateDay: real-world detection 4x worse than benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000847,"raw_usage":{"total_tokens":3692,"prompt_tokens":958,"completion_tokens":2734,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2636}},"tokens_in":574,"tokens_out":2734,"duration_ms":17129,"temperature":1.0,"reasoning_tokens":2636,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:16:06.828841+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the same models on a HATEDAY subsample where hateful tweets are oversampled to match the prevalence of academic datasets; if average precision climbs to near 40%, the reported overestimation is largely a prevalence artifact, and if it stays near 9.4%, the paper's conclusion stands.","supporting_citations":[],"review_version":1}