{"id":"6789c7b2-94ec-4d9e-b0fa-25d9e83c4ecd","arxiv_id":"1908.04531","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper creates a new annotated Danish dataset for offensive language detection and reports macro F1 scores up to 0.73 for targeted insult classification.","lead":"This paper introduces the first Danish-language dataset for detecting offensive and hateful speech on social media, built from Facebook and Reddit comments. It also evaluates several standard machine learning models for English and Danish, reporting around 0.70 macro F1 for Danish offensive language detection.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 7 reveals profanity-bearing posts in the Danish test set labeled NOT, contradicting the paper's own final annotation rule; reported F1 scores rest on labels the authors themselves flag as wrong.","rationale":"The reader's weakest assumption is annotation reliability, and my stress-test confirms that this is the load-bearing concern. I found an even more specific piece of evidence than low inter-annotator agreement: the paper's own Section 7 analysis gives concrete examples of test-set posts that contain profanity but are labeled NOT, directly contradicting the final annotation rule stated in Section 3.1. This internal inconsistency is not a matter of disputed definitions or outside consensus; the authors themselves acknowledge that those test labels are wrong according to their stated guidelines. If the gold test set contains known mislabels, the reported macro F1 scores cannot be interpreted as genuine performance estimates. The dataset may still be a valuable resource, and the models may still be reasonable baselines, but the headline numbers in the abstract are not trustworthy until the test labels are verified. This does not change the reader's CONDITIONAL verdict: the conditions should be (1) re-annotation or at least an audit of the test set, and (2) release of the dataset and code. My concern is the same class of issue the reader identified, so I agree with the reader's weakest-assumption analysis. Verdict remains conditional.","tokens_in":12749,"tokens_out":3549,"duration_ms":35208,"concrete_test":"Take a stratified random sample of 300 Danish test posts, oversampling the OFF/TIN classes, and have two independent fluent Danish annotators label them with the final guidelines from Section 3.1. Compute pairwise Jaccard/Cohen's kappa and, in particular, count posts containing profanity that are labeled NOT. If kappa is below 0.7 or the profanity-bearing NOT count is non-negligible, the reported macro F1 values cannot be taken as valid and should be recomputed after re-annotation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claims (macro F1 0.70/0.73/0.63 for Danish) are computed against a test set whose labels are not consistent with the stated annotation schema. In Section 3.1 the authors adopt the rule that \"any post containing any form of profanity should automatically be labeled as offensive.\" Yet Section 7 reports that the classifier labels profanity-bearing posts as offensive while the test set labels them NOT, giving \"Are you fucking serious?\" and \"Fuck I cried in this scene\" as examples that \"according to annotation guidelines should be classed as offensive.\" This is a direct admission that at least some test labels violate the final guidelines. The reliability problem is compounded by the annotation process: only the first 100 posts were double-annotated, with Jaccard indices of only 41.9% (A), 39.1% (B), and 42.8% (C); after guideline refinement, the remaining ~3,500 posts were annotated by a single author. No re-annotation of the 100 warm-up posts or a second pass on the full set is described. Since the gold labels are used to compute precision, recall, and F1, known label noise and known guideline violations make the reported numbers an unreliable estimate of real-world performance. The abstract's rounded 0.70/0.73/0.63 figures are therefore not well supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a Danish dataset of 3,600 user-generated comments from Reddit and Facebook, annotated following the OffensEval schema for three sub-tasks: offensive language identification (A), categorization of offensive language type (B), and offensive language target identification (C). It compares logistic regression and several BiLSTM-based classifiers on this dataset and on the English OLID dataset, reporting macro-averaged F1 scores of 0.70 for Danish sub-task A, 0.73 for sub-task B, and 0.63 for sub-task C. The paper also analyzes misclassifications and claims to provide the first Danish dataset of its kind.","tokens_in":12980,"tokens_out":3628,"duration_ms":36551,"significance":"If the dataset and reported results are reliable, the contribution is useful for under-resourced Danish NLP: it provides a new resource, follows an established annotation schema, compares multiple standard models, and includes a data statement. The availability of the classifiers and dataset under CC-BY is a strength, and the error analysis in Section 7 is a useful addition. However, the central quantitative claims rest on gold labels whose reliability is not established and that are, by the authors' own admission, inconsistent with the final annotation guidelines in at least some test instances. The contribution is therefore significant in intention but the specific performance numbers are not currently well supported.","major_comments":[{"comment":"Section 7 states that the classifier labels profanity-bearing content as offensive while the test set labels such posts NOT, giving 'Are you fucking serious?' and 'Fuck I cried in this scene' as examples that 'according to annotation guidelines should be classed as offensive.' This directly contradicts the final annotation rule in Section 3.1 that any post containing any form of profanity should automatically be labeled offensive. Since these test labels are the gold labels used to compute Table 4 and the class-wise scores in Table 5, the reported Danish sub-task A macro F1 of 0.699 is computed against a test set that the authors themselves identify as containing wrong labels. The reported number is therefore not a trustworthy estimate of performance under the final annotation schema.","section":"Section 7, Sub-task A"},{"comment":"Inter-annotator agreement was measured only on the first 100 posts, with Jaccard indices of 41.9% for sub-task A, 39.1% for sub-task B, and 42.8% for sub-task C. After the guidelines were refined, the remaining approximately 3,500 posts were annotated by a single author, and no final agreement study, second pass, or re-annotation of the warm-up posts is reported. Given the low agreement on the initial set, the reliability of the final labels is not established. The classifier scores in Tables 4, 7, and 10 should be interpreted with this fundamental uncertainty in mind.","section":"Section 3.1"},{"comment":"In the error analysis for sub-task B, the authors report that a test-set sample containing 'HillaryForPrison' is labeled as untargeted, although they characterize it as a clear targeted insult. This is another direct example of a test-set label violating the authors' own understanding of the annotation schema. It affects the Danish sub-task B evaluation reported in Table 7 and indicates that label noise is not confined to sub-task A, further undermining the reliability of the reported F1 scores.","section":"Section 7, Sub-task B"}],"minor_comments":[{"comment":"The text says 'Recall and precision scores are lower for UNT than TIN (Table 5)', but Table 5 reports sub-task A results; the relevant table is Table 8.","section":"Section 6, Sub-task B (English)"},{"comment":"The terms 'macro averaged F1-score', 'F1macro', and 'macro F1' are used interchangeably; please standardize the notation.","section":"Throughout"},{"comment":"There are several typos and possible OCR artifacts, such as 'T op-level features', 'the the' in the Background section, and 'Futher' in Section 6. A careful proofreading pass is needed.","section":"Section 4"},{"comment":"The baseline rows labeled 'All NOT' report macro F1 values; it would be helpful to state explicitly how the macro average is computed for a one-class baseline.","section":"Tables 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue is the inconsistency between the final annotation guidelines and the test-set labels. This is fixable within the scope of the paper by re-annotating the test set (and ideally the full dataset) according to the final guidelines, re-running the experiments, and reporting updated numbers. Please also ensure that the dataset is made available to reviewers so that the label quality can be independently assessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know about this paper for the dataset, not for the model numbers. It introduces the first Danish corpus for offensive language detection: 3,600 comments from Reddit and Facebook, annotated with the OLID scheme for offensiveness, targeted vs. untargeted, and target type. That fills a real gap for a low-resource language, and the data statement is a nice touch.\n\nThe classification work is standard—logistic regression and BiLSTMs—so novelty is concentrated in the resource. The error analysis in Section 7 is honest and useful; it identifies obfuscation and keyword bias as failure modes.\n\nThe soft spots are real and mostly about label reliability. Only the first 100 posts were double-annotated, with Jaccard indices around 40%, which is low. After revising guidelines, a single annotator labeled the remaining ~3,500 posts, with no reported agreement. More importantly, Section 7 itself gives two test-set examples—\"Are you fucking serious?\" and \"Fuck I cried in this scene\"—that are labeled NOT but, under the paper's own rule that any profanity means offensive, should be OFF. So the gold labels are internally inconsistent, and the reported F1 numbers (0.70/0.73/0.63) are computed against that noisy test set. The effect is likely small in aggregate, but it means the exact numbers shouldn't be taken as reliable estimates of real-world performance.\n\nOther issues: no error bars or significance tests; the dataset isn't publicly available yet (pending a shared task), which limits reproducibility; and the abstract's mention of sharing across languages isn't actually tested—models are trained separately per language.\n\nNone of this is fatal. The resource is genuinely valuable for Danish NLP, and the problems are fixable in revision: re-annotate the test set, or at least quantify how many labels violate the guidelines, and release the data. I'd send this to peer review with a request for major revision on annotation quality. The paper deserves referee time because the dataset fills a clear gap, and the authors are transparent about limitations—they flag the mislabeled examples themselves, which is more than many papers do.","headline":"First Danish offensive-language dataset, but the weak annotation reliability means the headline F1 scores should be read as provisional until the labels are cleaned or re-checked.","tokens_in":13484,"tokens_out":1866,"would_cite":true,"duration_ms":21602,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"First Danish dataset for offensive-language detection, annotated for type and target, supports classifiers reaching macro F1 of 0.70.","keywords":["offensive language detection","hate speech","Danish language","dataset annotation","social media","natural language processing","deep learning classifiers","multilingual NLP"],"falsifier":"Take a random sample of roughly 300 Danish posts from the corpus, have several independent Danish-speaking annotators label them with the same guidelines, and measure agreement with the published labels; if the original labels cannot be reproduced or agreement is far below the pilot Jaccard values, the reported F1 scores do not reflect real detection quality.","tokens_in":12534,"feed_emoji":"💬","tokens_out":6516,"duration_ms":59559,"temperature":0.7,"pith_summary":"The paper sets out to close a gap: offensive-language detection has mostly been built for English, and no Danish dataset existed for it. It constructs a 3,600-comment Danish corpus from Reddit and Facebook, annotated with the same three-level scheme used for English, and trains four classifiers on both languages. The central claim is that this scheme transfers to Danish and that automatic detection reaches usable levels: macro F1 of 0.70 for distinguishing offensive posts, 0.73 for detecting whether an offensive post is targeted, and 0.63 for identifying the target type. If that is right, Danish-language platforms and moderation teams gain a first practical resource for detecting hate speech and cyberbullying.","feed_headline":"First Danish hate-speech dataset trains detectors to F1 0.70","feed_subtitle":"A 3,600-comment Danish corpus from Reddit and Facebook lets models flag offensive posts and their targets","key_machinery":"The load-bearing object is a three-level annotation scheme for offensive language: first decide whether a post is offensive, then whether the offence is targeted at someone, then whether the target is an individual, a group, or something else. This scheme comes from the English-language task that the paper adopts, and the paper's Danish corpus instantiates it for the first time. The argument is carried by a suite of classifiers — a logistic regression baseline and three BiLSTM variants using learned embeddings, fixed pre-trained embeddings, or augmented features such as sentiment scores and n-grams — evaluated on both the English and Danish datasets under the same three subtasks.","core_discovery":"The paper's contribution is the first Danish dataset for offensive-language detection, together with evidence that detectors built on it work at levels close to English. The corpus consists of 3,600 user-generated comments gathered from Danish Reddit and Facebook, labeled according to three subtasks: offensive versus not offensive, targeted versus untargeted, and target type (individual, group, other). The best Danish systems reach macro F1 scores of 0.70, 0.73, and 0.63 on these three subtasks, while the English systems reach 0.74, 0.62, and 0.56. The result is presented as showing that a shared annotation scheme, combined with pre-trained word embeddings and a mix of lexical and linguistic features, supports automatic offensive-language detection in Danish despite a small, imbalanced training set.","pith_inferences":["We infer that the reported F1 scores should be treated as upper bounds until the labels themselves are validated: only 100 posts were double-annotated, with Jaccard agreement of 39–43%, and the remaining 3,500 posts were labeled by one person after the guidelines were revised.","The same low agreement on the pilot set suggests the annotation guidelines are ambiguous for Danish, especially for context-dependent posts; a follow-up with a larger annotator pool and adjudication would clarify how much of the measured performance is learning the task versus learning one annotator's interpretation.","Because Danish is closely related to Norwegian and Swedish, the dataset could seed cross-lingual transfer for offensive-language detection in other low-resource Scandinavian settings, though this is beyond what the paper tests."],"forward_implications":["Danish content moderation can use these models to flag offensive comments at accuracy levels comparable to English systems, with a concrete best macro F1 of 0.70 for the binary task.","The released corpus gives Danish NLP a benchmark for future work on offensive and hateful language, including cross-platform tests since the data comes from two different platforms.","The finding that a fixed pre-trained embedding model is best for English but worst for Danish suggests no single architecture dominates, so practical deployments should tune per language and task.","Class imbalance is the main bottleneck: with only about 12% offensive posts in Danish, recall for the minority classes is low, so methods that rebalance training data are likely to matter more than model choice.","The same annotation scheme and model family used for English yields workable Danish results, supporting the broader idea that this task can be exported to other under-resourced languages."],"supporting_citations":[{"why":"Supplies the three-subtask annotation guidelines and the English dataset used to build and compare the classifiers.","marker":"[1]"},{"why":"Provides the logistic-regression feature set and the additional English hate-speech dataset used as training data.","marker":"[3]"},{"why":"Documents the absence of an existing Danish offensive-language dataset, which motivates the corpus construction.","marker":"[5]"},{"why":"Defines the shared-task subtasks (offensive identification, targetedness, target type) that structure the dataset and evaluation.","marker":"[18]"},{"why":"Provides the Jaccard index used to measure the low double-annotation agreement on the pilot set.","marker":"[19]"},{"why":"Supplies the pre-trained FastText embeddings used by the best English model and by the Danish BiLSTM variants.","marker":"[24]"}],"fun_headline_variants":["First Danish hate speech dataset boosts detection to F1 0.70","Danish hate speech detection: new dataset, F1 up to 0.73","First Danish Reddit and Facebook corpus enables hate speech detection","Danish offensive language detector hits F1 0.73 on targeted posts","New Danish dataset trains hate speech models to F1 0.70"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset labels are assumed to be reliable enough to train and evaluate the detectors, but only 100 posts were double-annotated with low agreement, and the remaining 3,500 were annotated by a single person after a guideline change.","fun_headline_variants_meta":{"raw":{"variants":["First Danish hate speech dataset boosts detection to F1 0.70","Danish hate speech detection: new dataset, F1 up to 0.73","First Danish Reddit and Facebook corpus enables hate speech detection","Danish offensive language detector hits F1 0.73 on targeted posts","New Danish dataset trains hate speech models to F1 0.70"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000773,"raw_usage":{"total_tokens":3462,"prompt_tokens":1027,"completion_tokens":2435,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":2340}},"tokens_in":643,"tokens_out":2435,"duration_ms":17752,"temperature":1.0,"reasoning_tokens":2340,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:39:05.458420+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly 300 Danish posts from the corpus, have several independent Danish-speaking annotators label them with the same guidelines, and measure agreement with the published labels; if the original labels cannot be reproduced or agreement is far below the pilot Jaccard values, the reported F1 scores do not reflect real detection quality.","supporting_citations":[{"cited_title":"Predicting the type and target of offensive posts in social media","cited_arxiv_id":null,"evidence_quote":"Supplies the three-subtask annotation guidelines and the English dataset used to build and compare the classifiers."},{"cited_title":"Automated hate speech detection and the problem of offensive language","cited_arxiv_id":null,"evidence_quote":"Provides the logistic-regression feature set and the additional English hate-speech dataset used as training data."},{"cited_title":"The Lacunae of Danish Natural Language Processing","cited_arxiv_id":null,"evidence_quote":"Documents the absence of an existing Danish offensive-language dataset, which motivates the corpus construction."},{"cited_title":"Semeval- 2019 task 6: Identifying and categorizing offensive langua ge in social media (offenseval)","cited_arxiv_id":null,"evidence_quote":"Defines the shared-task subtasks (offensive identification, targetedness, target type) that structure the dataset and evaluation."},{"cited_title":"Similarity measures in scientometr ic research: The jaccard index versus salton’s cosine formula","cited_arxiv_id":null,"evidence_quote":"Provides the Jaccard index used to measure the low double-annotation agreement on the pilot set."},{"cited_title":"Advances in pre- training distributed word representations","cited_arxiv_id":null,"evidence_quote":"Supplies the pre-trained FastText embeddings used by the best English model and by the Danish BiLSTM variants."}],"review_version":1}