{"id":"64cd3778-7271-4225-81c7-d712b0947af2","arxiv_id":"2505.11969","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper presents a new 10,000-tweet Arabic hate speech corpus with seven target categories and reports transformer benchmark results, but its novelty and reliability claims are undermined by internal inconsistencies.","lead":"This paper introduces a dataset of 10,000 Arabic tweets annotated for offensive content and hate speech targets like religion, gender, and ethnicity, with benchmark results for Arabic language models. However, the claim of being the first multilabel dataset is contradicted by the authors' own earlier work, and the sample selection is biased.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The novelty claim is false on its face: Table 1 lists the authors' own 2024 LREC-COLING multilabel dataset, yet the paper says first-time multilabel annotation.","rationale":"The reader's weakest assumption concerned ASAD-based sampling, which is a legitimate reliability risk. However, the strongest claim in the paper is the novelty of the multilabel corpus, and that claim is contradicted directly by the paper's own Table 1 and its citation of Zaghouani, Mubarak, and Biswas 2024, a prior multilabel Arabic hate speech dataset by the same group. This makes false novelty the most load-bearing concern because it invalidates the stated contribution without requiring speculation about classifier noise. The dataset may still be useful as an additional resource, but the paper cannot claim 'first time' or 'unique' in its current form. The inter-annotator agreement and table inconsistencies are secondary but reinforce the need for revision. I agree with the reader's overall rejection verdict, though my principal concern differs from the weakest assumption listed.","tokens_in":7881,"tokens_out":6240,"duration_ms":63143,"concrete_test":"Download the current Zenodo corpus (DOI 10.5281/zenodo.14669917) and the 'So hateful!' LREC-COLING 2024 dataset; compute the tweet-ID intersection and compare the label schemas. If a substantial fraction of the 10,000 tweets appears in the 15,965-tweet prior corpus, or if both use the same seven-category target schema without documented intentional differences, the uniqueness/first-time claim is unsupported. Also check whether the Zenodo record links to the prior dataset as a previous version.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim of uniqueness is internally contradicted. In Related Works, Table 1 lists 'Zaghouani, Mubarak, and Biswas 2024' with 'Multilabel: Yes', and the text then states: 'Annotating each tweet to all the possible labels has been proposed for the first time in our dataset.' The cited work is the same group's LREC-COLING 2024 paper 'So hateful! Building a multi-label hate speech annotated Arabic dataset' (15,965 tweets). This is not a disagreement with outside consensus; it is the paper contradicting its own evidence. At a minimum, the current dataset would need to be positioned as an extension or re-annotation with a documented delta over the prior resource, and no such framing is provided. Additional inconsistencies, such as the abstract reporting IAA of 0.86 and 0.71 while the body reports Fleiss' kappa of 0.8143 without measuring target-level agreement, and Table 2 counts summing to 10,048 with percentages off by a factor of 100, further weaken reliability claims, but the false novelty claim alone undermines the stated contribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an annotated Arabic Twitter corpus of 10,000 tweets for hate speech analysis. Each tweet is labeled as offensive/clean, and offensive tweets are further labeled with one or more of seven hate speech target categories. The authors describe a pipeline using the ASAD tool to preselect tweets, a multi-annotator annotation procedure with a manager resolving conflicts, and inter-annotator agreement measured with Fleiss' kappa. They also present fine-tuning experiments with AraBERTv2, CamelBERT, and XLM-RoBERTa for both offensive/clean classification and hate speech target classification, reporting that AraBERTv2 performs best.","tokens_in":8131,"tokens_out":4596,"duration_ms":42073,"significance":"If the corpus were reliable and genuinely novel, it would be a useful resource for Arabic hate speech research, especially given the public Zenodo release, the CC BY 4.0 license, the code repository, and the inclusion of multilingual annotators from different Arab countries. The paper also provides a practical evaluation of three transformer models. However, the central claims of uniqueness and dataset reliability are undercut by internal inconsistencies: the novelty claim contradicts the paper's own related-work table, the abstract reports agreement values that the body does not support, and the dataset description contains numerical inconsistencies. These issues affect the main contribution, so the significance of the work as presented cannot be accepted without major revision or re-framing.","major_comments":[{"comment":"The claim in the Novelty section that 'Annotating each tweet to all the possible labels has been proposed for the first time in our dataset' is directly contradicted by Table 1, which lists Zaghouani, Mubarak, and Biswas 2024 with 'Multilabel: Yes' and a size of 15,965 tweets, and by the same group's LREC-COLING 2024 paper 'So hateful! Building a multi-label hate speech annotated Arabic dataset.' Because the cited prior work is in the same reference list and shares two authors, this is an internal inconsistency rather than a disagreement with outside consensus. The authors must either withdraw the first-time claim or position the new corpus as an extension or re-annotation with a clearly documented delta in size, annotation scheme, and content.","section":"Related Works and Novelty"},{"comment":"The abstract reports inter-annotator agreement of 0.86 for offensive content and 0.71 for multiple hate speech targets, but the IAA section reports a Fleiss' kappa of 0.8143 for the binary offensive/not-offensive task and explicitly states 'we didn't measure for hate speech target as we combine all possible annotation label to maximize the target group.' The 0.71 value in the abstract is therefore unsupported by any reported measurement, and the 0.86 value does not match the 0.8143 reported in the body. These numbers must be reconciled or removed.","section":"Abstract and Inter-Annotator Agreement (IAA)"},{"comment":"The counts in Table 2 do not add to the stated dataset size: 6036 + 3719 + 63 + 26 + 20 + 184 = 10,048, not 10,000. Additionally, the percentages for the small categories are off by a factor of 100 (e.g., 63/10,000 = 0.63%, not 0.006%). This undermines the reliability of the reported class distribution and must be corrected with the actual counts and denominator.","section":"Table 2"},{"comment":"The sampling procedure is described as choosing 4,000 tweets from the highest ASAD confidence band (80–100%), 4,000 from the average band (60–79%), and 2,000 from the low band (1–39%), followed by sentiment-based selection of 4,000 positive, 4,000 negative, and 2,000 neutral tweets. This is not a random sample of Arabic tweets, and the paper does not validate ASAD's confidence scores against human judgments. As a result, the offensiveness distribution in Table 2 cannot be interpreted as representative of Arabic Twitter, and models trained on this corpus may not generalize. The authors should either provide a validation of the ASAD confidence bands or explicitly describe the sampling design as purposive and discuss its limits.","section":"Data Collection"},{"comment":"The annotation procedure states that when annotators disagree on the target group, 'we combine all the target group' rather than adjudicating to a single gold label. This makes the target labels a union of annotator choices, not a resolved consensus. Combined with the explicit statement that target-level IAA was not measured, there is no evidence that the seven target categories are reliable. The target distribution in Table 3 is therefore difficult to interpret as ground truth, and the downstream target classification results rest on label definitions that have not been validated.","section":"Annotation Procedure"}],"minor_comments":[{"comment":"The caption of Figure 2 says 'Data Collection' but the figure is described in the text as a glimpse of the annotation guidelines; the caption should be corrected.","section":"Figure 2"},{"comment":"In the paragraph on Arabic hate speech datasets, the text attributes a 3,075-tweet collection to Alshaalan and Al-Khalifa, but Table 1 assigns the same size and 'Gulf countries' dialect to Alsafari, Sadaoui, and Mouhoub; this attribution mismatch should be fixed.","section":"Related Works"},{"comment":"The reference for Antoun, Baly, and Hajj has '????' as the publication year; the year should be replaced with the correct LREC 2020 date.","section":"References"},{"comment":"There are several typos and grammatical errors, including 'Methodolgy' (should be 'Methodology'), 'we didn't distuinsh the data' (should be 'distinguish'), 'between between' in Related Works, 'techiniques', 'wheather', 'techqniue', 'transforemr', and 'mahine'. A careful proofreading pass is recommended.","section":"Overall"}],"recommendation":"reject","confidential_remarks":"The novelty claim appears to overlap substantially with the authors' own LREC-COLING 2024 multilabel dataset, yet the manuscript does not disclose this overlap or spell out what is new in the present corpus. This is a novelty-disclosure concern that the editor may wish to weigh separately from the technical critique."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this paper offers a newly released 10k-tweet Arabic hate speech corpus with seven target labels, but its central novelty claim is undercut by the authors' own prior work, and the reliability reporting has inconsistencies that need fixing.\n\nWhat's actually new: the specific dataset—10,000 tweets, annotated for offensiveness and hate speech targets with multiple labels per tweet, released on Zenodo under CC BY 4.0—is new. The authors also provide annotation guidelines and code, and they run baseline transformers (AraBERT, CamelBERT, XLM-R) with reasonable results. That is a useful resource, incremental over what already exists.\n\nThe problems: the paper says \"Annotating each tweet to all the possible labels has been proposed for the first time in our dataset,\" while Table 1 lists the authors' own LREC-COLING 2024 paper (\"So hateful!\") as a multilabel dataset of 15,965 tweets. That is a direct internal contradiction. Either the 2024 dataset wasn't multilabel, or this claim is false. The authors need to position this as an extension, quantify the overlap, and document what's different.\n\nSecond, the IAA numbers don't line up. The abstract reports 0.86 for offensive content and 0.71 for multiple targets; the body says Fleiss' kappa is 0.8143 and that they didn't measure agreement on hate speech targets. These are different quantities, but the paper doesn't reconcile them.\n\nThird, Table 2's counts sum to 10,048, not 10,000, and the percentages for the minor categories are written as proportions (e.g., 63 tweets as .006 instead of 0.63%). That's the kind of error that makes one question the dataset's curation.\n\nFourth, the sampling is not random: tweets were preselected by ASAD confidence bands. That may make the corpus useful as a focused resource, but it excludes the \"clean\" tail and should be discussed as a bias, not as a representative sample.\n\nNone of this is fatal to the resource itself. The dataset and code are released and could be useful for Arabic NLP. But the paper's claims and reporting need serious revision. I would send it to peer review, because the resource deserves a proper vetting, and a good referee can help the authors fix these issues. I wouldn't cite it in its current form.","headline":"The released corpus is real, but the 'first-time multilabel' claim contradicts the authors' own 2024 dataset, and the IAA and percentage reporting are inconsistent.","tokens_in":8625,"tokens_out":3309,"would_cite":false,"duration_ms":30340,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes a 10,000-tweet Arabic hate speech corpus whose distinctive contribution is that every offensive tweet is annotated with all hate speech targets it contains, not just a single category.","keywords":["Arabic hate speech","multilabel annotation","Twitter corpus","offensive language detection","inter-annotator agreement","AraBERT","dialectal Arabic","hate speech targets"],"falsifier":"Re-annotate a random sample of the released 10,000 tweets under the published guidelines with fresh annotators and compare their offensive/clean labels and target sets with the released labels; a match far below the paper's reported agreement levels would refute the corpus's reliability. The check should also reconcile the two agreement values stated in the paper, 0.86 in the abstract and 0.8143 in the body.","tokens_in":7700,"feed_emoji":"🏷️","tokens_out":10031,"duration_ms":86924,"temperature":0.7,"pith_summary":"The paper's contribution is a resource: 10,000 Arabic tweets annotated for whether each tweet is offensive and, when it is, for every hate speech target it hits among seven categories: political ideology, origin/ethnicity/country, religion/sect, social class/profession, physical appearance/disability, gender, and other. The authors argue that this multilabel scheme, in which all target labels proposed by annotators are merged rather than forcing one label, is the first of its kind for Arabic hate speech. They report annotator agreement for the offensive/clean decision and for targets, and they show that a fine-tuned AraBERT model outperforms two other transformer baselines. If the corpus is reliable, it gives researchers a reusable benchmark for detecting not just whether an Arabic tweet is hateful but whom it targets.","feed_headline":"New dataset labels every hate speech target in 10,000 Arabic tweets","feed_subtitle":"It marks all applicable targets per tweet and sets Arabic hate speech baselines up to 0.79 micro-F1.","key_machinery":"The dataset itself is the central object: a 10,000-tweet Arabic Twitter corpus. The load-bearing design decision is multilabel annotation: instead of choosing one target per offensive tweet, the guidelines ask annotators to identify all applicable groups among seven categories, and the final label set for a tweet is the union of the labels assigned by annotators. The other key mechanism is the pre-filtering pipeline, in which offensive content is selected with ASAD's confidence scores across high, medium, and low bands and balanced by sentiment, so the final sample is intended to be diverse rather than a random draw from Twitter. Fine-tuned transformer models, primarily AraBERTv2, serve as a sanity check on label quality.","core_discovery":"The central claim is that hateful Arabic tweets often hit several groups at once, and the paper provides what it says is the first Arabic dataset designed to capture that multiplicity by instructing annotators to mark all applicable target labels and then merging disagreeing labels into one set per tweet. The construction pipeline starts from 60 million tweets collected over roughly two months, downsamples through ASAD's offensive-language confidence tiers and sentiment strata, and yields 10,000 tweets. In the resulting annotations, 60.36% of tweets are labelled hateful or offensive; among those, political ideology/sports and other targets are the most frequent (28.76% and 28.51%), followed by origin/ethnicity/country (26.20%) and religion/sect (13.51%). On the annotation-quality check, the paper reports a Fleiss' kappa of 0.8143 for offensive versus clean annotation and 0.71 for hate speech target annotation, and AraBERTv2 reaches micro-F1 scores of 0.7865 for offensive/clean and 0.6889 for target classification.","pith_inferences":["Because the sample is stratified by an automatic classifier rather than randomly sampled from Twitter, the corpus is best read as a benchmark for offensive speech that an automatic filter would flag, not as a neutral estimate of hate speech prevalence in Arabic Twitter.","The same 'mark all applicable targets and merge disagreeing labels' guideline could transfer to other low-resource languages; if adopted, it would make cross-language comparisons of hate speech targets feasible.","A testable extension would check whether target-label agreement improves when annotators share the same dialect region, since the paper motivates its annotator mix precisely by dialect-driven differences in interpretation."],"forward_implications":["A model trained on this corpus can predict multiple targets for a single hateful tweet, matching the way real posts attack more than one identity at once.","The reported baselines set a concrete reference point for Arabic multilabel hate speech detection: 0.7865 micro-F1 for offensive versus clean and 0.6889 for target classification.","The label distribution shows strong imbalance, with political/ideological and 'other' targets dominating while gender and physical-appearance targets are rare, so practical systems must handle low-frequency target classes.","Releasing the annotations and guidelines under an open license gives other researchers a documented gold standard for dialectally diverse Arabic social media content."],"supporting_citations":[{"why":"Supplies the ASAD tool whose offensive-language and sentiment confidence scores drive the tweet sampling strata.","marker":"Hassan et al. 2021"},{"why":"Prior multilabel Arabic hate speech dataset that this corpus extends; the comparison establishes the claimed novelty.","marker":"Zaghouani, Mubarak, and Biswas 2024"},{"why":"Earlier Arabic Twitter hate speech corpus with normal/abusive/hate labels that this work builds on.","marker":"Mulki et al. 2019"},{"why":"Earlier Arabic hate speech corpus and SVM/CNN/BERT baselines against which the new dataset is positioned.","marker":"Albadi, Kurdi, and Mishra 2018"},{"why":"Reports a micro-F1 of 0.7899 for Arabic hate speech detection, the closest prior baseline for the paper's transformer results.","marker":"Alsafari, Sadaoui, and Mouhoub 2020"},{"why":"Supplies the AraBERTv2 pre-trained model that achieves the best reported results in the paper.","marker":"Antoun, Baly, and Hajj"},{"why":"Supplies CamelBERT, the second transformer baseline used in the evaluation.","marker":"Inoue et al. 2021"},{"why":"Supplies XLM-RoBERTa, the third transformer baseline used in the evaluation.","marker":"Conneau 2019"}],"fun_headline_variants":["Arabic hate speech dataset marks all targets per tweet, 10k samples","New Arabic hate speech corpus tags every group hit in 10k tweets","Multi-label Arabic hate speech dataset: 10k tweets, AraBERTv2 baseline","First Arabic multi-target hate speech dataset releases baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that ASAD's automatic offensive-content confidence scores are accurate enough to pick out genuinely offensive and genuinely clean Arabic tweets; if those scores are noisy or biased by dialect, the 10,000-tweet benchmark will not be a reliable sample for training or evaluating hate speech systems.","fun_headline_variants_meta":{"raw":{"variants":["Arabic hate speech dataset marks all targets per tweet, 10k samples","New Arabic hate speech corpus tags every group hit in 10k tweets","Multi-label Arabic hate speech dataset: 10k tweets, AraBERTv2 baseline","First Arabic multi-target hate speech dataset releases baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1415,"prompt_tokens":925,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":412}},"tokens_in":541,"tokens_out":490,"duration_ms":4673,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:43:28.881530+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of the released 10,000 tweets under the published guidelines with fresh annotators and compare their offensive/clean labels and target sets with the released labels; a match far below the paper's reported agreement levels would refute the corpus's reliability. The check should also reconcile the two agreement values stated in the paper, 0.86 in the abstract and 0.8143 in the body.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier Arabic hate speech corpus and SVM/CNN/BERT baselines against which the new dataset is positioned."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ASAD tool whose offensive-language and sentiment confidence scores drive the tweet sampling strata."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier Arabic Twitter hate speech corpus with normal/abusive/hate labels that this work builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports a micro-F1 of 0.7899 for Arabic hate speech detection, the closest prior baseline for the paper's transformer results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies CamelBERT, the second transformer baseline used in the evaluation."}],"review_version":1}