{"id":"368e4b17-f759-44b8-b263-f48732279fe5","arxiv_id":"2412.04046","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new publicly available dataset of 3,320 tweets from 2020 to 2022 labels hostility toward UK MPs and the targeted identity type, with benchmarks for BERT, RoBERTa, LLaMA, and GPT models.","lead":"Researchers created a manually annotated dataset of 3,320 tweets aimed at UK Members of Parliament, labeling whether each tweet is hostile and, if so, whether it targets race, gender, religion, or none. They also tested several AI models on the detection task and analyzed topics and language patterns in the abuse.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Descriptive hostility comparisons rely on a stratified convenience sample that cannot support population-level claims about party and identity differences.","rationale":"The reader's weakest assumption identifies the same sample representativeness issue that I consider most load-bearing. The paper's descriptive findings about hostility by party and identity are presented as substantive insights, yet they rest on a convenience sample of 18 hand-selected MPs and their five highest-activity days, with fixed quotas of hostile and non-hostile tweets. This design cannot support prevalence claims. The dataset itself, however, remains a useful resource with documented annotation and reproducible baselines, so the reader's conditional verdict is appropriate. My concern does not change that verdict; it reinforces the need for the authors to either reframe the descriptive claims as sample-specific or validate them on a representative sample.","tokens_in":17021,"tokens_out":7955,"duration_ms":82436,"concrete_test":"Obtain from the full collection C the total number of tweets per MP-day and the Gorrell-classifier hostility score for every tweet. Recompute Figures 2–3 using inverse-probability weights: each annotated tweet is weighted by the inverse of its sampling probability (17/N_h for classifier-hostile, 20/N_nh for classifier-nonhostile, within each MP-day). If the party/identity differences in hostility rates and identity-mix proportions persist with confidence intervals excluding zero, the sampling concern is mitigated; if they shrink, flip, or become non-significant, the descriptive claims should be withdrawn or explicitly restated as sample-specific. As a second check, sample the same number of tweets from random non-peak days for the same MPs and compare the weighted hostility rates; large differences would confirm that peak-day selection distorts the prevalence estimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central dataset contribution is credible, but one of its claimed insights — that certain parties and identity groups receive more hostility (e.g., 'Conservative Party receive more race-based hostility', 'non-white and non-Christian MPs face significantly higher levels', Figures 2–3) — is not supported by the sampling design. In the Data Sampling section, 18 MPs are hand-selected, and for each MP only the 5 highest-posting-activity days are used; on each such day, exactly 17 classifier-positive and 20 classifier-negative tweets are sampled. This fixes the number of hostile tweets per MP-day, so raw counts in Figures 2 and 3 are artifacts of the sampling quotas rather than estimates of hostility rates. Moreover, peak-activity days are likely to cluster around controversies, so they do not represent typical abuse; the paper's claim that this ensures a 'long temporal span' is not guaranteed. Finally, the initial hostility screening uses the Gorrell et al. (2020) classifier; if its false-positive/false-negative rates vary by party or identity, the annotated set inherits that bias. These descriptive claims are a stated contribution, so the concern is load-bearing for the paper's conclusions, though not for the dataset's reuse as a benchmark.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a manually annotated dataset of 3,320 English tweets directed at UK MPs, collected between November 2020 and December 2022. Tweets are labeled for hostility (binary) and, when hostile, for targeted identity characteristics (race, gender, religion, none), with per-annotator confidence scores; three gold-label sets are derived. The authors describe their sampling and annotation protocol, present linguistic (BOW, LIWC) and topic (BERTopic) analyses, and benchmark several pretrained language models and large language models on binary hostility detection and multi-class identity classification in flat and hierarchical settings. The dataset is released publicly on Zenodo.","tokens_in":17224,"tokens_out":6039,"duration_ms":60987,"significance":"If the dataset is used as a benchmark for UK-specific political hostility detection, it fills a genuine gap: existing UK political abuse datasets either lack hostility labels or cover only Islamophobia. The annotation pipeline is thorough (48 annotators, three per tweet, training and testing, hostility Fleiss kappa 0.68-0.79), and the public release with confidence scores and multiple gold sets is a useful resource. The model evaluation provides reproducible baselines. The descriptive claims about party- and identity-based differences are not, however, supported by the sampling design as currently presented.","major_comments":[{"comment":"The sampling design fixes, for each of the 18 hand-selected MPs, exactly 17 classifier-positive and 20 classifier-negative tweets per MP on each of the 5 highest-posting-activity days. Consequently, the raw counts in Figures 2 and 3 and the statements in the Dataset section ('MPs belonging to the Conservative Party receive more race-based hostility'; 'non-white and non-Christian MPs face significantly higher levels') are not estimates of hostility rates in the population of UK MPs: the quota on hostile tweets makes the number of hostile tweets per MP-day an artifact of sampling, and the hand-picked MPs and peak-activity days are not a probability sample. The authors should either analyze the full collection or a proper random sample with appropriate denominators, or explicitly restrict these claims to the annotated sample and support them with appropriate statistical tests.","section":"Data Sampling; Dataset (Figures 2-3)"},{"comment":"The initial screening uses the Gorrell et al. (2020) abusive-language classifier to select the 17 'hostile' tweets per MP-day. The manual annotations therefore apply to a set that is conditional on this classifier's decisions. If the classifier's error rates vary by party or identity group, the annotated dataset and every descriptive statistic derived from it inherit that selection bias. The manuscript does not report any validation of this screening classifier on the 2020-2022 period or any analysis of its error distribution across parties or identities. Please quantify the screening precision and recall on a sample and discuss the implications, or weaken the descriptive claims accordingly.","section":"Data Sampling"},{"comment":"The claim that sampling the '5 different highest posting activity days for each MP' ensures 'a long temporal span' is not substantiated. Peak-activity days are likely to cluster around discrete political events such as elections, scandals, or policy announcements, and the authors do not report the actual date distribution of the sampled tweets over November 2020 to December 2022. Without this information, the conclusion that the two-year span provides a broader range of topics and improves generalizability is not supported. Please report the temporal distribution of the sample and, if necessary, adjust the sampling design or the conclusions.","section":"Data Sampling"}],"minor_comments":[{"comment":"Please specify how many annotations were corrected by experts in the race/religion confusion cases, and whether the correction was applied before or after computing agreement statistics.","section":"Data Annotation"},{"comment":"The value 'moral 1.51' lacks a leading zero and is inconsistent with the other correlations; it should presumably be 0.151.","section":"Table 5"},{"comment":"The example tweets are numbered 'Tweet 7' twice; renumber the examples to avoid ambiguity.","section":"Data Characterisation"},{"comment":"The exact prompts used for LLaMA and GPT are not given; include the full prompt templates in the appendix for reproducibility.","section":"Experimental Set-up"},{"comment":"The BERT citation is given as 'Kenton and Toutanova 2019'; the standard citation is Devlin et al. (2019).","section":"References"},{"comment":"The availability section mentions both an anonymous review URL and a Zenodo record; the final version should state one canonical URL for the dataset.","section":"Dataset Availability"}],"recommendation":"major_revision","confidential_remarks":"The dataset and benchmark evaluation are solid and constitute a useful contribution. The main issue is the gap between the sampling design and the paper's descriptive conclusions about party- and identity-based differences in hostility. If the authors reframe those findings as sample-specific and add the requested temporal-distribution and screening-validation analyses, I would support acceptance. The concern should not be read as fatal to the resource itself, which remains reusable as a benchmark regardless of the descriptive limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is the real contribution here: 3,320 tweets spanning two years, manually annotated for hostility and targeted identity (race, gender, religion, none), with confidence scores and a few intersectional labels. It fills a concrete gap—existing UK political datasets either lack identity labels or focus on a single type of abuse. The annotation pipeline is careful: 48 annotators, 3 per tweet, training and testing, and Fleiss' kappa of 0.68–0.79 for hostility, which is creditable for this task. Releasing it on Zenodo with tweet IDs makes it reproducible and immediately useful.\n\nThe soft spots are real but not fatal. The sampling design is the main issue. The authors hand-pick 18 MPs, then take only the 5 highest-posting-activity days per MP, and on each day sample a fixed quota of 17 classifier-positive and 20 classifier-negative tweets. That means Figures 2 and 3 describe the composition of this convenience sample, not the population of hostility towards UK MPs. The claim that 'the Conservative Party receive more race-based hostility' or that 'non-white and non-Christian MPs face significantly higher levels' goes beyond what the data can support, because the MP selection is not representative and the activity-day selection likely clusters around controversies. The authors present these as general findings. They should either reframe them as sample-specific observations or validate on a representative sample. The 'long temporal span' point is also weaker than claimed: five peak days per MP do not guarantee coverage of typical abuse over two years.\n\nThe small religion category (36–52 examples) makes the multi-class results noisy, and the authors might have done per-class error analysis rather than only macro-F1. The initial screening by the Gorrell classifier could in principle bias which tweets end up in the set, but the manual labels are independent, so the benchmark value of the dataset is intact. This is not a circularity problem.\n\nWho benefits: computational social scientists studying online abuse of politicians, and NLP researchers needing a UK-specific benchmark. It deserves a serious referee, but the authors should be pressed to temper the descriptive conclusions. I'd accept it with revisions rather than desk reject.","headline":"A solid, much-needed dataset for UK political hostility with careful annotation, but the headline descriptive claims about party and identity differences are not supported by the sampling design.","tokens_in":603,"tokens_out":627,"would_cite":true,"duration_ms":27796,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new dataset maps two years of hostile tweets aimed at UK MPs and tests which models can spot them.","keywords":["UK politics","hostility detection","MP abuse","Twitter dataset","identity-based hate","annotation","large language models","hate speech"],"falsifier":"Re-run the analysis on a random sample of tweets from all 568 MPs with active accounts across the full two-year period, instead of the 18 selected MPs and their five highest-activity days. If the Conservative Party no longer shows a higher rate of race-based hostility than Labour, the paper's headline descriptive claim fails. A second check: re-annotate the same tweets with a fresh pool of annotators and recompute Fleiss' kappa; if identity-label agreement falls below substantial (0.4), the identity labels do not support the comparative findings.","tokens_in":16788,"feed_emoji":"🗳️","tokens_out":5244,"duration_ms":46204,"temperature":0.7,"pith_summary":"This paper constructs a publicly available dataset of 3,320 tweets directed at UK MPs over two years, each manually annotated for hostility and, when hostile, for the identity targeted: race, gender, religion, or none. The authors argue this fills a gap, because existing UK political hostility datasets either lack manual hostility labels, cover only a short period, or focus on a single identity type. They show that hostility tracks contemporaneous issues such as Brexit, illegal immigration, and the cost-of-living crisis, that Conservative MPs in their sample receive more race-based hostility, and that MPs from racial and religious minorities receive more identity-based hostility. They then benchmark five models, finding that a domain-adapted RoBERTa reaches the best macro F1 (73.03) on binary hostility detection, while GPT-3.5 with in-prompt definitions leads identity classification in a hierarchical setup (55.98 macro F1).","feed_headline":"Two years of hostile tweets to UK MPs become a dataset","feed_subtitle":"3,320 annotated tweets reveal race-based hostility as the top target, plus model benchmarks.","key_machinery":"The central object is the annotation taxonomy and dataset construction pipeline. An umbrella 'hostile' definition merges hate, abuse, and toxicity; a hierarchical task structure first asks hostile or not, then religion, gender, race, or none; and each label carries a 1–5 confidence score. Three gold-label sets are derived: majority vote (Set 1), confidence-filtered (Set 2), and intersectionality-preserving (Set 3). The sampling pipeline selects 18 MPs balanced by party and identity, takes their five highest-posting-activity days, and oversamples tweets flagged by an abusive-language classifier, producing 3,320 tweets. This machinery is what lets the paper claim the dataset is identity-aware, temporally diverse, and quality-controlled.","core_discovery":"The central claim is that the two-year, identity-labelled dataset is a step-change resource for UK political hostility detection, enabling both descriptive analysis and model training that prior resources could not support. In the paper's own framing, it 'bridges the gap' left by datasets that lack hostility labels, cover short windows, or address only Islamophobia. The authors demonstrate the resource's value with three gold-label sets, a confidence-aware annotation scheme, and a comparison of fine-tuned transformers (BERT, RoBERTa, RoBERTa-Hate) and zero-shot LLMs (LLaMA-3-8B, GPT-3.5) on binary hostility identification and multi-class identity classification.","pith_inferences":["The 18-MP sample, skewed to non-white and female MPs, means the paper's descriptive statistics describe the sample, not the population; a population-weighted sample could confirm or overturn the Conservative-party finding.","The taxonomy's umbrella 'hostile' definition could transfer to other countries, but the paper's own Brexit and immigration examples suggest models will need country-specific retraining.","Annotator confusion between race and religion labels for Muslim and Jewish targets implies downstream models may inherit that confusion; expert post-correction was needed, so automated systems should expect ambiguity at that boundary.","The dataset's intersectional labels are few (43), so claims about intersectional hostility are suggestive, not statistically powerful."],"forward_implications":["The released dataset gives researchers a UK-specific, identity-labeled resource for training and evaluating hostility detectors.","Confidence-filtered labels (Set 2) improve model performance, so future annotation efforts should record and use per-annotation confidence.","Because hostility tracks contemporaneous issues, classifiers trained on this two-year period will likely need periodic updating as new issues emerge.","Adding label definitions to LLM prompts yields large gains in identity classification, suggesting prompt design is a key lever for zero-shot abuse detection.","Race-based hostility is the most common identity category in the sample, with illegal immigration discussions carrying twice the hostility of non-hostile tweets on that topic."],"supporting_citations":[{"why":"Supplies the abusive-language classifier used to oversample hostile tweets for manual annotation.","marker":"Gorrell et al. 2020"},{"why":"Provides the Streaming API collection method used to gather over 30 million tweets from all MP accounts.","marker":"Bakir, Farrell, and Bontcheva 2024"},{"why":"Existing UK political dataset with abuse labels but no identity characteristics, motivating the new resource.","marker":"Ward and McLoughlin 2020"},{"why":"Existing UK political dataset focused only on Islamophobia, which the new dataset broadens to multiple identities.","marker":"Vidgen and Yasseri 2020"},{"why":"Large UK political dataset with automatically generated hate labels but no manual identity annotations, used as a comparison baseline.","marker":"Agarwal et al. 2021"},{"why":"Justifies the two-year collection span as improving temporal generalisability of classifiers trained on the data.","marker":"Jin et al. 2023"},{"why":"Supplies the hierarchical classification structure (hostile then identity) analogous to OffensEval.","marker":"Zampieri et al. 2019"}],"fun_headline_variants":["3,320 hostile tweets to UK MPs now a public dataset","Identity-targeted abuse of UK MPs: new dataset maps it","UK MPs face race-targeted abuse: 3,320-tweet dataset","New dataset labels racist, sexist abuse aimed at UK MPs","3,320 tweets reveal UK MPs are targeted by race-based hate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that 18 hand-selected MPs and their five highest-posting-activity days represent hostility towards UK MPs as a whole; if those MPs or days are unrepresentative, the aggregate party and identity comparisons do not generalize.","fun_headline_variants_meta":{"raw":{"variants":["3,320 hostile tweets to UK MPs now a public dataset","Identity-targeted abuse of UK MPs: new dataset maps it","UK MPs face race-targeted abuse: 3,320-tweet dataset","New dataset labels racist, sexist abuse aimed at UK MPs","3,320 tweets reveal UK MPs are targeted by race-based hate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001277,"raw_usage":{"total_tokens":5205,"prompt_tokens":914,"completion_tokens":4291,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":4201}},"tokens_in":530,"tokens_out":4291,"duration_ms":31653,"temperature":1.0,"reasoning_tokens":4201,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:49:00.732117+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis on a random sample of tweets from all 568 MPs with active accounts across the full two-year period, instead of the 18 selected MPs and their five highest-activity days. If the Conservative Party no longer shows a higher rate of race-based hostility than Labour, the paper's headline descriptive claim fails. A second check: re-annotate the same tweets with a fresh pool of annotators and recompute Fleiss' kappa; if identity-label agreement falls below substantial (0.4), the identity labels do not support the comparative findings.","supporting_citations":[{"cited_title":"E.; Roberts, I.; Greenwood, M","cited_arxiv_id":null,"evidence_quote":"Supplies the abusive-language classifier used to oversample hostile tweets for manual annotation."},{"cited_title":"E.; Farrell, T.; and Bontcheva, K","cited_arxiv_id":null,"evidence_quote":"Provides the Streaming API collection method used to gather over 30 million tweets from all MP accounts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Existing UK political dataset with abuse labels but no identity characteristics, motivating the new resource."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Existing UK political dataset focused only on Islamophobia, which the new dataset broadens to multiple identities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large UK political dataset with automatically generated hate labels but no manual identity annotations, used as a comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the hierarchical classification structure (hostile then identity) analogous to OffensEval."}],"review_version":1}