REVIEW 3 major objections 4 minor 33 references
Evaluating Large Language Models for Antisemitic Incident Classification
T0 review · 3 major / 4 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Large language models can classify antisemitic incidents from reports with fine-grained labels, but still need clear definitions for rhetoric and examples for actions to work reliably.
desk verdict Solid empirical task paper with useful prompt ablations; the synthetic-negative FP claim is the softest support for "potential," but the rhetoric/action findings and resources still stand. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The fine-grained hateful-event taxonomy (coarse categories of Targeting versus Expression, plus nine non-exclusive types spanning assault, vandalism, historical tropes, bullying, etc.) together with controlled prompt variants that inject definitions or examples; these components turn raw short reports into structured, multi-label predictions that distinguish speech-like from action-like harm.
What would settle it
Re-run the binary detection experiments after replacing the synthetic negatives with a large set of real, human-written campus or local-news stories that mention Jewish life or Israel but contain no antisemitic incident; if false-positive rates jump sharply or the relative benefit of definitions versus examples disappears, the central performance claims collapse.
Extended reading notes
Core claim
Large language models, especially the stronger closed model tested, can perform fine-grained classification of antisemitic event reports at levels that already support human monitoring, but accuracy remains uneven across label types and is substantially improved by supplying term definitions for rhetoric-oriented harms and in-context examples for action-oriented harms.
Load-bearing premise
The synthetic set of non-antisemitic Jewish-related reports, generated by the same model family from a handful of seed phrases, is a valid negative class for measuring false positives.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes fine-grained hateful event detection (distinct from hate-speech detection) and evaluates GPT-4o and Llama-3.2-3B-Instruct on antisemitic incident reports drawn from AMCHA, ADL-HEAT, a GPT-4o-generated Synthetic contrast set, and a Campus-News scrape. Using controlled prompt ablations (NoCtx, Assumed, Def, Assumed-Def, Assumed-IcE), it reports that GPT-4o shows higher precision and more stable performance than Llama, that definitions most improve rhetoric-oriented fine-grained types while in-context examples most improve action-oriented types (Table 7, Figure 5), and that the best setup can surface candidate incidents from campus newspapers for human review. The authors conclude that LLMs have potential for the task yet require substantial improvement and closer collaboration with domain experts.
Significance. If the empirical patterns hold under stronger validation, the work supplies a useful new task formulation, publicly releasable resources (AMCHA subset, Synthetic seeds, Campus-News), and concrete prompt-design guidance that practitioners monitoring hate incidents could adopt immediately. The differential effect of definitions versus examples is a clean, actionable finding supported by per-type F1 tables and error examples; the campus case study, though preliminary, illustrates a realistic screening pipeline. These contributions are valuable for both NLP evaluation methodology and civil-society monitoring tools, provided the negative-class and annotation-reliability issues are addressed.
major comments (3)
- [§3.3, Table 4] §3.3 and Table 4: The Synthetic negative class is generated by GPT-4o itself from only 12 hand-crafted positive seed phrases. The reported 0 % false-positive rate for GPT-4o (versus non-zero for Llama) therefore cannot be treated as independent evidence of superior precision; stylistic and construct-relevance artifacts acknowledged by the authors make the binary-detection half of the “potential” claim circular. A non-synthetic, human-authored negative set is required before the precision comparison can support the central narrative.
- [§5.4, Table 9] §5.4 and Table 9: The human-annotated Campus-News sample contains only 19 positive articles. Precision/recall figures are consequently too noisy to underwrite the claim that the pipeline “supports early monitoring and intervention.” Either expand the annotated set substantially or relegate the case study to a purely qualitative demonstration.
- [§3.1, §5.3] §3.1 and §5.3: No inter-annotator agreement statistics are reported for the AMCHA fine-grained labels (or for the Campus-News consensus coding). Given the acknowledged subjectivity of antisemitism definitions and the mutual exclusivity of Targeting versus Expression, reliability numbers are load-bearing for interpreting the modest F1 scores and the rhetoric/action differential.
minor comments (4)
- [Figure 1] Figure 1 caption and surrounding text: the Perspective API scores are presented without version or date; given that the tool is being sunset, a brief note on the exact model used would aid reproducibility.
- [Table 1] Table 1 frequencies sum >100 % because types are multi-label; an explicit note in the caption would prevent misreading by readers unfamiliar with multi-label settings.
- [§4.1] §4.1 prompt templates: the exact Wikipedia definition string and the randomly chosen in-context examples should be listed in an appendix or repository file so that the Assumed-IcE condition is fully reproducible.
- Throughout: occasional typographic inconsistencies (e.g., “hatefulevents” missing space, “Assumed-IcE” vs “Ice”) should be cleaned for the camera-ready version.
Circularity Check
No circularity: purely empirical LLM evaluation with no derivation, fitted parameters, or load-bearing self-citation chain that reduces claims to inputs by construction.
full rationale
The paper introduces a new task (fine-grained hateful event detection), assembles/annotates datasets (AMCHA, ADL-HEAT, Synthetic, Campus-News), and reports empirical classification metrics (binary detection rates, weighted F1, per-type F1) under prompt ablations for GPT-4o and Llama-3.2-3B-Instruct. There are no equations, first-principles derivations, uniqueness theorems, or ansatzes. Performance numbers (e.g., Table 4 binary rates, Table 7/Figure 5 type F1s, Campus-News case study) are direct measurements of model outputs against expert or constructed labels; they are not obtained by fitting a free parameter on a subset and then 'predicting' a statistically forced quantity. Synthetic negatives are generated by GPT-4o from 12 hand-crafted seed phrases and used only as a negative class for false-positive measurement; the paper itself flags stylistic/construct risks and calls for further validation, so the 0% FP figure is presented with caveats rather than as a self-justifying prediction. Prompt definitions and in-context examples are external inputs supplied by the authors, not circular outputs. Self-citations (if any) are ordinary background and do not underwrite the central empirical claims. The evaluation is therefore self-contained against its own benchmarks; no step reduces by construction to its inputs.
Assumptions & free parameters
free parameters (2)
- temperature for synthetic generation =
0.9
- number of in-context examples per type =
1
assumptions (4)
- domain assumption AMCHA and ADL-HEAT gold labels are treated as ground truth for antisemitic incidents
- domain assumption Minimalist Wikipedia/IHRA definition of antisemitism plus selected AMCHA types (excluding pure anti-Israel categories) is an appropriate evaluation taxonomy
- ad hoc to paper Synthetic reports generated from 12 positive Jewish/Israeli seed phrases are valid non-antisemitic negatives
- domain assumption Prompting (zero/few-shot) is the appropriate evaluation regime rather than fine-tuning
invented entities (2)
-
hateful event detection task (fine-grained)
independent evidence
-
rhetoric-oriented vs action-oriented type grouping
Cite this review
Pith. "Pith review of Evaluating Large Language Models for Antisemitic Incident Classification." pith.science (2026). https://pith.science/paper/LAMGZ343
@misc{pith2026260704890,
author = {Pith},
title = {Pith review of: Evaluating Large Language Models for Antisemitic Incident Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/LAMGZ343}},
note = {Machine review of arXiv:2607.04890}
}
read the original abstract
Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI's GPT-4o and Meta's Llama-3.2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault). A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI's ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.
Figures
Reference graph
Works this paper leans on
-
[1]
Accessed: 2024-5-13
Quantifying Hate: A Year of Anti-Semitism on Twitter. Accessed: 2024-5-13. ADL.2024a. 46%ofAdultsWorldwideHoldSignificantAntisemiticBeliefs,ADLPollFinds. Accessed: 2024-5-13. ADL. 2024b. Antisemitic Attitudes in America
2024
-
[2]
Accessed: 2024-5-13. ADL
2024
-
[3]
Accessed: 2025-12-20
Portrait of Antisemitic Experiences in the U.S., 2024-2025. Accessed: 2025-12-20. Ali, Moonis and Savvas Zannettou
2024
-
[4]
InProceedings of the International AAAI Conference on Web and Social Media, volume 19, pages 122–145
What’s in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of 17 LLM-Generated Text Annotations. InProceedings of the International AAAI Conference on Web and Social Media, volume 19, pages 122–145. Bagavathi,Arunkumar,PedramBashiri,ShannonReid,MatthewPhillips,andSiddharthKrishnan.2019. Examining unt...
2019
-
[5]
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1664–1674, Hong Kong, China. Association for Computational Linguistics. B...
2019
-
[6]
Accessed: 2025-12-20
Victim of Boulder Firebombing Attack Dies of Wounds.The New York Times. Accessed: 2025-12-20. Chandra, Mohit, Dheeraj Pailla, Himanshu Bhatia, Aadilmehdi Sanchawala, Manish Gupta, Manish Shrivastava, and Ponnurangam Kumaraguru
2025
-
[7]
Subverting the Jewtocracy
“Subverting the Jewtocracy”: Online Antisemitism Detection Using Multimodal Deep Learning. InProceedings of the 13th ACM Web Science Conference 2021, WebSci ’21, page 148–157, New York, NY, USA. Association for Computing Machinery. Chew, Peter A
2021
-
[8]
InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363, Online and Punta Cana, Dominican Republic
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Ewing, Giselle Ruhiyyih
2021
Show all 33 references
-
[9]
Accessed: 2025-12-20
Shapiro: Nothing will ’deter me from proudly and openly practicing my faith,’ after police say alleged arsonist targeted governor over Palestine.Politico. Accessed: 2025-12-20. FBI
2025
-
[10]
Accessed: 2025-12-20
Hate Crime Data Explorer: 2024 Hate Crime Statistics. Accessed: 2025-12-20. Feldman, David and Volovici, Marc. 2023.Antisemitism, Islamophobia and the Politics of Definition. Palgrave Critical Studies of Antisemitism and Racism. Springer International Publishing. Felkner, Virg...
2024
-
[11]
InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14104–14115, Bangkok, Thailand
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14104–14115, Bangkok, Thailand. Association for Computational...
2023
-
[12]
IHRA and JDA: Examining definitions of antisemitism in 2021.Fathom. IHRA
2021
-
[13]
Accessed: 2024-5-13
Working Definition of Antisemitism. Accessed: 2024-5-13. Jasper, Clint, Danuta Kozaki, and Kimberley Price
2024
-
[14]
Accessed: 2025-12-20
Australian Jews speak out about antisemitism after Bondi Beach shooting.ABC News. Accessed: 2025-12-20. JDA
2025
-
[15]
Accessed: 2024-5-13
The Jerusalem Declaration On Antisemitism. Accessed: 2024-5-13. Jewish Federation
2024
-
[16]
https://www
Fbi data: 69% of religion-based hate crimes targeted jews. https://www. jewishfederations.org/blog/all/fbi-data-497668. Accessed: 2025-12-20. Jiang, Nan-Jiang and Marie-Catherine de Marneffe
2025
-
[17]
Can GPT-4 detect subcategories of hatred? In 2024 IEEE Digital Platforms and Societal Harms (DPSH), pages 1–6. IEEE. Mustafa, Raza Ul and Nathalie Japkowicz
2024
-
[18]
Nefriana, Rr, Muheng Yan, Rebecca Hwa, and Yu-Ru Lin
Monitoring the evolution of antisemitic discourse on extremist social media using BERT.arXiv preprint arXiv:2403.05548. Nefriana, Rr, Muheng Yan, Rebecca Hwa, and Yu-Ru Lin
-
[19]
Leader-driven or Leaderless: How Partici- pation Structure Sustains Engagement and Shapes Narratives in Online Hate Communities.arXiv preprint arXiv:2512.12441. Nexus
-
[20]
Accessed: 2024-5-3
The Nexus Project - Israel and Antisemitism.https://nexusproject.us/. Accessed: 2024-5-3. Ozalp, Sefa, Matthew L. Williams, Pete Burnap, Han Liu, and Mohamed Mostafa
2024
-
[21]
InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 35356–35385
Evaluating Large Language Models for Detecting Antisemitism. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 35356–35385. Penslar, Derek
2025
-
[22]
Accessed: 2025-12-20
Antisemitic and anti-Israeli attacks rise since October 7, 2023.Reuters. Accessed: 2025-12-20. Ron, Gal, Effi Levi, Odelia Oshri, and Shaul Shenhav
2023
-
[23]
InThe 7th Workshop on Online Abuse and Harms (WOAH), pages 215–220, Toronto, Canada
Factoring hate speech: A new annotation framework to study hate speech in social media. InThe 7th Workshop on Online Abuse and Harms (WOAH), pages 215–220, Toronto, Canada. Association for Computational Linguistics. Sap,Maarten,SaadiaGabriel,LianhuiQin,DanJurafsky,NoahA.Smith,...
2020
-
[24]
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United Sta...
2022
-
[25]
InProceedingsofthe63rdAnnualMeeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5687–5709
Making FETCH! Happen: FindingEmergentDogWhistlesThroughCommonHabitats. InProceedingsofthe63rdAnnualMeeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5687–5709. Schmidt,AnnaandMichaelWiegand.2017. ASurveyonHateSpeechDetectionusingNaturalLan...
2017
-
[26]
Jewish Museum Is Part of Global Surge in Antisemitism.The New York Times
Slaying Outside D.C. Jewish Museum Is Part of Global Surge in Antisemitism.The New York Times. Accessed: 2025-12-20. Smedt, Tom De
2025
-
[27]
Sutherland,Callum.2025
Codes, patterns and shapes of contemporary online antisemitism and conspiracy narratives – an annotation guide and labeled german-language dataset in the context of covid-19.Proceedings of the International AAAI Conference on Web and Social Media, 17(1):1082–1092. Sutherland,C...
2025
-
[28]
Too Many People Are Making Excuses.The New York Times
Antisemitism Is an Urgent Problem. Too Many People Are Making Excuses.The New York Times. Accessed: 2025-12-20. Tripodi, Rocco, Massimo Warglien, Simon Levis Sullam, and Deborah Paci
2025
-
[29]
InProceedingsofthe1stInternationalWorkshop on Computational Approaches to Historical Language Change, pages 115–125, Florence, Italy
Tracing Antisemitic Language ThroughDiachronicEmbeddingProjections: France1789-1914. InProceedingsofthe1stInternationalWorkshop on Computational Approaches to Historical Language Change, pages 115–125, Florence, Italy. Association for Computational Linguistics. U.S. Department...
1914
-
[30]
Accessed: 2024-5-13
Learn About Hate Crimes. Accessed: 2024-5-13. Vargas, Francielle, Isabelle Carvalho, Fabiana Rodrigues de Góes, Thiago Pardo, and Fabrício Benevenuto
2024
-
[31]
InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2289–2303
Introducing CAD: the Contextual Abuse Dataset. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2289–2303. Warner, William and Julia Hirschberg
2021
-
[32]
In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024), pages 1–10, Trento
Leveraging annotator disagreement for text classification. In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024), pages 1–10, Trento. Association for Computational Linguistics. Yin, Wenjie and Arkaitz Zubiaga
2024
-
[33]
Heil Hitler
A legal approach to hate speech – operationalizing the EU‘s legal framework against the expression of hatred as an NLP task. In Proceedings of the Natural Legal Language Processing Workshop 2022, pages 53–64, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computatio...
2022
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.