REVIEW 2 major objections 4 minor 80 references
Fine-grained taxonomies of antisemitism raise LLM recall but cut precision; a full 550-page lexicon adds no further gain.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Fine-grained taxonomic prompts substantially raise antisemitism-detection recall across four LLMs while lowering precision; larger conceptual resources add no further quantitative gain and post-Holocaust forms remain hardest.
T0 review reviewed 2026-07-11 challenge →
load-bearing objection Clean multi-model result: fine-grained taxonomies raise antisemitism recall (often with large effect sizes) at the cost of precision, and a 550-page lexicon adds nothing quantitative over a compact list of concepts. the 2 major comments →
You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Compact fine-grained taxonomic representations of antisemitism substantially improve detection recall across four state-of-the-art LLMs while reducing precision; providing a much larger conceptual resource (the complete 550-page Lexicon) yields no additional quantitative benefit over the compact taxonomy alone; post-Holocaust antisemitism remains the most persistent detection and classification challenge under every model and configuration.
What carries the argument
The four prompting configurations (BASE, IHRA, STRUCT, STRUCT+EX, plus LEXICON for Gemini) that inject external conceptual representations of antisemitism at inference time, together with the mapping of both annotation schemes onto four higher-level content groups that permit cross-dataset comparison of recall and explanation references.
Load-bearing premise
The authors' mapping of the Decoding annotation codes onto Lexicon chapters and IHRA sections, and the further collapse of those labels into four content groups, preserves the distinctions needed for the recall and reference-distribution claims.
What would settle it
Re-running the same model-prompt combinations after an independent re-mapping of Decoding codes that forbids multi-section assignments and re-tests whether STRUCT still doubles recall for Israel-related antisemitism and whether post-Holocaust content remains hardest.
If this is right
- Moderation systems that inject a short, fine-grained concept list will catch more antisemitic posts than those that rely on the IHRA definition alone, at the cost of more false positives.
- Adding long definitional documents or many examples is unlikely to raise F1 once the relevant concepts have already been activated.
- Detection performance will remain uneven across forms of antisemitism, with post-Holocaust and subtle justificatory cases continuing to lag.
- Explanation quality will continue to show over-production of references and surface-cue reliance even when classification is correct.
- Abstention options will stay under-used relative to the uncertainty models express in free-text explanations.
Where Pith is reading between the lines
- The same compact-taxonomy advantage may transfer to other historically layered hate categories whose surface forms are sparse relative to their conceptual range.
- If models mainly need activation cues rather than elaboration, retrieval of a short concept list may be more efficient than full-document long-context prompting for moderation.
- Cross-lingual tests on post-Holocaust narratives rooted outside English-language discourse would clarify whether the persistent difficulty is data-distributional or conceptual.
- Calibrating the Unsure option to match uncertainty language already present in explanations could reduce overconfident false positives without new training.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a controlled multi-model study of how external conceptual representations affect LLM detection and explanation of antisemitism. Using the Bloomington and Decoding expert-annotated datasets, it compares four prompting regimes (BASE, IHRA definition, fine-grained STRUCT taxonomy drawn from the Decoding Antisemitism Lexicon, and STRUCT+EX) across Gemini-2.5-Flash, Claude Sonnet 4.6, GPT-5.4 and LLaMA-3.3-70B, plus a full 550-page Lexicon condition for Gemini. Binary detection metrics (bootstrap + inverse-probability weighting) show that the compact fine-grained taxonomy substantially raises recall—often with large Cohen’s h—while lowering precision; the full Lexicon yields no further quantitative gain. Content-group analyses and explanation mapping reveal that post-Holocaust antisemitism remains hardest, that models systematically over-produce conceptual references, and that explanations frequently rely on lexical cues or accept distorted premises. The work therefore claims both the utility of structured conceptual activation and clear remaining limits of current LLM reasoning for this historically complex phenomenon.
Significance. If the empirical patterns hold, the paper supplies a timely and practically useful result for both computational antisemitism research and the broader design of knowledge-augmented moderation systems. It is the first systematic head-to-head comparison of definitional, taxonomic, example-augmented and long-context conceptual resources on the same antisemitism task, and the null result for the 550-page Lexicon is especially informative. Methodological strengths include dual expert datasets, proper resampling and multiple-testing correction, effect-size reporting, and a qualitative error analysis that surfaces concrete failure modes (sarcasm, justificatory framing, overconfidence). These contributions are concrete and falsifiable; they do not rest on untested theoretical machinery.
major comments (2)
- [§3.3 / Table 4] §3.3 and Table 4: Unsure answers are collapsed to No for all precision/recall/F1 calculations. Abstention rates are non-negligible and configuration-dependent (Gemini and GPT reach 6–24 % on Decoding negatives), and the qualitative analysis (§4.4.1) shows models frequently verbalise uncertainty while still emitting a binary label. Reporting the primary metrics both with Unsure treated as a third class and under the current collapse would make the overconfidence claim more robust.
- [§4.4 / Limitations] §4.4 and Limitations: The large-context null result rests on a single model (Gemini) and a Lexicon whose illustrative examples partially overlap the Decoding evaluation set. Although the authors correctly note strong recall gains on the independent Bloomington set and the absence of obvious positional bias, a short ablation that removes or masks the overlapping examples (or reports performance stratified by chapter provenance) would remove residual circularity concerns for the claim that “supplying substantially larger conceptual resources yields no additional quantitative benefit.”
minor comments (4)
- [Figures 1–2] Figure 1 and Figure 2: confidence intervals and arrow annotations are helpful, but the y-axis scales differ slightly across panels; a uniform scale would ease visual comparison of absolute recall levels.
- [Appendix A.2.3 / Table 3] Table 3 (content-group mapping) is clear, yet the many-to-many Decoding-to-IHRA assignments are only summarised; a short supplementary table listing the exact multi-section codes would aid reproducibility.
- [Related Work] §2.1–2.2: a few recent concurrent works on long-context RAG for hate-speech policy documents are missing; adding them would better situate the Lexicon experiment.
- [Throughout] Minor typographical inconsistencies appear in author names (Mihaljević vs. Mihaljevi´c) and in the arXiv identifier formatting; these should be standardised before camera-ready.
Circularity Check
No significant circularity: empirical multi-model comparison of prompting configurations measured against independent expert labels
full rationale
This is a controlled empirical study, not a theoretical derivation. The load-bearing claims (fine-grained STRUCT/STRUCT+EX taxonomies raise recall while lowering precision relative to BASE/IHRA; the full 550-page Lexicon yields no further quantitative gain; post-Holocaust forms remain hardest) are established by running four LLMs on two expert-annotated datasets and computing precision/recall/F1 (Table 4, Figure 1) against held-out human Yes/non-Yes labels. Those binary metrics do not depend on the authors' many-to-many code-to-chapter/section mapping or the four content-group collapse, which are used only for secondary analyses. The acknowledged overlap of some Lexicon examples with the Decoding corpus is a potential leakage concern, not a circular reduction: the same recall gains appear on the independent Bloomington set, and the paper reports the null large-context result rather than claiming a forced prediction. No self-definitional equations, fitted-parameter-as-prediction, uniqueness theorems, or ansatz-smuggling via self-citation appear. The derivation chain is simply 'prompt with resource X o measure detection metrics on expert labels'.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Expert annotations on Bloomington (IHRA sections) and Decoding (46 codes) constitute a sufficiently reliable gold standard for binary and multi-label evaluation.
- ad hoc to paper The authors’ mapping from Decoding codes to Lexicon chapters / IHRA sections and the subsequent collapse into four content groups preserve the distinctions needed for comparative recall and reference-distribution claims.
- domain assumption Speaker-intent instructions and the Unsure option are sufficient to handle quotation, sarcasm and condemnation.
invented entities (1)
-
Four higher-level content groups (Aggressive+Hate, Classic+Power, Israel-Related, Post-Holocaust)
no independent evidence
Cite this review
Pith. "Pith review of You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism." pith.science (2026). https://pith.science/paper/DKVPHD5A
@misc{pith2026260704945,
author = {Pith},
title = {Pith review of: You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism},
year = {2026},
howpublished = {\url{https://pith.science/paper/DKVPHD5A}},
note = {Machine review of arXiv:2607.04945}
}
read the original abstract
LLMs enable the integration of external conceptual resources at inference time, creating new opportunities for detecting ideologically and historically complex phenomena such as antisemitism. We investigate how different forms of conceptual grounding affect antisemitism detection and explanation behavior across four state-of-the-art LLMs. Using two expert-annotated datasets, we compare definitional, fine-grained taxonomic, example-augmented, and large-context representations of antisemitism. We find that fine-grained taxonomic representations substantially improve recall, while simultaneously reducing precision. Surprisingly, supplying substantially larger conceptual resources yields no additional quantitative benefit. Post-Holocaust antisemitism poses the most persistent challenge across models and configurations. Analysis of explanations further reveals systematic limitations including overproduction of conceptual references, reliance on lexical cues, overconfidence, and difficulties with subtle or justificatory forms of antisemitism. Our findings highlight both the potential and the remaining limitations of conceptually grounded LLMs for antisemitism detection and reasoning.
Figures
Reference graph
Works this paper leans on
-
[1]
Decoding Antisemitism: A Guide to Identifying Antisemitism Online , author =
-
[2]
Towards an Applicable Definition of Antisemitism , author=
Annotating Antisemitic Online Content. Towards an Applicable Definition of Antisemitism , author=. 2019 , eprint=
2019
-
[3]
2016 , institution =
The Working Definition of Antisemitism , howpublished =. 2016 , institution =
2016
-
[4]
Comerford, Milo and Gerster, Lea , title =. 2021 , type =. doi:10.2838/671381 , isbn =
-
[5]
, year =
Finkelstein, Joel and Paresky, Pamela and Goldenberg, Alex and Zannettou, Savvas and Jussim, Lee and Riggleman, Denver and Farmer, John and Goldenberg, Paul and Donohue, Jack and Modi, Malav H. , year =. Antisemitic
-
[6]
Israel Journal of Foreign Affairs , volume =
Porat, Dina , title =. Israel Journal of Foreign Affairs , volume =. 2011 , publisher =
2011
-
[7]
Judenhass im
Schwarz-Friesel, Monika , year =. Judenhass im
-
[8]
2025 , url =
Generating Hate: Anti-. 2025 , url =
2025
-
[9]
How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content , url =
Mihaljević, Helena and Steffen, Elisabeth , year =. How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content , url =. Proceedings of the 2nd Workshop on Computational Linguistics for Political Text Analysis (
-
[10]
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
Mendelsohn, Julia and Le Bras, Ronan and Choi, Yejin and Sap, Maarten. From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models. Proceedings of the 61st Annual Meeting of the ACL (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.845
-
[11]
Proceedings of the National Academy of Sciences , volume =
Fabrizio Gilardi and Meysam Alizadeh and Maël Kubli , title =. Proceedings of the National Academy of Sciences , volume =. 2023 , doi =
2023
-
[12]
Arviv, Eyal and Hanouna, Simo and Tsur, Oren , year =. It’s a Thin Line Between Love and Hate: Using the Echo in Modeling Dynamics of Racist Online Communities , volume =. doi:10.1609/icwsm.v15i1.18041 , journal =
-
[13]
Probing LLM s for hate speech detection: strengths and vulnerabilities
Roy, Sarthak and Harshvardhan, Ashish and Mukherjee, Animesh and Saha, Punyajoy. Probing LLM s for hate speech detection: strengths and vulnerabilities. Findings of the ACL: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.407
-
[14]
Li, Lingyao and Fan, Lizhou and Atreja, Shubham and Hemphill, Libby , year =. "HOT" ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social Media , volume =. ACM Transactions on the Web , publisher =. doi:10.1145/3643829 , number =
-
[15]
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
Melis, Matteo and Lapesa, Gabriella and Assenmacher, Dennis. A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance. Proceedings of the The 9th Workshop on Online Abuse and Harms (WOAH). 2025
2025
-
[16]
Steffen, Elisabeth and Pustet, Milena and Mihaljević, Helena , year =. 8. Algorithms Against Antisemitism?: Towards The Automated Detection of Antisemitic Content Online , url =. doi:10.11647/obp.0406.08 , series =
-
[17]
Becker, Matthias J. and Ascone, Laura and Bolton, Matthew and Bundzíková, Veronika and Chapelan, Alexis and Haupeltshofer, Pia and Krugel, Alexa and Kurjan, Iael and Mihaljević, Helena and Munnes, Stefan and Placzynta, Karolina and Pustet, Milena and Salhi, Mohamed and Scheiber, Markus and Tschiskale, Victor , title =
-
[18]
Decoding Antisemitism: An
Pustet, Milena and Mihaljević, Helena , title =. Decoding Antisemitism: An
-
[19]
Chapelan, Alexis and Ascone, Laura and Becker, Matthias J. and Bolton, Matthew and Haupeltshofer, Pia and Krasni, Jan and Krugel, Alexa and Mihaljević, Helena and Placzynta, Karolina and Pustet, Milena and Scheiber, Markus and Steffen, Elisabeth and Troschke, Hagen and Tschiskale, Victor and Vincent, Chloé , title =
-
[20]
2021 , howpublished =
Becker, Matthias J and Troschke, Hagen and Allington, Daniel , title =. 2021 , howpublished =
2021
-
[21]
Detecting Anti-Jewish Messages on Social Media
Jikeli, Gunther and Awasthi, Deepika and Axelrod, David and Miehling, Daniel and Wagh, Pauravi and Joeng, Weejoeng , urldate =. Detecting Anti-Jewish Messages on Social Media. Building an Annotated Corpus That Can Serve as A Preliminary Gold Standard , url =. Workshop Proceedings of the 15th International
-
[22]
Proceedings of the International AAAI Conference on Web and Social Media , author =
Codes, Patterns and Shapes of Contemporary Online Antisemitism and Conspiracy Narratives – an Annotation Guide and Labeled German-Language Dataset in the Context of. Proceedings of the International AAAI Conference on Web and Social Media , author =. 2023 , publisher =. doi:10.1609/icwsm.v17i1.22216 , pages =
-
[23]
Chandra, Mohit and Pailla, Dheeraj and Bhatia, Himanshu and Sanchawala, Aadilmehdi and Gupta, Manish and Shrivastava, Manish and Kumaraguru, Ponnurangam , date =. "Subverting the jewtocracy": Online antisemitism detection using multimodal deep learning , url =. Proceedings of the 13th. doi:10.1145/3447535.3462502 , series =
-
[24]
Computational and mathematical organization theory , volume=
Differences between antisemitic and non-antisemitic English language tweets , author=. Computational and mathematical organization theory , volume=. 2024 , publisher=
2024
-
[25]
doi:10.5281/zenodo.14448399 , url =
Jikeli, Gunther and Karali, Sameer and Miehling, Daniel and Soemer, Katharina , title =. doi:10.5281/zenodo.14448399 , url =
-
[26]
Pustet, Milena and Steffen, Elisabeth and Mihaljević, Helena. Detection of Conspiracy Theories Beyond Keyword Bias in G erman-Language Telegram Using Large Language Models. Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024). 2024. doi:10.18653/v1/2024.woah-1.2
-
[27]
Guo, Keyan and Hu, Alexander and Mu, Jaden and Shi, Ziheng and Zhao, Ziming and Vishwamitra, Nishant and Hu, Hongxin , month = jan, year =. An. doi:10.48550/arXiv.2401.03346 , urldate =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2401.03346
-
[28]
A Group-Specific Approach to NLP for Hate Speech Detection , url=
Karina Halevy , year=. A Group-Specific Approach to NLP for Hate Speech Detection , url=
-
[29]
2024 , publisher =
Eureka: Evaluating and Understanding Large Foundation Models , author=. 2024 , publisher =
2024
-
[30]
Evaluating Large Language Models for Detecting Antisemitism
Patel, Jay and Mehta, Hrudayangam and Blackburn, Jeremy. Evaluating Large Language Models for Detecting Antisemitism. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.1792
-
[31]
2024 , eprint=
Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media , author=. 2024 , eprint=
2024
-
[32]
The Authoritarian Personality , pages=
Prejudice in the Interview Material , author=. The Authoritarian Personality , pages=. 1950 , publisher=
1950
-
[33]
, title =
Adorno, Theodor W. , title =. 1950 , publisher =
1950
-
[34]
2015 , publisher =
The Definition of Anti-Semitism , author =. 2015 , publisher =
2015
-
[35]
The IHRA Working Definition of Antisemitism , booktitle =
Weitzman, Mark , editor =. The IHRA Working Definition of Antisemitism , booktitle =. 2019 , publisher =
2019
-
[36]
Berlin Declaration , author =
-
[37]
Who we are , author =
-
[38]
Collective Efficacy and the Role of Community Organisations in Challenging Online Hate Speech , author =
Antisemitism on Twitter. Collective Efficacy and the Role of Community Organisations in Challenging Online Hate Speech , author =. Social Media + Society , volume =. 2020 , publisher =
2020
-
[39]
Chandra, Mohit and Pailla, Dheeraj and Bhatia, Himanshu and Sanchawala, Aadilmehdi and Gupta, Manish and Shrivastava, Manish and Kumaraguru, Ponnurangam , booktitle =
-
[40]
2022 , note =
How Platforms Rate on Hate , author =. 2022 , note =
2022
-
[41]
2022 , note =
Online Antisemitisme in 2020 , author =. 2022 , note =
2020
-
[42]
Wistrich , title =
Robert S. Wistrich , title =. 1994 , address =
1994
-
[43]
M o M o E : Mixture of Moderation Experts Framework for AI -Assisted Online Governance
Goyal, Agam and Zhan, Xianyang and Chen, Yilun and Saha, Koustuv and Chandrasekharan, Eshwar. M o M o E : Mixture of Moderation Experts Framework for AI -Assisted Online Governance. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.638
-
[44]
2025 , publisher=
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author =. 2025 , publisher=
2025
-
[45]
2024 , publisher=
The. 2024 , publisher=
2024
-
[46]
2026 , institution =
2026
-
[47]
Proceedings of the ACM on Human-Computer Interaction , month = apr, articleno =
Ma, Renkai and Kou, Yubo , title =. Proceedings of the ACM on Human-Computer Interaction , month = apr, articleno =. 2023 , issue_date =. doi:10.1145/3579477 , abstract =
-
[48]
Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , articleno =
Kolla, Mahi and Salunkhe, Siddharth and Chandrasekharan, Eshwar and Saha, Koustuv , title =. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , articleno =. 2024 , publisher =. doi:10.1145/3613905.3650828 , abstract =
-
[49]
2023 , eprint=
An In-depth Look at Gemini's Language Abilities , author=. 2023 , eprint=
2023
-
[50]
Watch Your Language: Investigating Content Moderation with Large Language Models , volume =
Kumar, Deepak and AbuHashem, Yousef Anees and Durumeric, Zakir , year =. Watch Your Language: Investigating Content Moderation with Large Language Models , volume =. doi:10.1609/icwsm.v18i1.31358 , journal =
-
[51]
Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =
Hartmann, David and Oueslati, Amin and Staufer, Dimitri and Pohlmann, Lena and Munzert, Simon and Heuer, Hendrik , title =. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , articleno =. 2025 , publisher =. doi:10.1145/3706598.3713998 , abstract =
-
[52]
Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning , url=
Tang, Ruixiang and Kong, Dehan and Huang, Longtao and Xue, Hui , year=. Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning , url=. doi:10.18653/v1/2023.findings-acl.284 , booktitle=
-
[53]
Di Bonaventura, Chiara and Siciliani, Lucia and Basile, Pierpaolo and Merono Penuela, Albert and Mcgillivray, Barbara , editor =. Is. Proceedings of the. 2024 , pages =
2024
-
[54]
doi:10.48550/arXiv.2512.25015 , urldate =
Agarwal, Siddhant and Dhuler, Adya and Ruhnke, Polly and Speisman, Melvin and Akhtar, Md Shad and Yadav, Shweta , month = dec, year =. doi:10.48550/arXiv.2512.25015 , urldate =
-
[55]
Zhong, Yang and Baghel, Bhiman Kumar , month = jun, year =. Multimodal. 2024. doi:10.1109/CVPRW63382.2024.00206 , abstract =
-
[56]
Shaikh, Omar and Zhang, Hongxin and Held, William and Bernstein, Michael and Yang, Diyi , editor =. On. Proceedings of the 61st ACL (. 2023 , pages =. doi:10.18653/v1/2023.acl-long.244 , abstract =
-
[57]
Yang, Yongjin and Kim, Joonkee and Kim, Yujin and Ho, Namgyu and Thorne, James and Yun, Se-Young , editor =. Findings of the ACL:. 2023 , pages =. doi:10.18653/v1/2023.findings-emnlp.365 , abstract =
-
[58]
Huang, Fan and Kwak, Haewoon and An, Jisun , year=. Is. doi:10.1145/3543873.3587368 , booktitle=
-
[59]
, title =
Adorno, Theodor W. , title =. 1999 , publisher =
1999
-
[60]
Comprehending and Confronting Antisemitism
Monika Schwarz-Friesel , title =. Comprehending and Confronting Antisemitism. A Multi-Faceted Approach , series =. 2020 , publisher =
2020
-
[61]
The Working Definition of Antisemitism
Porat, Dina , editor =. The Working Definition of Antisemitism. A 2018 Perception , booktitle =. 2020 , publisher =
2018
-
[62]
2014 , address =
Qualitative Content Analysis: Theoretical Foundation, Basic Procedures and Software Solution , author =. 2014 , address =
2014
-
[63]
Proceedings of the International AAAI Conference on Web and Social Media , author=
What’s in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations , volume=. Proceedings of the International AAAI Conference on Web and Social Media , author=. 2025 , pages=. doi:10.1609/icwsm.v19i1.35807 , number=
-
[64]
https://www.adl.org/sites/default/files/pdfs/2025-05/j7-annual-report-on-antisemitism-2025.pdf
J7 Annual Report on Antisemitism 2025 , url = "https://www.adl.org/sites/default/files/pdfs/2025-05/j7-annual-report-on-antisemitism-2025.pdf", institution =
2025
-
[65]
Kölner Zeitschrift für Soziologie und Sozialpsychologie , year =
Bergmann, Werner and Erb, Rainer , title =. Kölner Zeitschrift für Soziologie und Sozialpsychologie , year =
-
[66]
The ADL Global 100: Index of Antisemitism , howpublished =
-
[68]
A comprehensive survey on integrating large language models with knowledge-based methods , volume=
Yang, Wenli and Some, Lilian and Bain, Michael and Kang, Byeong , year=. A comprehensive survey on integrating large language models with knowledge-based methods , volume=. doi:10.1016/j.knosys.2025.113503 , journal=
-
[69]
Wang, Mingyang and Stoll, Alisa and Lange, Lukas and Adel, Heike and Schuetze, Hinrich and Strötgen, Jannik , editor =. Bring. Proceedings of the. 2025 , pages =. doi:10.18653/v1/2025.l2m2-1.12 , urldate =
-
[70]
Large. Data Sci. Eng. , author =. 2025 , keywords =. doi:10.1007/s41019-025-00285-y , language =
-
[71]
Zhang, Min and He, Jianfeng and Ji, Taoran and Lu, Chang-Tien. Don ' t Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLM s in Implicit Hate Speech Detection. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.652
-
[72]
Markov, Ilia and Daelemans, Walter , editor =. The. Proceedings of the. 2022 , pages =
2022
-
[73]
Antisemitic Messages? A Guide to High-Quality Annotation and a Labeled Dataset of Tweets
Jikeli, Gunther and Karali, Sameer and Miehling, Daniel and Soemer, Katharina , month = apr, year =. Antisemitic. doi:10.48550/arXiv.2304.14599 , abstract =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2304.14599
-
[74]
Sociological Methods & Research , volume =
Youngjin Chae and Thomas Davidson , title =. Sociological Methods & Research , volume =
-
[75]
Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =
Reynolds, Laria and McDonell, Kyle , title =. Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =. 2021 , isbn =
2021
-
[76]
Can Prompting LLM s Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
Ghorbanpour, Faeze and Dementieva, Daryna and Fraser, Alexander. Can Prompting LLM s Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study. Proceedings of the The 9th Workshop on Online Abuse and Harms (WOAH). 2025
2025
-
[77]
2025 , eprint=
LLM Inference Enhanced by External Knowledge: A Survey , author=. 2025 , eprint=
2025
-
[78]
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
ElSherief, Mai and Ziems, Caleb and Muchlinski, David and Anupindi, Vaishnavi and Seybolt, Jordyn and De Choudhury, Munmun and Yang, Diyi. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.29
-
[79]
An In-depth Analysis of Implicit and Subtle Hate Speech Messages
Ocampo, Nicol \'a s Benjam \'i n and Sviridova, Ekaterina and Cabrio, Elena and Villata, Serena. An In-depth Analysis of Implicit and Subtle Hate Speech Messages. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. doi:10.18653/v1/2023.eacl-main.147
-
[80]
Do LLM s Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
Madhusudhan, Nishanth and Madhusudhan, Sathwik Tejaswi and Yadav, Vikas and Hashemi, Masoud. Do LLM s Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models. Proceedings of the 31st International Conference on Computational Linguistics. 2025
2025
-
[81]
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
Wen, Bingbing and Howe, Bill and Wang, Lucy Lu. Characterizing LLM Abstention Behavior in Science QA with Context Perturbations. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.197
This paper was first reviewed by grok-4.5 on July 11, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.