Pith. sign in

REVIEW 2 major objections 16 references

Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech

T0 review · 2 major / 0 minor · reviewed 2026-05-22 · grok-4.3

Pith's one-line read LLMs recognize semi-coded Islamophobic slurs that standard systems miss

desk verdict The abstract claims LLMs understand specific OOV slurs but gives no prompts, metrics, samples, or validation, so the main result cannot be assessed. read the letter →

arxiv 2503.18273 v3 submitted 2025-03-24 cs.LG

classification cs.LG
keywords Islamophobiahatespeechdetectionlargelanguagemodelssocialmediaanalysissemi-codedtoxicityscoringtopicmodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines semi-coded Islamophobic terms such as muzrat, pislam, mudslime, mohammedan, and muzzies that circulate on platforms including 4Chan, Gab, and Telegram. It demonstrates that large language models can interpret these out-of-vocabulary slurs in context. Toxicity analysis indicates these posts score higher than other hate speech categories, and topic modeling shows the discourse spans political and far-right movements while targeting Muslim immigrants. Readers would care because this suggests LLMs could augment detection of language that appears neutral or ambiguous.

What carries the argument

Large language models applied to out-of-vocabulary slurs, Google Perspective API for toxicity scoring, and BERT topic modeling for discourse extraction

What would settle it

A comparison study where human annotators classify the same posts for hate speech and show low agreement with LLM classifications would challenge the claim that LLMs understand the slurs.

Watch

Extended reading notes

Core claim

LLMs understand these Out-Of-Vocabulary slurs; Islamophobic posts receive higher toxicity scores than Antisemitism; topic modeling extracts various topics showing discourse in political, conspiratorial, and far-right movements particularly directed against Muslim immigrants. Further improvements in moderation strategies and algorithmic detection are necessary.

Load-bearing premise

The listed terms function as Islamophobic slurs in the sampled contexts and LLM outputs constitute reliable evidence of understanding without validation against human judgments.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript claims to perform a large-scale analysis of semi-coded Islamophobic terms such as (muzrat, pislam, mudslime, mohammedan, muzzies) on platforms including 4Chan, Gab, and Telegram. It uses LLMs to demonstrate understanding of these OOV terms, Google Perspective API to show higher toxicity scores than other hate speech categories such as Antisemitism, and BERT topic modeling to extract topics spanning political, conspiratorial, and far-right movements particularly directed against Muslim immigrants. The conclusion states that LLMs understand these slurs but further improvements in moderation are needed.

Significance. The topic of detecting semi-coded hate speech is relevant to computational social science and content moderation. However, because the manuscript supplies no data, methods, sample sizes, prompts, metrics, or results, it is not possible to assess whether any contribution would hold or advance the field.

major comments (2)
  1. [Abstract] Abstract: The claims that LLMs understand the listed OOV slurs, that Islamophobic posts receive higher toxicity scores, and that topic modeling reveals specific discourse patterns are asserted without any reported prompts, evaluation criteria, quantitative metrics (e.g., accuracy, F1), sample sizes, or statistical results, so the central empirical findings lack visible support.
  2. [Abstract] Abstract: The premise that the listed terms function as Islamophobic slurs in the sampled contexts is taken as given without any annotation details, contextual examples, or validation against human judgments, which is load-bearing for the analysis of semi-coded hate speech.

Simulated Author's Rebuttal

2 responses · 1 unresolved

We thank the referee for the review and the emphasis on empirical transparency. The comments correctly identify that the provided abstract asserts findings without accompanying details. We address each point below. As only the abstract is available, our ability to supply the requested specifics is limited.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The claims that LLMs understand the listed OOV slurs, that Islamophobic posts receive higher toxicity scores, and that topic modeling reveals specific discourse patterns are asserted without any reported prompts, evaluation criteria, quantitative metrics (e.g., accuracy, F1), sample sizes, or statistical results, so the central empirical findings lack visible support.

    Authors: We agree that the abstract presents the claims without the supporting methodological details, metrics, or sample sizes. Abstracts are by nature concise, but the absence of any reference to evaluation criteria or results does leave the findings without visible support in the provided text. We will revise the abstract to incorporate a high-level statement of the evaluation approach and key quantitative outcomes where space permits. revision: yes

  2. Referee: [Abstract] Abstract: The premise that the listed terms function as Islamophobic slurs in the sampled contexts is taken as given without any annotation details, contextual examples, or validation against human judgments, which is load-bearing for the analysis of semi-coded hate speech.

    Authors: We agree that the abstract assumes the listed terms are Islamophobic slurs without supplying annotation procedures, examples, or human validation. This premise is indeed central, and its lack of support in the abstract is a valid concern. We will revise the abstract to include a short statement on the basis for identifying these terms as semi-coded hate speech. revision: yes

standing simulated objections not resolved
  • Absence of data, methods, sample sizes, prompts, metrics, results, annotation details, contextual examples, and human validation in the manuscript as provided (limited to the abstract).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; purely descriptive empirical study with no derivations

full rationale

The paper is a descriptive empirical analysis that applies off-the-shelf LLMs, Google Perspective API, and BERT topic modeling to a corpus of social media posts. No equations, parameter fitting, predictions derived from inputs, self-citations, or uniqueness theorems appear in the provided abstract or description. The central claim that LLMs 'understand' the listed terms is presented as an observation from model outputs rather than a derived result that reduces to the inputs by construction. This is the normal case of a non-circular empirical paper.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is an applied empirical study with no free parameters, axioms, or new invented entities described in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech." pith.science (2026). https://pith.science/paper/2503.18273

@misc{pith2026250318273,
  author       = {Pith},
  title        = {Pith review of: Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2503.18273}},
  note         = {Machine review of arXiv:2503.18273}
}
read the original abstract

In recent years, Islamophobia has gained significant traction across Western societies, fueled by the rise of digital communication networks. This paper performs a large-scale analysis of specialized, semi-coded Islamophobic terms such as (muzrat, pislam, mudslime, mohammedan, muzzies) floated on extremist social platforms, i.e., 4Chan, Gab, Telegram, etc. Many of these terms appear lexically neutral or ambiguous outside of specific contexts, making them difficult for both human moderators and automated systems to reliably identify as hate speech. First, we use Large Language Models (LLMs) to show their ability to understand these terms. Second, Google Perspective API suggests that Islamophobic posts tend to receive higher toxicity scores than other categories of hate speech like Antisemitism. Finally, we use BERT topic modeling approach to extract different topics and Islamophobic discourse on these social platforms. Our findings indicate that LLMs understand these Out-Of-Vocabulary (OOV) slurs; however, further improvements in moderation strategies and algorithmic detection are necessary to address such discourse effectively. Our topic modeling also indicates that Islamophobic text is found across various political, conspiratorial, and far-right movements and is particularly directed against Muslim immigrants. Taken altogether, we performed one of the first studies on Islamophobic semi-coded terms and shed a global light on Islamophobia.

Figures

Figures reproduced from arXiv: 2503.18273 by the authors.

Figure 1
Figure 1. Toxicity comparison of Islamophobia and Antisemitism. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. 100 most occurring words in Antisemitism. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 2
Figure 2. 100 most occurring words in Islamophobia. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

16 extracted references · 16 canonical work pages

  1. [1]

    The language of islamophobia in internet articles,

    H. Mohideen and S. Mohideen, “The language of islamophobia in internet articles,” Intellectual Discourse , vol. 16, no. 1, 2008

  2. [2]

    Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,

    Y. Qu, X. Shen, X. He, M. Backes, S. Zannettou, and Y. Zhang, “Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,” in Proceedings of the 2023 ACM SIGSAC conference on computer and communications security , pp. 3403–3417, 2023

  3. [3]

    How developments in natural language processing help us in understanding human behaviour,

    R. Mihalcea, L. Biester, R. L. Boyd, Z. Jin, V . Perez-Rosas, S. Wilson, and J. W. Pennebaker, “How developments in natural language processing help us in understanding human behaviour,” Nature Human Behaviour , vol. 8, no. 10, pp. 1877–1889, 2024

  4. [4]

    Howard, Freedom of expression and religious hate speech in Europe

    E. Howard, Freedom of expression and religious hate speech in Europe. Routledge, 2017

  5. [5]

    Hate speech detection and reclaimed language: Mitigating false positives and compounded discrimination,

    E. Zsisku, A. Zubiaga, and H. Dubossarsky, “Hate speech detection and reclaimed language: Mitigating false positives and compounded discrimination,” in Proceedings of the 16th ACM Web Science Conference, pp. 241–249, 2024

  6. [6]

    Hate speech on twitter: A pragmatic approach to collect hateful and offensive expressions and perform hate speech detection,

    H. Watanabe, M. Bouazizi, and T. Ohtsuki, “Hate speech on twitter: A pragmatic approach to collect hateful and offensive expressions and perform hate speech detection,” IEEE access, vol. 6, pp. 13825– 13835, 2018

  7. [7]

    Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decision making,

    P. Burnap and M. L. Williams, “Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decision making,” Policy & internet , vol. 7, no. 2, pp. 223–242, 2015. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 8

  8. [8]

    Is- lamophobia content detection using natural language processing,

    A. Jaleel, M. Anwar, F. Ali, R. Mukhtar, and M. Farooq, “Is- lamophobia content detection using natural language processing,” Journal of Computing & Biomedical Informatics , vol. 4, no. 02, pp. 88–97, 2023

Show all 16 references
  1. [9]

    Enhancing automated hate speech detection: Addressing islamophobia and freedom of speech in online discussions,

    E. Aldreabi and J. Blackburn, “Enhancing automated hate speech detection: Addressing islamophobia and freedom of speech in online discussions,” in Proceedings of the International Conference on Advances in Social Networks Analysis and Mining , pp. 644–651, 2023

  2. [10]

    ‘is this a hate speech?’the difficulty in combating radicalisation in coded communications on social media platforms,

    B. Farrand, “‘is this a hate speech?’the difficulty in combating radicalisation in coded communications on social media platforms,” European Journal on Criminal Policy and Research , vol. 29, no. 3, pp. 477–493, 2023

  3. [11]

    Challenges of hate speech detection in social media: Data scarcity, and leveraging external resources,

    G. Kov ´acs, P. Alonso, and R. Saini, “Challenges of hate speech detection in social media: Data scarcity, and leveraging external resources,” SN Computer Science , vol. 2, no. 2, p. 95, 2021

  4. [12]

    Hate speech detec- tion and racial bias mitigation in social media based on bert model,

    M. Mozafari, R. Farahbakhsh, and N. Crespi, “Hate speech detec- tion and racial bias mitigation in social media based on bert model,” PloS one , vol. 15, no. 8, p. e0237861, 2020

  5. [13]

    Monitoring the evolution of antisemitic hate speech on extremist social media,

    R. U. Mustafa and N. Japkowicz, “Monitoring the evolution of antisemitic hate speech on extremist social media,” in 2024 IEEE Digital Platforms and Societal Harms (DPSH) , pp. 1–8, IEEE, 2024

  6. [14]

    Coded term discovery for online hate speech detection,

    D. Kikkisetti, R. Mustafa, W. Melillo, R. Corizzo, Z. Boukouvalas, J. Gill, and N. Japkowicz, “Coded term discovery for online hate speech detection,” in 2024 IEEE 11th International Conference on Data Science and Advanced Analytics (DSAA) , pp. 1–10, IEEE, 2024

  7. [15]

    Perspective api,

    G. Jigsaw, “Perspective api,” 2025. Accessed: March 16, 2025

  8. [16]

    Undocumented aapi: Profiles of those ineligible for immigration relief,

    D. Millet, “Undocumented aapi: Profiles of those ineligible for immigration relief,” 2022. Accessed: March 6, 2025

Pith tools

Reviewed May 22, 2026 · model on record in the stance chip above.