REVIEW 3 major objections 5 minor 72 references
Fine-tuning general-purpose LLM embeddings alone reaches state-of-the-art implicit hate speech detection.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Fine-tuning large general-purpose text embeddings with a simple instruction yields state-of-the-art implicit hate speech detection, with up to 20.35 point cross-dataset F1 gains.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Fine-tuning general-purpose LLM embeddings is a simple, reproducible recipe for implicit hate speech, and the cross-dataset scaling finding is real; the claimed SOTA hinges on unverified baseline comparability. the 3 major comments →
Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is that specializing generalist LLM embedding models—Stella, Jasper, NV-Embed, and E5—by fine-tuning them with a fixed instruction template ("Instruct: classify the following in no hate or hate.\nQuery: ...") and a two-layer MLP head outperforms task-specific implicit-hate-speech pipelines. Adding context and emotion modules to BERT-based classifiers produces only marginal gains, whereas the fine-tuned embeddings reach new state-of-the-art results on IHC and strong results across datasets. The largest claimed improvement is a 20.35 percentage-point F1-macro gain on ToxiGen after fine-tuning on IHC, and the pattern across models is that larger embedders generalize better
What carries the argument
The load-bearing mechanism is instruction-templated fine-tuning of generalist LLM embedding models. Each model receives the fixed instruction prepended to the tweet, produces token-level embeddings that are combined (normalized sum over tokens, or mean pooling for NV-Embed), and the resulting vector is classified by a two-layer MLP with LeakyReLU. NV-Embed is fine-tuned with LoRA; the other models are fully fine-tuned. This replaces the external-knowledge, emotion, and contrastive-sampling components of prior methods with the embeddings' own pretrained semantic structure.
Load-bearing premise
The claimed state-of-the-art rests on the assumption that the cited baselines used the same train/test splits and the same hate/not-hate label definitions as this paper; if they did not, the percentage-point gaps are not directly comparable.
What would settle it
Rerun LAHN and ConPrompt under this paper's exact splits (60/20/20 for IHC and DynaHate, 80/10/10 for SBIC, 70/10/20 for ToxiGen) and label binarization; if their F1-macro scores match or exceed the fine-tuned embeddings, the claimed state-of-the-art and the 20.35-point gain disappear.
If this is right
- External-knowledge and emotion-fusion modules can be dropped without sacrificing accuracy; content-only fine-tuned embeddings match or beat them.
- Cross-dataset generalization scales with model size: the 7-billion-parameter NV-Embed gives the largest transfer gains.
- Linear probing of a large embedder can rival full fine-tuning of a smaller one, making deployment a cost-performance decision rather than a pure accuracy decision.
- Prior contrastive and prompt-based methods become the comparison points rather than the ceiling, since a generic recipe now exceeds them.
- Instruction-tuned generative LLMs used as encoders remain below fine-tuned embedding models, so classification-specific embeddings are the more promising direction for this task.
Where Pith is reading between the lines
- The reported percentage-point gaps may overstate the true improvement if prior baselines used different train/test splits or different hate/not-hate label binarization; rerunning LAHN and ConPrompt under this paper's exact protocol would settle that question.
- The scaling trend suggests that even larger or domain-adapted generalist embedders would push cross-dataset scores further—a testable prediction by repeating the same recipe on newer models.
- The paper's own target-bias analysis implies that deployment of such models needs fairness auditing over protected groups and topics, since hate probability varies with the group named in the text.
- The same instruction-templated fine-tuning recipe could plausibly transfer to other subtle text classification tasks such as irony, microaggression, or coded-language detection, since the mechanism is generic specialization of embeddings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses implicit hate speech (IHS) detection by comparing two families of classifiers: BERT-based models augmented with emotion and LLM-generated context via four fusion strategies, and general-purpose LLM-based embedding models (E5, Stella, Jasper, NV-Embed) fine-tuned with a fixed instruction template and an MLP head. Experiments on IHC, DynaHate, SBIC, and ToxiGen report in-dataset and cross-dataset F1-macro scores. The central claim is that fine-tuning recent general-purpose embedding models achieves state-of-the-art performance, with up to 1.10 percentage point improvement in in-dataset evaluation and up to 20.35 percentage points in cross-dataset evaluation over published baselines such as LAHN, ConPrompt, ImpCon, SharedCon, and CCL.
Significance. If the empirical claims hold, the practical contribution is substantial: a simple, general recipe (embedding model + MLP head + one fixed instruction) that outperforms specialized contrastive-learning and external-knowledge pipelines for IHS detection, without explicit knowledge injection. The manuscript has genuine strengths: the internal evaluation protocol is clearly described, results are averaged over five seeds with standard deviations, the instruction template is held fixed, code is released, and the paper also documents a useful negative result that BERT-based fusion with emotion/context yields only marginal gains. The cross-dataset generalization trend across model sizes is interesting. However, the headline state-of-the-art claims rest on comparisons with published numbers whose splits, preprocessing, and test subsets are not verified to match the authors' protocol, and the contamination check is too thin to rule out a simpler explanation for the cross-dataset gains. These issues are fixable but load-bearing for the central claim.
major comments (3)
- [Section 4.2 / Table 4] The central SOTA claim is an empirical comparison against published baselines (LAHN, ConPrompt, ImpCon, SharedCon, CCL) in Table 4, but no evidence is provided that the baseline numbers were obtained with the same train/validation/test splits and label preprocessing. Section 4.2 states 60/20/20 (IHC, DynaHate), 80/10/10 (SBIC), and 70/10/20 (ToxiGen), plus thresholds for SBIC and ToxiGen, yet it is not documented that each cited work used these exact splits, the same implicit-hate-only subset of IHC, or the same ToxiGen test subset. If the splits or binarization differ, the reported gains (1.10 pp on IHC; 20.35 pp on IHC→ToxiGen) are uninterpretable. Please rerun the baselines under the same protocol or, at minimum, provide a per-dataset alignment table showing split sizes, filter rules, and test sets used by each baseline, and report significance tests where the margins are small.
- [Section 4.1 vs Section 4.2] The ToxiGen setup is internally inconsistent. Section 4.1 says 'We use the split provided by the authors' and gives 8960/1792/940 train/validation/test samples; Section 4.2 lists '70/10/20 for ToxiGen'. These ratios do not match (8960/11692 ≈ 76.6% train, not 70%). Please state exactly which split was used for Table 3 and for the cross-dataset evaluations in Table 4; the discrepancy directly affects reproducibility and the comparability with Table 4 baselines.
- [Section 4.4, Note on data contamination] The one-sentence contamination check is too weak for the paper's generalization claim. The tested embedding models are trained on web-scale text, and absence of these datasets in model cards is not evidence that they were not seen. Please provide a stronger contamination analysis (e.g., perplexity or nearest-neighbor checks on the test sets, or a discussion of known training corpora) or explicitly state as a limitation that contamination cannot be ruled out. Without this, the cross-dataset gains are harder to attribute to the fine-tuning procedure rather than to pretraining exposure.
minor comments (5)
- [Table 4, IHC→SBIC block] FT E5 and FT Stella report identical F1-M = 72.60 (1.35). This looks like a copy-paste error; please verify the values and standard deviations.
- [Section 4.1 / Table 2] The IHC label counts need clarification. The text says explicit hate samples are excluded, but Table 2's total (18,666) equals only implicit hate + not hate. State the original dataset counts and the filtered counts so readers can compare with baselines.
- [Section 3.2] The instruction 'classify the following in no hate or hate' is ungrammatical; consider 'classify the following as no hate or hate' and state explicitly that this wording was held constant across all models.
- [Section 4.4] The phrase 'loses -0.35 p.p.' is ambiguous; rephrase as 'is 0.35 p.p. below/above' depending on the intended direction of comparison with ConPrompt.
- [Section 1 / Conclusion] The claim 'we are the first to use general-purpose LLM-based embedding models for IHS detection' is a strong novelty claim; consider softening it to avoid being falsified by earlier or concurrent work.
Circularity Check
No significant circularity: the reported results are genuine held-out evaluations of fine-tuned general-purpose embedding models against externally published baselines.
full rationale
The paper makes no derived claim that reduces by construction to its inputs. Its central claim is empirical: fine-tuning general-purpose LLM embedding backbones (E5, Stella, Jasper, NV-Embed) with a two-layer MLP on IHS training data and evaluating on held-out test splits. The instruction template ('Instruct: classify the following in no hate or hate.\nQuery:') is a fixed constant, chosen a priori and not fitted to test labels. The MLP weights are learned on the training portion, and all reported F1-macro values come from separate in-dataset or cross-dataset test sets. No equation in the paper defines a predicted quantity in terms of a fitted parameter; the only fitted parameters are the fine-tuned model and MLP weights, and they are evaluated on data not used for fitting. The cited baselines (LAHN, ConPrompt, SharedCon, CCL, ImpCon, HARE) are external works with independently published numbers, and the paper compares against those numbers rather than deriving them from its own assumptions. There are no self-citations by the present authors, so no load-bearing self-citation chain exists. The skeptical concern about whether the cited baselines used identical train/test splits and label preprocessing is a legitimate threat to the interpretability of the SOTA comparison, but it is a comparability/validity issue, not circularity: the paper's own measurements are genuine held-out evaluations, and no claimed result is forced by definition. The Appendix G limitations and the data-contamination note are cautions about deployment and leakage, not admissions that a result is defined into existence. Therefore no circular step can be exhibited, and the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- instruction template =
Instruct: classify the following in no hate or hate. Query:
- fine-tuning hyperparameters =
lr=2e-6, batch=16 (8 for NV-Embed), epochs=4, wd=0.5
- LoRA rank/alpha for NV-Embed =
r=16, alpha=32
axioms (3)
- domain assumption The four datasets, with the label preprocessing adopted from prior work (IHC uses implicit hate only as the positive class; SBIC uses offensiveness >= 0.5; ToxiGen uses toxicity sum > 5.5), are valid operationalizations of implicit hate speech.
- domain assumption F1-macro is the appropriate primary metric for comparing against the cited state-of-the-art.
- domain assumption The pretrained embedding models were not trained on the evaluation datasets.
Cite this review
Pith. "Pith review of Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets." pith.science (2026). https://pith.science/paper/7HH5KRNN
@misc{pith2026250820750,
author = {Pith},
title = {Pith review of: Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets},
year = {2026},
howpublished = {\url{https://pith.science/paper/7HH5KRNN}},
note = {Machine review of arXiv:2508.20750}
}
read the original abstract
Implicit hate speech (IHS) is indirect language that conveys prejudice or hatred through subtle cues, sarcasm or coded terminology. IHS is challenging to detect as it does not include explicit derogatory or inflammatory words. To address this challenge, task-specific pipelines can be complemented with external knowledge or additional information such as context, emotions and sentiment data. In this paper, we show that, by solely fine-tuning recent general-purpose embedding models based on large language models (LLMs), such as Stella, Jasper, NV-Embed and E5, we achieve state-of-the-art performance. Experiments on multiple IHS datasets show up to 1.10 percentage points improvements for in-dataset, and up to 20.35 percentage points improvements in cross-dataset evaluation, in terms of F1-macro score.
Figures
Reference graph
Works this paper leans on
-
[1]
Oprea, Steven Wilson, and Walid Magdy
Ibrahim Abu Farha, Silviu V. Oprea, Steven Wilson, and Walid Magdy. 2022. SemEval-2022 Task 6: iSarcasmEval, Intended Sarcasm Detection in English and Arabic. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022). Association for Computational Linguistics, Seattle, United States, 802–814. doi:10.18653/v1/2022.semeval-1.111
-
[2]
Hyeseon Ahn, Youngwook Kim, Jungin Kim, and Yo-Sub Han. 2024. Shared- Con: Implicit Hate Speech Detection using Shared Semantics. In Findings of the Association for Computational Linguistics ACL 2024 . Association for Com- putational Linguistics, Bangkok, Thailand and virtual meeting, 10444–10455. doi:10.18653/v1/2024.findings-acl.622
-
[3]
Lopez Monroy, Luis Gonzalez, David E
Mario Aragon, Adrian P. Lopez Monroy, Luis Gonzalez, David E. Losada, and Manuel Montes. 2023. DisorBERT: A Double Domain Adaptation Model for Detecting Signs of Mental Disorders in Social Media. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, To...
-
[4]
Md Rabiul Awal, Rui Cao, Roy Ka-Wei Lee, and Sandra Mitrović. 2021. Angry- BERT: Joint Learning Target and Emotion for Hate Speech Detection. InAdvances in Knowledge Discovery and Data Mining: 25th Pacific-Asia Conference, PAKDD 2021, Virtual Event, May 11–14, 2021, Proceedings, Part I . Springer-Verlag, Berlin, Heidelberg, 701–713. doi:10.1007/978-3-030-...
-
[5]
Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco M. Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019. SemEval- 2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter. InProceedings of the 13th International Workshop on Semantic Evaluation. Association for Computational L...
-
[6]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learn- ing at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computat...
-
[7]
Ocampo, Elena Cabrio, and Serena Villata
Greta Damo, Nicolás B. Ocampo, Elena Cabrio, and Serena Villata. 2024. Unveiling the Hate: Generating Faithful and Plausible Explanations for Implicit and Subtle Hate Speech Detection. In Natural Language Processing and Information Systems . Springer Nature Switzerland, Cham, 211–225
work page 2024
-
[8]
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial Bias in Hate Speech and Abusive Language Detection Datasets. In Proceedings of Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets the Third Workshop on Abusive Language Online . Association for Computational Linguistics, Florence, Italy, 25–...
-
[9]
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated Hate Speech Detection and the Problem of Offensive Language. Proceedings of the International AAAI Conference on Web and Social Media 11, 1 (2017), 512–515. doi:10.1609/icwsm.v11i1.14955
-
[10]
Fabio Del Vigna, Andrea Cimino, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi. 2017. Hate me, hate me not: Hate speech detection on Facebook. In Proceedings of the first Italian conference on cybersecurity (ITASEC17) . 86–95
work page 2017
-
[11]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Com...
-
[12]
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Sey- bolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Com- putational Linguistics, Online and Punta Cana, Domi...
-
[13]
Agneta Fischer, Eran Halperin, Daphna Canetti, and Alba Jasini. 2018. Why We Hate. Emotion Review 10, 4 (2018), 309–320
work page 2018
-
[14]
Antigoni Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018. Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior. Proceedings of the International AAAI Conference on Web and Social Media 12, 1 (Jun. 2018). doi:10.1...
-
[15]
Lei Gao, Alexis Kuppersmith, and Ruihong Huang. 2017. Recognizing Explicit and Implicit Hate Speech Using a Weakly Supervised Two-path Bootstrapping Approach. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Asian Federation of Natural Lan- guage Processing, Taipei, Taiwan, 774–782. https...
work page 2017
-
[16]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al
-
[17]
Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794 [cs.CL] https://arxiv.org/abs/2203.05794
Pith/arXiv arXiv 2022
-
[18]
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics,...
-
[19]
Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . https://openreview.net/forum?id=nZeVKeeFYf9
work page 2022
-
[20]
Fan Huang, Haewoon Kwak, and Jisun An. 2023. Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech. In Companion Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23 Companion) . Association for Computing Machinery, New York, NY, USA, 294–297. doi:10.1145/3543873.3587368
arXiv 2023
-
[21]
Jafari, Guanlin Li, Praboda Rajapaksha, Reza Farahbakhsh, and Noel Crespi
Amir R. Jafari, Guanlin Li, Praboda Rajapaksha, Reza Farahbakhsh, and Noel Crespi. 2023. Fine-Grained Emotions Influence on Implicit Hate Speech Detection. IEEE Access 11 (2023), 105330–105343. doi:10.1109/ACCESS.2023.3318863
-
[22]
Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra S
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra S. Chaplot, et al . 2023. Mistral 7B. arXiv:2310.06825 [cs.CL] https: //arxiv.org/abs/2310.06825
Pith/arXiv arXiv 2023
-
[23]
Tianming Jiang. 2025. Learn from Failure: Causality-guided Contrastive Learning for Generalizable Implicit Hate Speech Detection. In Proceedings of the 31st Inter- national Conference on Computational Linguistics. Association for Computational Linguistics, Abu Dhabi, UAE, 8858–8867. https://aclanthology.org/2025.coling- main.593/
work page 2025
-
[24]
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019. A Just and Comprehensive Strategy for Using NLP to Address Online Abuse. In Pro- ceedings of the 57th Annual Meeting of the Association for Computational Lin- guistics. Association for Computational Linguistics, Florence, Italy, 3658–3666. doi:10.18653/v1/P19-1357
-
[25]
Jaehoon Kim, Seungwan Jin, Sohyun Park, Someen Park, and Kyungsik Han. 2024. Label-aware Hard Negative Sampling Strategies with Momentum Contrastive Learning for Implicit Hate Speech Detection. In Findings of the Association for Computational Linguistics: ACL 2024 . Association for Computational Linguistics, Bangkok, Thailand, 16177–16188. doi:10.18653/v1...
-
[26]
Youngwook Kim, Shinwoo Park, and Yo-Sub Han. 2022. Generalizable Implicit Hate Speech Detection using Contrastive Learning. In Proceedings of the 29th International Conference on Computational Linguistics . International Committee on Computational Linguistics, Gyeongju, Republic of Korea, 6667–6679. https: //aclanthology.org/2022.coling-1.579
work page 2022
-
[27]
Youngwook Kim, Shinwoo Park, Youngsoo Namgoong, and Yo-Sub Han. 2023. ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech Detection. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 10964–10980. doi:10.18653/v1/2023.findings-emnlp.731
-
[28]
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, and Ali Farhadi. 2024. Matryoshka Representation Learning. arXiv:2205.13147 [cs.LG] https://arxiv.org/abs/2205.13147
Pith/arXiv arXiv 2024
-
[29]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025. NV-Embed: Improved Techniques for Train- ing LLMs as Generalist Embedding Models. InThe Thirteenth International Confer- ence on Learning Representations . https://openreview.net/forum?id=lgsyLSsDRe
work page 2025
-
[30]
Lingyao Li, Lizhou Fan, Shubham Atreja, and Libby Hemphill. 2024. “HOT” ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social Media. ACM Trans. Web 18, 2 (2024), 36 pages. doi:10.1145/3643829
doi:10.1145/3643829 2024
-
[31]
Jessica Lin. 2022. Leveraging World Knowledge in Implicit Hate Speech Detec- tion. In Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) . Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid), 31–39. doi:10.18653/v1/2022.nlp4pi-1.4
-
[32]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL] https://arxiv.org/abs/1907.11692
Pith/arXiv arXiv 2019
-
[33]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations . https://openreview.net/ forum?id=Bkg6RiCqY7
work page 2019
-
[34]
Rijul Magu and Jiebo Luo. 2018. Determining Code Words in Euphemistic Hate Speech Using Word Embedding Networks. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2) . Association for Computational Linguistics, Brussels, Belgium, 93–100. doi:10.18653/v1/W18-5112
-
[35]
Sarah Masud, Ashutosh Bajpai, and Tanmoy Chakraborty. 2024. Focal inferential infusion coupled with tractable density discrimination for implicit hate detection. Natural Language Processing (2024), 1–27. doi:10.1017/nlp.2024.60
-
[36]
Leland McInnes, John Healy, and James Melville. 2020. UMAP: Uni- form Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426 [stat.ML] https://arxiv.org/abs/1802.03426
Pith/arXiv arXiv 2020
-
[37]
Changrong Min, Hongfei Lin, Ximing Li, He Zhao, Junyu Lu, Liang Yang, and Bo Xu. 2023. Finding hate speech with auxiliary emotion detection from self- training multi-label learning perspective. Information Fusion 96 (2023), 214–223. doi:10.1016/j.inffus.2023.03.015
-
[38]
Khouloud Mnassri, Praboda Rajapaksha, Reza Farahbakhsh, and Noel Crespi
-
[39]
Marzieh Mozafari, Reza Farahbakhsh, and Noel Crespi. 2020. A BERT-Based Transfer Learning Approach for Hate Speech Detection in Online Social Media. In Complex Networks and Their Applications VIII: Volume 1 Proceedings of the Eighth International Conference on Complex Networks and Their Applications COMPLEX NETWORKS 2019 8. Springer International Publishi...
work page 2020
-
[40]
Marzieh Mozafari, Reza Farahbakhsh, and Noël Crespi. 2020. Hate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model. PloS one 15, 8 (2020), e0237861
work page 2020
-
[41]
Ocampo, Elena Cabrio, and Serena Villata
Nicolás B. Ocampo, Elena Cabrio, and Serena Villata. 2023. Unmasking the Hidden Meaning: Bridging Implicit and Explicit Hate Speech Embedding Representations. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 6626–6637. doi:10.18653/v1/2023.findings-emnlp.441
-
[42]
Juan M. Pérez, Franco M. Luque, Demian Zayat, Martín Kondratzky, Agustín Moro, Pablo S. Serrati, Joaquín Zajac, Paula Miguel, Natalia Debandi, Agustín Gravano, et al. 2023. Assessing the Impact of Contextual Information in Hate Speech Detection. IEEE Access 11 (2023), 30575–30590. doi:10.1109/ACCESS.2023. 3258973
-
[43]
Flor M. Plaza-Del-Arco, M. Dolores Molina-González, L. Alfonso Ureña-López, and María T. Martín-Valdivia. 2021. A Multi-Task Learning Approach to Hate Speech Detection Leveraging Sentiment Analysis. IEEE Access 9 (2021), 112478– 112489. doi:10.1109/ACCESS.2021.3103697
-
[44]
pysentimiento: A Python Toolkit for Opinion Mining and Social NLP tasks
Juan M. Pérez, Mariela Rajngewerc, Juan Carlos Giudici, Damián A. Furman, Franco Luque, Laura Alonso Alemany, and María Vanina Martínez. 2023. py- sentimiento: A Python Toolkit for Opinion Mining and Social NLP tasks. arXiv:2106.09462 [cs.CL] https://arxiv.org/abs/2106.09462
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[45]
Jing Qian, Mai ElSherief, Elizabeth Belding, and William Y. Wang. 2019. Learning to Decipher Hate Symbols. InProceedings of the 2019 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, ...
-
[46]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1, Article 140 (Jan. 2020), 67 pages
work page 2020
-
[47]
Hind Saleh, Areej Alhothali, and Kawthar Moria. 2023. Detection of Hate Speech using BERT and Hate Speech Word Embedding with Deep Model. Applied Artificial Intelligence 37, 1 (2023), 2166719
work page 2023
-
[48]
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 5477–5490. doi:10.18653/v1/2020.acl-main.486
-
[49]
Anna Schmidt and Michael Wiegand. 2017. A Survey on Hate Speech Detec- tion using Natural Language Processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media . Association for Com- putational Linguistics, Valencia, Spain, 1–10. doi:10.18653/v1/W17-1101
-
[50]
Rohit Sridhar and Diyi Yang. 2022. Explaining Toxic Text via Knowledge En- hanced Text Generation. In Proceedings of the 2022 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Seattle, United States, 811–826. doi:10.18653/v1/2022.naacl-main.59
-
[51]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, et al
-
[52]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad F. Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features. ...
Pith/arXiv arXiv 2025
-
[53]
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021. Learn- ing from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection. In Proceedings of the 59th Annual Meeting of the Association for Com- putational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . As...
-
[54]
arXiv:2403.08295 [cs.CL] https://arxiv.org/abs/2403.08295
Gemma: Open Models Based on Gemini Research and Technology. arXiv:2403.08295 [cs.CL] https://arxiv.org/abs/2403.08295
-
[55]
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Multilingual E5 Text Embeddings: A Technical Report. arXiv:2402.05672 [cs.CL] https://arxiv.org/abs/2402.05672
Pith/arXiv arXiv 2024
-
[56]
William Warner and Julia Hirschberg. 2012. Detecting Hate Speech on the World Wide Web. InProceedings of the Second Workshop on Language in Social Media. Association for Computational Linguistics, Montréal, Canada, 19–26. https://aclanthology.org/W12-2103
work page 2012
-
[57]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024. Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv:2212.03533 [cs.CL] https://arxiv.org/abs/2212. 03533
Pith/arXiv arXiv 2024
-
[58]
Zeerak Waseem and Dirk Hovy. 2016. Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. In Proceedings of the NAACL Student Research Workshop. Association for Computational Linguistics, San Diego, California, 88–93. doi:10.18653/v1/N16-2013
-
[59]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, et al. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https://arxiv.org/abs/2505.09388
Pith/arXiv arXiv 2025
-
[60]
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017. Understanding Abuse: A Typology of Abusive Language Detection Subtasks. In Proceedings of the First Workshop on Abusive Language Online . Association for Computational Linguistics, Vancouver, BC, Canada, 78–84. doi:10.18653/v1/W17- 3012
doi:10.18653/v1/w17- 2017
-
[61]
Yongjin Yang, Joonkee Kim, Yujin Kim, Namgyu Ho, James Thorne, and Se- Young Yun. 2023. HARE: Explainable Hate Speech Detection with Step-by- Step Reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 5490–5505. doi:10.18653/v1/2023.findings-emnlp.365
-
[62]
Ziyuan Yang, Ming Yan, Yingyu Chen, Hui Wang, Zexin Lu, and Yi Zhang
-
[63]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, et al. 2024. Qwen2 Technical Report. arXiv:2407.10671 [cs.CL] https://arxiv.org/abs/2407.10671
Pith/arXiv arXiv 2024
-
[64]
Dun Zhang, Jiacheng Li, Ziyang Zeng, and Fulong Wang. 2025. Jasper and Stella: distillation of SOTA embedding models. arXiv:2412.19048 [cs.IR] https: //arxiv.org/abs/2412.19048
Pith/arXiv arXiv 2025
-
[65]
Min Zhang, Jianfeng He, Taoran Ji, and Chang-Tien Lu. 2024. Don’t Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Bangkok,...
-
[66]
Trustworthy Hate Speech Detection Through Visual Augmentation
Trustworthy Hate Speech Detection Through Visual Augmentation. arXiv:2409.13557 [cs.CV] https://arxiv.org/abs/2409.13557
work page internal anchor Pith review Pith/arXiv arXiv
-
[67]
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, and other
-
[68]
In The Twelfth International Conference on Learning Representations
KoLA: Carefully Benchmarking World Knowledge of Large Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=AqN23oqraW
-
[71]
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Proce...
-
[72]
Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2023. Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks. arXiv:2304.10145 [cs.AI] https://arxiv.org/abs/2304.10145 Appendices for "Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets" A Context generation with Llama2...
Pith/arXiv arXiv 2023
-
[2023]
In ICC 2023 - IEEE International Conference on Communications
Hate Speech and Offensive Language Detection Using an Emotion-Aware Shared Encoder. In ICC 2023 - IEEE International Conference on Communications . 2852–2857. doi:10.1109/ICC45041.2023.10279690
arXiv 2023
-
[2024]
arXiv:2407.21783 [cs.AI] https://arxiv.org/ abs/2407.21783
The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/ abs/2407.21783
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.