Pith. sign in

REVIEW 4 major objections 4 minor 77 references

MM-HSD: Multi-Modal Hate Speech Detection in Videos

T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read 0.874 macro-F1 for video hate detection with on-screen text as query

desk verdict Solid engineering paper with a plausible new finding—on-screen text as query in cross-modal attention—but the configuration-selection evidence is internally inconsistent, so the headline M-F1 needs a reproducibility fix before I'd trust it. read the letter →

arxiv 2508.20546 v1 pith:XK5UUTAB submitted 2025-08-28 cs.MM cs.AI

classification cs.MMcs.AI
keywords hatespeechdetectionmultimodalfusioncross-modalattentionvideoclassificationon-screentextOCRaudio-visualanalysisMM
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MM-HSD is a four-modality model for detecting hate speech in videos, combining the speech transcript, audio, video frames, and on-screen text. The paper's central claim is that using cross-modal attention with on-screen text as the query and the concatenated transcript, audio, and video embeddings as the key gives the best fusion configuration, reaching a macro-F1 of 0.874 on the HateMM benchmark and outperforming the previous best published score of 0.848. This matters because hate in videos often appears in a non-obvious modality: letting a sparse, potentially hate-carrying channel like on-screen text pull context from richer modalities captures dependencies that simple concatenation or late fusion can miss. The authors also present the first systematic query/key comparison for cross-modal attention in video hate speech detection, and ablate every component to show that each modality and the attention block contribute.

What carries the argument

The load-bearing mechanism is Cross-Modal Attention (CMA), an attention operation of the form $\mathrm{softmax}(Q K^\top / \sqrt{d_k})V$ that lets one modality query a sequence of features from other modalities. In MM-HSD the query is the on-screen text embedding and the key/value is the concatenation, along the sequence dimension, of the transcript, audio, and video embeddings; the CMA output is then concatenated with separately encoded unimodal outputs before the final classification head. The design decouples visual context from on-screen text by using a vision transformer for frames and a separate OCR module for the text in the frame, so the query channel is not redundant with the video channel. The role of the CMA block is to let a sparse, sometimes absent, on-screen-text signal absorb contextual cues from the other three modalities and to do so at the embedding level rather than after classification.

What would settle it

Re-run MM-HSD on HateMM while choosing the attention arrangement from the average validation macro-F1 over several random seeds instead of one seed; if the chosen arrangement changes or the test macro-F1 falls below the 0.848 of the previous best method, the claim that on-screen text is the best query is not supported. A second check is to fix the arrangement and evaluate on a different video hate dataset, such as MultiHateClip, where failing to beat a transcript-only model would undercut the generality of the four-modality design.

Watch

Extended reading notes

Core claim

The paper reports that on the HateMM dataset, a tetra-modal pipeline — speech-transcript text embeddings, audio embeddings from a self-supervised speech encoder, video-frame embeddings from a vision transformer, and on-screen text extracted by OCR — fused by a cross-modal attention (CMA) block plus unimodal encoders reaches an unbiased accuracy (mean of class-wise recalls) of 0.878, macro-F1 of 0.874, and hate-class F1 of 0.853 over five runs. The configuration that delivers this is CMA applied to raw embeddings with on-screen text as query and the concatenation of transcript, audio, and video as key, with the CMA output concatenated to the four unimodal encoder outputs before classification. Across the full grid of query/key combinations, this O-query/TAV-key choice is the best, while using on-screen text as part of the key generally hurts performance. Removing any single modality from the full model drops macro-F1 to between 0.815 and 0.845, and replacing CMA by plain concatenation drops it to 0.842, so the paper's claim is that every component contributes and the attention asymmetry is load-bearing.

Load-bearing premise

The key premise is that the arrangement chosen by trying many alternatives on the validation portion with single-seed runs — letting on-screen text pull information from transcript, audio, and video — is also the best arrangement on the held-out test set; if that search got lucky on validation, the reported 0.874 macro-F1 could be too high.

Editorial extensions

If this is right

  • Video hate-speech systems should treat on-screen text as a distinct modality: OCR captures hate conveyed in frames that transcript-only or image-only models cannot see.
  • Early cross-modal attention plus late fusion of the unimodal encoders beats either alone: the full model reaches 0.874 macro-F1 against 0.846 for CMA used standalone and 0.842 for concatenation without CMA.
  • The attention asymmetry is a design choice that matters: making on-screen text the query and transcript, audio, and video the key is the best arrangement among all tested query-key pairs.
  • Every modality pulls its weight in the tetra-modal setting; removing any one of transcript, audio, video, or on-screen text lowers macro-F1, with transcript removal causing the largest drop.
  • Deployment is plausible once features are precomputed: the final classifier has 4.6 million parameters even though the upstream feature extractors are much larger.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the query-asymmetry result transfers, it suggests a general heuristic for multimodal fusion: let the modality most likely to carry the target signal be the attention query and let richer context modalities be the keys, rather than always defaulting to text as the query.
  • Because the configuration search over query-key pairs was run on validation macro-F1 with single-seed runs, the true advantage of the O-query/TAV-key choice over its nearest alternatives may be smaller; averaging configuration selection over seeds or folds is a direct robustness check the paper leaves open.
  • The same architecture is naturally extendable to frame-level hate localisation by making the attention temporally aware, which would also expose which frames drive a video-level hate verdict.
  • A testable extension is to swap the transcript channel for speech-to-text output in another language (or a dataset without on-screen text) to see whether the OCR-query design remains optimal, or whether the query should become the modality that is most sparse or most hate-bearing in that new setting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper presents MM-HSD, a multi-modal hate speech detection model for videos that integrates four modalities: speech transcript (T), audio (A), video frames (V), and on-screen text (O). The proposed architecture applies cross-modal attention (CMA) to the raw modality embeddings and concatenates the CMA output with separately encoded per-modality representations before classification. The authors report that the best configuration uses on-screen text as the attention query and the concatenation of transcript, audio, and video as the key, and they report a macro-F1 of 0.874 on the HateMM benchmark, outperforming published baselines. The paper also contributes an ablation study of modality subsets, query-key configurations, early versus late fusion, and computational efficiency, and it releases the code.

Significance. If the reported benchmark result holds, this is a useful and timely empirical contribution to video-based hate speech detection, a relatively understudied task. The inclusion of on-screen text as a separate modality, the systematic query-key analysis, the released code, and the explicit comparison against several recent baselines are strengths that would make the paper valuable to the community. The reported gap to prior state-of-the-art macro-F1 (0.874 versus 0.848) is practically meaningful. However, the central empirical claim is currently clouded by internal inconsistencies in the reported results and in the experimental protocol, so the significance of the contribution cannot be fully assessed until these issues are resolved.

major comments (4)
  1. [Section 5.5 and Appendix A, Tables 6 and 7, versus Table 3] The results in Tables 6 and 7 are inconsistent with Table 3 for the same selected configuration. For query O and key TAV, Table 6 reports M-F1 = 0.877 and Table 7 reports M-F1 = 0.884, while Table 3 lists CMA-LF† at 0.837 (std 0.024) and CMA-S† at 0.846 (std 0.006) for the same configuration. Since the captions identify Table 6 as late-fusion CMA and Table 7 as raw-input CMA, the reader expects these values to agree with the corresponding rows of Table 3; a gap of 0.038–0.047 is far larger than the reported standard deviations and cannot be attributed to seed variation. The reader therefore cannot determine which architecture actually produced Tables 6 and 7, and the headline 0.874 M-F1 and the design claim that O-as-query is best rest on an unidentified model. Please clarify the exact model behind each table and provide multi-seed mean and standard deviation for the configuration search.
  2. [Section 5.1 and Section 5.6] The reported data-split counts do not add up. Section 5.1 states that 5-fold cross-validation is performed on 85% of the data, with each fold containing 698 training and 175 validation samples, while Section 5.6 states that inference is performed over 155 test samples. For the 1083 labeled videos in HateMM, these numbers imply different totals: 873 videos for train plus validation, 920.55 for 85% of the data, and 1028 for 698+175+155. Please report the exact number of videos used after preprocessing, describe how the split was constructed, and state how many videos were discarded and why. Without this information, the comparison in Table 2 cannot be reproduced.
  3. [Section 5.5 and Section 4.1] The query-key configuration search is conducted with a single seed and on validation folds, and the selected configuration is then adopted for MM-HSD. Section 5.5 states that the experiments in Tables 6 and 7 were performed in a single-seed setup. If those tables are actually MM-HSD runs, the final configuration is selected from 28 single-seed validation results on the same benchmark, with no multi-seed comparison of alternative configurations. If they are CMA-LF and CMA-S runs, the transfer of the O-query/TAV-key choice to MM-HSD is not directly supported. Either way, the claim that on-screen text is the best query and the headline 0.874 M-F1 need multi-seed mean and standard deviation for at least the top configurations, and the selection procedure should be described in a way that makes clear whether test-set information was used.
  4. [Table 2 and Section 5.1] The state-of-the-art comparison in Table 2 is based on scores taken from the original papers rather than re-running baselines under the identical protocol. Because the paper uses a specific 85/15 split with 5-fold cross-validation, differences in data partitioning, preprocessing, and evaluation code can materially affect macro-F1. Please either re-run the prior baselines under the same split and report mean and standard deviation, or explicitly state the splits used by each baseline and justify why the comparison is fair. At minimum, a significance test should be reported for the difference between MM-HSD (0.874, std 0.009) and the previous best HCC1 result (0.848), and the comparison with the single-run result of TCE-DBF [71] should be qualified accordingly.
minor comments (4)
  1. [Appendix A, paragraph before Table 6] The appendix describes Table 7 as 'early fusion (model II)', but Figure 2 labels model II as 'Late Fusion with CMA as Additional Modality' and model III as 'Early Fusion with CMA as Unique Feature Extractor'. Please make the model labels consistent between the figure and the appendix.
  2. [Section 5.4, sentence citing Table 3] The sentence 'M-F1 score of 0.878 (MM-HSD) against 0.846 (w/o CMA)' appears to report accuracy values, since Table 3 lists MM-HSD M-F1 = 0.874 and w/o CMA M-F1 = 0.842. Please correct the metric labels.
  3. [Table 7, TVA row] The value '0.82813' in the P(H) column appears to be a typo; it should likely be a three-decimal value consistent with the other entries.
  4. [Section 5.5, correlation statement] The reported 'very strong positive correlation of 91%' between CMA-S and CMA-LF performance gains does not state the type of correlation coefficient or its statistical significance. Please add this detail.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the central claim is an empirical benchmark result with a held-out test evaluation; the internal table inconsistencies create a reproducibility concern, not circular reasoning.

full rationale

The paper does not contain a mathematical derivation chain whose output is equivalent to its input. The central claim is an empirical benchmark result on HateMM: MM-HSD achieves an M-F1 of 0.874 on a held-out test split (15% of the data), with the model trained and validated via 5-fold cross-validation on the remaining 85% (Section 5.1). The query-key configuration (on-screen text as query, concatenated transcript/audio/video as key) is selected by systematic comparison on validation folds (Section 5.5, Appendix A, Tables 6 and 7), and the final performance is then reported on the test set. This is standard model selection followed by out-of-sample evaluation, not a fitted parameter renamed as a prediction. The paper contains no self-citations: all prior-work references, such as [12], [40], [66], and [71], are external. The authors explicitly note that the query-key search used a single seed, which is a limitation on the strength of the configuration-selection claim but not a circular step. A separate concern is that Table 3 reports CMA-S as 0.846 and CMA-LF as 0.837, while Appendix Tables 6 and 7 list the same named configurations as 0.877 and 0.884, respectively; this inconsistency undermines reproducibility and the clarity of the model identity underlying the headline result, but it does not make any claim equivalent to its inputs by construction or by definition. Therefore, no specific circular step can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 6 free parameters · 3 assumptions · 0 invented entities

The central claim depends on hyperparameters and a configuration choice selected via validation on the same benchmark, and on the adequacy of frozen pre-trained encoders. No new theoretical entities are introduced; the model is an engineering combination of existing components.

free parameters (6)
  • learning rate = chosen from {1e-3, 1e-4, 1e-5} via validation
    Hyperparameter tuned on validation folds; affects convergence and final performance.
  • L1 penalty (elastic net) = chosen from {1e-3, 1e-4, 1e-5} via validation
    Regularization strength selected by validation performance.
  • L2 penalty (elastic net) = chosen from {1e-4, 1e-5, 1e-6} via validation
    Regularization strength selected by validation performance.
  • dropout rate = chosen from {0.3, 0.4, 0.5} via validation
    Regularization hyperparameter selected on validation data.
  • early stopping patience = chosen from {5, 10} via validation
    Controls training length and model selection.
  • query/key configuration = O as query, TAV as key
    Selected from a large set of combinations based on validation M-F1 on the same dataset; a discrete design choice fitted to the validation data.
assumptions (3)
  • domain assumption HateMM labels are correct and representative of video hate speech.
    The model is trained and evaluated solely on HateMM; if annotations are noisy or the dataset is not representative, the reported performance and design conclusions may not generalize.
  • domain assumption Pre-trained feature extractors (Detoxify, ViT, wav2vec2, Whisper, PaddleOCR) produce embeddings that preserve hate-relevant information without fine-tuning.
    The authors avoid fine-tuning unimodal extractors due to small dataset size, assuming frozen representations are sufficient for the downstream task.
  • domain assumption Cross-modal attention with modalities concatenated along the sequence dimension is an appropriate way to model inter-modal dependencies.
    The CMA mechanism assumes that concatenated sequences from different modalities can be processed as a single sequence, which may not account for modal-specific distributions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MM-HSD: Multi-Modal Hate Speech Detection in Videos." pith.science (2026). https://pith.science/paper/XK5UUTAB

@misc{pith2026250820546,
  author       = {Pith},
  title        = {Pith review of: MM-HSD: Multi-Modal Hate Speech Detection in Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XK5UUTAB}},
  note         = {Machine review of arXiv:2508.20546}
}
read the original abstract

While hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully capture inter-modal dependencies. Moreover, previous work often omits relevant modalities such as on-screen text and audio, which may contain subtle hateful content and thus provide essential cues, both individually and in combination with others. In this paper, we present MM-HSD, a multi-modal model for HSD in videos that integrates video frames, audio, and text derived from speech transcripts and from frames (i.e.~on-screen text) together with features extracted by Cross-Modal Attention (CMA). We are the first to use CMA as an early feature extractor for HSD in videos, to systematically compare query/key configurations, and to evaluate the interactions between different modalities in the CMA block. Our approach leads to improved performance when on-screen text is used as a query and the rest of the modalities serve as a key. Experiments on the HateMM dataset show that MM-HSD outperforms state-of-the-art methods on M-F1 score (0.874), using concatenation of transcript, audio, video, on-screen text, and CMA for feature extraction on raw embeddings of the modalities. The code is available at https://github.com/idiap/mm-hsd

Figures

Figures reproduced from arXiv: 2508.20546 by the authors.

Figure 1
Figure 1. Sample frames extracted from videos classified as [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Our tetra-modal HSD architecture. The four modalities are audio transcripts, audio signal, video frames, and [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Impact of modality combinations on the M-F1 score. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 52 canonical work pages

  1. [71]

    Haitao Xiong, Wei Jiao, and Yuanyuan Cai. 2024. TCE-DBF: Textual Context Enhanced Dynamic Bimodal Fusion for Hate Video Detection. Data Technologies and Applications (2024)

  2. [1]

    Cleber Alcântara, Viviane Moreira, and Diego Feijo. 2020. Offensive Video Detection: Dataset and Baseline Results. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 4309–4319

  3. [2]

    Ahlam Alrehili. 2019. Automatic Hate Speech Detection on Social Media: A Brief Survey. In Proceedings of the 2019 IEEE/ACS 16th International Conference on Computer Systems and Applications (AICCSA) . IEEE, 1–6

  4. [3]

    Jinmyeong An, Wonjun Lee, Yejin Jeon, Jungseul Ok, Yunsu Kim, and Gary G. Lee. 2024. An Investigation into Explainable Audio Hate Speech Detection. In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. Association for Computational Linguistics, Kyoto, Japan, 533–543. doi:10.18653/v1/2024.sigdial-1.45

  5. [4]

    Iván Arcos and Paolo Rosso. 2024. Sexism Identification on TikTok: A Mul- timodal AI Approach with Text, Audio, and Video. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 15th International Conference of the CLEF Association, CLEF 2024, Grenoble, France, September 9–12, 2024, Pro- ceedings, Part I (Grenoble, France). Springer-Ver...

  6. [5]

    Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 12449–12460

  7. [6]

    Kirtilekha Bhesra and Akshay Agarwal. 2025. A Multi-modal Framework to Counter Hate Speeches. In Pattern Recognition, Apostolos Antonacopoulos, Subhasis Chaudhuri, Rama Chellappa, Cheng-Lin Liu, Saumik Bhattacharya, and Umapada Pal (Eds.). Springer Nature Switzerland, Cham, 197–207

  8. [7]

    Shukla, and Akshay Agarwal

    Kirtilekha Bhesra, Shivam A. Shukla, and Akshay Agarwal. 2024. Audio vs. Text: Identify a Powerful Modality for Effective Hate Speech Detection. In The Second Tiny Papers Track at ICLR 2024

Show all 77 references
  1. [8]

    Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2021. HateBERT: Retraining BERT for Abusive Language Detection in English. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021) , Aida Mostafazadeh Davani, Douwe Kiela, Mathias Lambert...

  2. [9]

    Silva, Deborah L

    Lu Cheng, Kai Shu, Siqi Wu, Yasin N. Silva, Deborah L. Hall, and Huan Liu. 2020. Unsupervised Cyberbullying Detection via Time-Informed Gaussian Mixture Model. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland...

  3. [10]

    Vishwakarma

    Anusha Chhabra and Dinesh K. Vishwakarma. 2023. A Literature Survey on Multimodal and Multilingual Automatic Hate Speech Identification. Multimedia Systems 29, 3 (2023), 1203–1230

  4. [11]

    Lu Chi, Guiyu Tian, Yadong Mu, and Qi Tian. 2019. Two-Stream Video Classifica- tion with Cross-Modality Attention. In Proceedings of the IEEE/CVF international conference on computer vision workshops

  5. [12]

    Mithun Das, Rohit Raj, Punyajoy Saha, Binny Mathew, Manish Gupta, and Animesh Mukherjee. 2023. HateMM: A Multi-Modal Dataset for Hate Video Classification. Proceedings of the International AAAI Conference on Web and Social Media 17, 1 (Jun. 2023), 1014–1023. doi:10.1609/icwsm....

  6. [13]

    Mithun Das, Rohit Raj, Punyajoy Saha, Binny Mathew, Manish Gupta, and Animesh Mukherjee. 2023. HateMM: A Multi-modal Dataset for Hate Video Classification. doi:10.5281/zenodo.7799469 MM-HSD: Multi-Modal Hate Speech Detection in Videos

  7. [14]

    Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial Bias in Hate Speech and Abusive Language Detection Datasets. In Proceedings of the Third Workshop on Abusive Language Online , Sarah T. Roberts, Joel Tetreault, Vinodkumar Prabhakaran, and Zeerak Waseem (E...

  8. [15]

    Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018. Hate Speech Dataset from a White Supremacy Forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2) , Darja Fišer, Ruihong Huang, Vinodkumar Prabhakaran, Rob Voigt, Zeerak Waseem, an...

  9. [16]

    Debele and Michael M

    Abreham G. Debele and Michael M. Woldeyohannis. 2022. Multimodal Amharic Hate Speech Detection Using Deep Learning. In 2022 International Conference on Information and Communication Technology for Development for Africa (ICT4DA) . IEEE, 102–107

  10. [17]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...

  11. [18]

    Duong, Remi Lebret, and Karl Aberer

    Chi T. Duong, Remi Lebret, and Karl Aberer. 2017. Multimodal Classification for Analysing Social Media. arXiv:1708.02099 [cs.CL]

  12. [19]

    Eniafe Festus Ayetiran and Özlem Özgöbek. 2024. A Review of Deep Learning Techniques for Multimodal Fake News and Harmful Languages Detection. IEEE Access 12 (2024), 76133–76153. doi:10.1109/access.2024.3406258

  13. [20]

    Jinmiao Fu, Shaoyuan Xu, Huidong Liu, Yang Liu, Ning Xie, Chien-Chih Wang, Jia Liu, Yi Sun, and Bryan Wang. 2022. CMA-CLIP: Cross-Modality Attention Clip for Text-Image Classification. In2022 IEEE International Conference on Image Processing (ICIP). 2846–2850. doi:10.1109/icip...

  14. [21]

    Ankita Gandhi, Param Ahir, Kinjal Adhvaryu, Pooja Shah, Ritika Lohiya, Erik Cambria, Soujanya Poria, and Amir Hussain. 2024. Hate Speech Detection: A Comprehensive Review of Recent Works. Expert Systems (2024)

  15. [22]

    Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. 2020. Ex- ploring Hate Speech Detection in Multimodal Publications. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1470–1478

  16. [23]

    Gongane, Mousami V

    Vaishali U. Gongane, Mousami V. Munot, and Alwin D. Anuse. 2022. Detection and Moderation of Detrimental Content on Social Media Platforms: Current Status and Future Directions. Social Network Analysis and Mining 12, 1 (2022), 129

  17. [24]

    Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu

    Satya K. Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. 2022. X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5006–5015

  18. [25]

    Jonatas Grosman. 2021. Fine-tuned XLSR-53 Large Model for Speech Recognition in English. https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53- english

  19. [26]

    Shrey Gupta, Pratyush Priyadarshi, and Manish Gupta. 2023. Hateful Comment Detection and Hate Target Type Prediction for Video Comments. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Man- agement. Acm, Birmingham United Kingdom, 3923–3927...

  20. [27]

    Hana, Said Al Faraby, and Arif Bramantoro

    Karimah M. Hana, Said Al Faraby, and Arif Bramantoro. 2020. Multi-Label Clas- sification of Indonesian Hate Speech on Twitter Using Support Vector Machines. In 2020 International Conference on Data Science and Its Applications (ICoDSA) . IEEE, 1–7

  21. [28]

    Laura Hanu and Unitary team. 2020. Detoxify. Github. https://github.com/unitaryai/detoxify

  22. [29]

    Hatebase. 2025. Hatebase is a collaborative, regionalized repository of multilin- gual hate speech. https://hatebase.org Accessed: 2025-02-27

  23. [30]

    Hee, Shivam Sharma, Rui Cao, Palash Nandi, Preslav Nakov, Tanmoy Chakraborty, and Roy Ka-Wei Lee

    Ming S. Hee, Shivam Sharma, Rui Cao, Palash Nandi, Preslav Nakov, Tanmoy Chakraborty, and Roy Ka-Wei Lee. 2024. Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models. In Findings of the Association for Computational Linguistics: EMNLP 202...

  24. [31]

    Hoque, M

    Eftekhar Hossain, Omar Sharif, Mohammed M. Hoque, M. A. Akber Dewan, Nazmul Siddique, and Md. A. Hossain. 2022. Identification of Multilingual Offense and Troll from Social Media Memes Using Weighted Ensemble of Multimodal Features. Journal of King Saud University - Computer a...

  25. [32]

    Mohd. I. Hossain Junaid, Faisal Hossain, and Rashedur M. Rahman. 2021. Bangla Hate Speech Detection in Videos Using Machine Learning. In 2021 IEEE 12th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON). IEEE, New York, NY, USA, 0347–0351. doi:...

  26. [33]

    Olga Jubany and Malin Roiha. 2016. Backgrounds, Experiences and Responses to Online Hate Speech: A Comparative Cross-Country Analysis. Online report. Barcelona: University of Barcelona (2016)

  27. [34]

    Rajeshwari Kandakatla. 2016. Identifying Offensive Videos on YouTube . Mas- ter’s thesis. Wright State University. http://rave.ohiolink.edu/etdc/view?acc_ num=wright1484751212961772 Available at OhioLINK Electronic Theses and Dissertations Center

  28. [35]

    Simon Kemp. 2025. Digital 2025: The State of Social Media in 2025. (February 2025). https://datareportal.com/reports/digital-2025-sub-section-state-of-social

  29. [36]

    Khan, and Muhammad K

    Hareem Kibriya, Ayesha Siddiqa, Wazir Z. Khan, and Muhammad K. Khan. 2024. Towards Safer Online Communities: Deep Learning and Explainable AI for Hate Speech Detection and Classification. Computers and Electrical Engineering 116 (2024), 109153

  30. [37]

    Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020. The Hateful Memes Chal- lenge: Detecting Hate Speech in Multimodal Memes. Advances in neural infor- mation processing systems 33 (2020), 2611–2624

  31. [38]

    Kiilu, George Okeyo, Richard Rimiru, and Kennedy Ogada

    Kelvin K. Kiilu, George Okeyo, Richard Rimiru, and Kennedy Ogada. 2018. Using Naïve Bayes Algorithm in Detection of Hate Tweets. International Journal of Scientific and Research Publications 8, 3 (2018), 99–107

  32. [39]

    Dirk Kindermann. 2023. Against ‘Hate Speech’. Journal of Applied Philosophy 40, 5 (2023), 813–835. doi:10.1111/japp.12648

  33. [40]

    Koushik, Diptesh Kanojia, and Helen Treharne

    Girish A. Koushik, Diptesh Kanojia, and Helen Treharne. 2025. Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content. In Companion Proceedings of the ACM on Web Conference 2025 (Sydney NSW, Australia) (WWW’25). Association for Comput...

  34. [41]

    Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, and Fan Zhou. 2025. Biting off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection. In The Web Conference 2025. https://openreview. net/forum?id=GrJYzmDzfW

  35. [42]

    Phillip Lippe, Nithin Holla, Shantanu Chandra, Santhosh Rajamanickam, Geor- gios Antoniou, Ekaterina Shutova, and Helen Yannakoudakis. 2020. A Multimodal Framework for the Detection of Hateful Memes. arXiv:2012.12871 [cs.CL]

  36. [43]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]

  37. [44]

    Zhiyu Ma, Shaowen Yao, Liwen Wu, Song Gao, and Yunqi Zhang. 2022. Hateful Memes Detection Based on Multi-Task Learning. Mathematics 10, 23 (2022), 4525

  38. [45]

    Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019. Hate Speech Detection: Challenges and Solutions. PloS one 14, 8 (2019)

  39. [46]

    Madukwe, Xiaoying Gao, and Bing Xue

    Kosisochukwu J. Madukwe, Xiaoying Gao, and Bing Xue. 2022. Token Replacement-Based Data Augmentation Methods for Hate Speech Detection. World Wide Web 25, 3 (2022), 1129–1150

  40. [47]

    Krishanu Maity, A. S. Poornash, Sriparna Saha, and Pushpak Bhattacharyya

  41. [48]

    Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee

    Binny Mathew, Punyajoy Saha, Seid M. Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. HateXplain: A Benchmark Dataset for Explain- able Hate Speech Detection. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 14867–14875

  42. [49]

    Sophia Koepke, and Zeynep Akata

    Otniel-Bogdan Mercea, Thomas Hummel, A. Sophia Koepke, and Zeynep Akata

  43. [50]

    Mullah and Wan M

    Nanlir S. Mullah and Wan M. N. W. Zainon. 2021. Advances in Machine Learning Algorithms for Hate Speech Detection in Social Media: A Review. IEEE Access 9 (2021), 88364–88376. doi:10.1109/access.2021.3089515

  44. [51]

    Noever and Samantha E

    David A. Noever and Samantha E. M. Noever. 2021. Reading Isn’t Believing: Adversarial Attacks On Multi-Modal Neurons. arXiv:2103.10480 [cs.LG]

  45. [52]

    PaddlePaddle. 2025. PaddleOCR: Multi-Language OCR System. https://github. com/PaddlePaddle/PaddleOCR. Accessed: 2025-04-11

  46. [53]

    Konstantinos Perifanos and Dionysis Goutsos. 2021. Multimodal Hate Speech Detection in Greek Social Media. Multimodal Technologies and Interaction 5, 7 (2021). doi:10.3390/mti5070034

  47. [54]

    Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Costantino Grana, Rita Cucchiara, and Elisa Ficarra. 2025. Semantically Condi- tioned Prompts for Visual Recognition Under Missing Modality Scenarios. In 2025 IEEE/CVF Winter Conference on Applica...

  48. [55]

    R. G. Praveen and Jahangir Alam. 2024. Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. 4803–4813. Berta Céspedes-Sarrias, Carl...

  49. [56]

    Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

    Alec Radford, Jong W. Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models from Natural Language Supervision. In International...

  50. [57]

    Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever

    Alec Radford, Jong W. Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust Speech Recognition via Large-Scale Weak Supervision. In International conference on machine learning . PMLR, 28492–28518

  51. [58]

    Aneri Rana and Sonali Jha. 2022. Emotion Based Hate Speech Detection using Multimodal Learning. arXiv:2202.06218 [cs.LG]

  52. [59]

    Anchal Rawat, Santosh Kumar, and Surender S. Samant. 2024. Hate Speech Detection in Social Media: Techniques, Recent Trends, and Future Challenges. Wiley Interdisciplinary Reviews: Computational Statistics 16, 2 (2024)

  53. [60]

    Roy, Asis K

    Pradeep K. Roy, Asis K. Tripathy, Tapan K. Das, and Xiao-Zhi Gao. 2020. A Frame- work for Hate Speech Detection Using Deep Convolutional Neural Network. IEEE Access 8 (2020), 204951–204962

  54. [61]

    Saleem, Kelly P

    Haji M. Saleem, Kelly P. Dillon, Susan Benesch, and Derek Ruths. 2017. A Web of Hate: Tackling Hateful Speech in Online Social Spaces. arXiv:1709.10159 [cs.CL]

  55. [62]

    Vlad Sandulescu. 2020. Detecting Hateful Memes Using a Multimodal Deep Ensemble. arXiv:2012.13235 [cs.LG]

  56. [63]

    Chakravarthi, Mihael Arcan, and Paul Buitelaar

    Shardul Suryawanshi, Bharathi R. Chakravarthi, Mihael Arcan, and Paul Buitelaar

  57. [64]

    George-Alexandru Vlad, George-Eduard Zaharia, Dumitru-Clementin Cercel, and Mihai Dascalu. 2020. UPB@DANKMEMES: Italian Memes Analysis-Employing Visual Models and Graph Convolutional Networks for Meme Identification and Hate Speech Detection. EV ALITA Evaluation of NLP and Spe...

  58. [65]

    Hongbo Wang, Junyu Lu, Yan Han, Kai Ma, Liang Yang, and Hongfei Lin. 2025. Towards Patronizing and Condescending Language in Chinese Videos: A Multi- modal Dataset and Detector. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS...

  59. [66]

    Tan, and Roy K

    Han Wang, Rui Y. Tan, and Roy K. Lee. 2025. Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection. InProceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(Www ’25). Association for Computing Machinery, New York, NY, USA, ...

  60. [67]

    Tan, Usman Naseem, and Roy K

    Han Wang, Rui Y. Tan, Usman Naseem, and Roy K. Lee. 2024. MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili. In Proceedings of the 32nd ACM International Conference on Multimedia (Melbourne VIC, Australia) (MM ’24). Association...

  61. [68]

    Wilson and Molly K

    Richard A. Wilson and Molly K. Land. 2020. Hate Speech on Social Media: Content Moderation in Context. Connecticut Law Review 52 (2020), 1029

  62. [69]

    Wu and Unnathi Bhandary

    Ching S. Wu and Unnathi Bhandary. 2020. Detection of Hate Speech in Videos Using Machine Learning. In 2020 International Conference on Computational Science and Computational Intelligence (CSCI) . IEEE, Las Vegas, NV, USA, 585–

  63. [70]

    Yang Wu, Pengwei Zhan, Yunjian Zhang, Liming Wang, and Zhen Xu. 2021. Mul- timodal Fusion with Co-Attention Networks for Fake News Detection. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli ...

  64. [72]

    Fan Yang, Xiaochang Peng, Gargi Ghosh, Reshef Shilon, Hao Ma, Eider Moore, and Goran Predovic. 2019. Exploring Deep Multimodal Fusion of Text and Photo for Hate Speech Classification. Association for Computational Linguistics, Florence, Italy. doi:10.18653/v1/W19-3502

  65. [73]

    Weibo Zhang, Guihua Liu, Zhuohua Li, and Fuqing Zhu. 2020. Hate- ful Memes Detection via Complementary Visual and Linguistic Networks. arXiv:2012.04977 [cs.CV] MM-HSD: Multi-Modal Hate Speech Detection in Videos APPENDIX A. EARLY VS. LATE FUSION Table 6 (model I), which corres...

  66. [590]

    doi:10.1109/csci51800.2020.00104

  67. [2020]

    In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying

    Multimodal Meme Dataset (MultiOFF) for Identifying Offensive Content in Image and Text. In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying. European Language Resources Association (ELRA), Marseille, France, 32–41. https://aclanthology.org/2020.trac-1.6

  68. [2022]

    In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part XX (Tel Aviv, Israel)

    Temporal and Cross-modal Attention for Audio-Visual Zero-Shot Learning. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part XX (Tel Aviv, Israel). Springer-Verlag, Berlin, Heidelberg, 488–505. doi:10.1007/978-3-03...

  69. [2024]

    In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.)

    ToxVidLM: A Multimodal Framework for Toxicity Detection in Code- Mixed Videos. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 11130–1114...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.