REVIEW 4 major objections 4 minor 77 references
MM-HSD: Multi-Modal Hate Speech Detection in Videos
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read 0.874 macro-F1 for video hate detection with on-screen text as query
desk verdict Solid engineering paper with a plausible new finding—on-screen text as query in cross-modal attention—but the configuration-selection evidence is internally inconsistent, so the headline M-F1 needs a reproducibility fix before I'd trust it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is Cross-Modal Attention (CMA), an attention operation of the form $\mathrm{softmax}(Q K^\top / \sqrt{d_k})V$ that lets one modality query a sequence of features from other modalities. In MM-HSD the query is the on-screen text embedding and the key/value is the concatenation, along the sequence dimension, of the transcript, audio, and video embeddings; the CMA output is then concatenated with separately encoded unimodal outputs before the final classification head. The design decouples visual context from on-screen text by using a vision transformer for frames and a separate OCR module for the text in the frame, so the query channel is not redundant with the video channel. The role of the CMA block is to let a sparse, sometimes absent, on-screen-text signal absorb contextual cues from the other three modalities and to do so at the embedding level rather than after classification.
What would settle it
Re-run MM-HSD on HateMM while choosing the attention arrangement from the average validation macro-F1 over several random seeds instead of one seed; if the chosen arrangement changes or the test macro-F1 falls below the 0.848 of the previous best method, the claim that on-screen text is the best query is not supported. A second check is to fix the arrangement and evaluate on a different video hate dataset, such as MultiHateClip, where failing to beat a transcript-only model would undercut the generality of the four-modality design.
Extended reading notes
Core claim
The paper reports that on the HateMM dataset, a tetra-modal pipeline — speech-transcript text embeddings, audio embeddings from a self-supervised speech encoder, video-frame embeddings from a vision transformer, and on-screen text extracted by OCR — fused by a cross-modal attention (CMA) block plus unimodal encoders reaches an unbiased accuracy (mean of class-wise recalls) of 0.878, macro-F1 of 0.874, and hate-class F1 of 0.853 over five runs. The configuration that delivers this is CMA applied to raw embeddings with on-screen text as query and the concatenation of transcript, audio, and video as key, with the CMA output concatenated to the four unimodal encoder outputs before classification. Across the full grid of query/key combinations, this O-query/TAV-key choice is the best, while using on-screen text as part of the key generally hurts performance. Removing any single modality from the full model drops macro-F1 to between 0.815 and 0.845, and replacing CMA by plain concatenation drops it to 0.842, so the paper's claim is that every component contributes and the attention asymmetry is load-bearing.
Load-bearing premise
The key premise is that the arrangement chosen by trying many alternatives on the validation portion with single-seed runs — letting on-screen text pull information from transcript, audio, and video — is also the best arrangement on the held-out test set; if that search got lucky on validation, the reported 0.874 macro-F1 could be too high.
Editorial extensions
If this is right
- Video hate-speech systems should treat on-screen text as a distinct modality: OCR captures hate conveyed in frames that transcript-only or image-only models cannot see.
- Early cross-modal attention plus late fusion of the unimodal encoders beats either alone: the full model reaches 0.874 macro-F1 against 0.846 for CMA used standalone and 0.842 for concatenation without CMA.
- The attention asymmetry is a design choice that matters: making on-screen text the query and transcript, audio, and video the key is the best arrangement among all tested query-key pairs.
- Every modality pulls its weight in the tetra-modal setting; removing any one of transcript, audio, video, or on-screen text lowers macro-F1, with transcript removal causing the largest drop.
- Deployment is plausible once features are precomputed: the final classifier has 4.6 million parameters even though the upstream feature extractors are much larger.
Reading between the lines
- If the query-asymmetry result transfers, it suggests a general heuristic for multimodal fusion: let the modality most likely to carry the target signal be the attention query and let richer context modalities be the keys, rather than always defaulting to text as the query.
- Because the configuration search over query-key pairs was run on validation macro-F1 with single-seed runs, the true advantage of the O-query/TAV-key choice over its nearest alternatives may be smaller; averaging configuration selection over seeds or folds is a direct robustness check the paper leaves open.
- The same architecture is naturally extendable to frame-level hate localisation by making the attention temporally aware, which would also expose which frames drive a video-level hate verdict.
- A testable extension is to swap the transcript channel for speech-to-text output in another language (or a dataset without on-screen text) to see whether the OCR-query design remains optimal, or whether the query should become the modality that is most sparse or most hate-bearing in that new setting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents MM-HSD, a multi-modal hate speech detection model for videos that integrates four modalities: speech transcript (T), audio (A), video frames (V), and on-screen text (O). The proposed architecture applies cross-modal attention (CMA) to the raw modality embeddings and concatenates the CMA output with separately encoded per-modality representations before classification. The authors report that the best configuration uses on-screen text as the attention query and the concatenation of transcript, audio, and video as the key, and they report a macro-F1 of 0.874 on the HateMM benchmark, outperforming published baselines. The paper also contributes an ablation study of modality subsets, query-key configurations, early versus late fusion, and computational efficiency, and it releases the code.
Significance. If the reported benchmark result holds, this is a useful and timely empirical contribution to video-based hate speech detection, a relatively understudied task. The inclusion of on-screen text as a separate modality, the systematic query-key analysis, the released code, and the explicit comparison against several recent baselines are strengths that would make the paper valuable to the community. The reported gap to prior state-of-the-art macro-F1 (0.874 versus 0.848) is practically meaningful. However, the central empirical claim is currently clouded by internal inconsistencies in the reported results and in the experimental protocol, so the significance of the contribution cannot be fully assessed until these issues are resolved.
major comments (4)
- [Section 5.5 and Appendix A, Tables 6 and 7, versus Table 3] The results in Tables 6 and 7 are inconsistent with Table 3 for the same selected configuration. For query O and key TAV, Table 6 reports M-F1 = 0.877 and Table 7 reports M-F1 = 0.884, while Table 3 lists CMA-LF† at 0.837 (std 0.024) and CMA-S† at 0.846 (std 0.006) for the same configuration. Since the captions identify Table 6 as late-fusion CMA and Table 7 as raw-input CMA, the reader expects these values to agree with the corresponding rows of Table 3; a gap of 0.038–0.047 is far larger than the reported standard deviations and cannot be attributed to seed variation. The reader therefore cannot determine which architecture actually produced Tables 6 and 7, and the headline 0.874 M-F1 and the design claim that O-as-query is best rest on an unidentified model. Please clarify the exact model behind each table and provide multi-seed mean and standard deviation for the configuration search.
- [Section 5.1 and Section 5.6] The reported data-split counts do not add up. Section 5.1 states that 5-fold cross-validation is performed on 85% of the data, with each fold containing 698 training and 175 validation samples, while Section 5.6 states that inference is performed over 155 test samples. For the 1083 labeled videos in HateMM, these numbers imply different totals: 873 videos for train plus validation, 920.55 for 85% of the data, and 1028 for 698+175+155. Please report the exact number of videos used after preprocessing, describe how the split was constructed, and state how many videos were discarded and why. Without this information, the comparison in Table 2 cannot be reproduced.
- [Section 5.5 and Section 4.1] The query-key configuration search is conducted with a single seed and on validation folds, and the selected configuration is then adopted for MM-HSD. Section 5.5 states that the experiments in Tables 6 and 7 were performed in a single-seed setup. If those tables are actually MM-HSD runs, the final configuration is selected from 28 single-seed validation results on the same benchmark, with no multi-seed comparison of alternative configurations. If they are CMA-LF and CMA-S runs, the transfer of the O-query/TAV-key choice to MM-HSD is not directly supported. Either way, the claim that on-screen text is the best query and the headline 0.874 M-F1 need multi-seed mean and standard deviation for at least the top configurations, and the selection procedure should be described in a way that makes clear whether test-set information was used.
- [Table 2 and Section 5.1] The state-of-the-art comparison in Table 2 is based on scores taken from the original papers rather than re-running baselines under the identical protocol. Because the paper uses a specific 85/15 split with 5-fold cross-validation, differences in data partitioning, preprocessing, and evaluation code can materially affect macro-F1. Please either re-run the prior baselines under the same split and report mean and standard deviation, or explicitly state the splits used by each baseline and justify why the comparison is fair. At minimum, a significance test should be reported for the difference between MM-HSD (0.874, std 0.009) and the previous best HCC1 result (0.848), and the comparison with the single-run result of TCE-DBF [71] should be qualified accordingly.
minor comments (4)
- [Appendix A, paragraph before Table 6] The appendix describes Table 7 as 'early fusion (model II)', but Figure 2 labels model II as 'Late Fusion with CMA as Additional Modality' and model III as 'Early Fusion with CMA as Unique Feature Extractor'. Please make the model labels consistent between the figure and the appendix.
- [Section 5.4, sentence citing Table 3] The sentence 'M-F1 score of 0.878 (MM-HSD) against 0.846 (w/o CMA)' appears to report accuracy values, since Table 3 lists MM-HSD M-F1 = 0.874 and w/o CMA M-F1 = 0.842. Please correct the metric labels.
- [Table 7, TVA row] The value '0.82813' in the P(H) column appears to be a typo; it should likely be a three-decimal value consistent with the other entries.
- [Section 5.5, correlation statement] The reported 'very strong positive correlation of 91%' between CMA-S and CMA-LF performance gains does not state the type of correlation coefficient or its statistical significance. Please add this detail.
Circularity Check
No circularity found: the central claim is an empirical benchmark result with a held-out test evaluation; the internal table inconsistencies create a reproducibility concern, not circular reasoning.
full rationale
The paper does not contain a mathematical derivation chain whose output is equivalent to its input. The central claim is an empirical benchmark result on HateMM: MM-HSD achieves an M-F1 of 0.874 on a held-out test split (15% of the data), with the model trained and validated via 5-fold cross-validation on the remaining 85% (Section 5.1). The query-key configuration (on-screen text as query, concatenated transcript/audio/video as key) is selected by systematic comparison on validation folds (Section 5.5, Appendix A, Tables 6 and 7), and the final performance is then reported on the test set. This is standard model selection followed by out-of-sample evaluation, not a fitted parameter renamed as a prediction. The paper contains no self-citations: all prior-work references, such as [12], [40], [66], and [71], are external. The authors explicitly note that the query-key search used a single seed, which is a limitation on the strength of the configuration-selection claim but not a circular step. A separate concern is that Table 3 reports CMA-S as 0.846 and CMA-LF as 0.837, while Appendix Tables 6 and 7 list the same named configurations as 0.877 and 0.884, respectively; this inconsistency undermines reproducibility and the clarity of the model identity underlying the headline result, but it does not make any claim equivalent to its inputs by construction or by definition. Therefore, no specific circular step can be exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (6)
- learning rate =
chosen from {1e-3, 1e-4, 1e-5} via validation
- L1 penalty (elastic net) =
chosen from {1e-3, 1e-4, 1e-5} via validation
- L2 penalty (elastic net) =
chosen from {1e-4, 1e-5, 1e-6} via validation
- dropout rate =
chosen from {0.3, 0.4, 0.5} via validation
- early stopping patience =
chosen from {5, 10} via validation
- query/key configuration =
O as query, TAV as key
assumptions (3)
- domain assumption HateMM labels are correct and representative of video hate speech.
- domain assumption Pre-trained feature extractors (Detoxify, ViT, wav2vec2, Whisper, PaddleOCR) produce embeddings that preserve hate-relevant information without fine-tuning.
- domain assumption Cross-modal attention with modalities concatenated along the sequence dimension is an appropriate way to model inter-modal dependencies.
Cite this review
Pith. "Pith review of MM-HSD: Multi-Modal Hate Speech Detection in Videos." pith.science (2026). https://pith.science/paper/XK5UUTAB
@misc{pith2026250820546,
author = {Pith},
title = {Pith review of: MM-HSD: Multi-Modal Hate Speech Detection in Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/XK5UUTAB}},
note = {Machine review of arXiv:2508.20546}
}
read the original abstract
While hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully capture inter-modal dependencies. Moreover, previous work often omits relevant modalities such as on-screen text and audio, which may contain subtle hateful content and thus provide essential cues, both individually and in combination with others. In this paper, we present MM-HSD, a multi-modal model for HSD in videos that integrates video frames, audio, and text derived from speech transcripts and from frames (i.e.~on-screen text) together with features extracted by Cross-Modal Attention (CMA). We are the first to use CMA as an early feature extractor for HSD in videos, to systematically compare query/key configurations, and to evaluate the interactions between different modalities in the CMA block. Our approach leads to improved performance when on-screen text is used as a query and the rest of the modalities serve as a key. Experiments on the HateMM dataset show that MM-HSD outperforms state-of-the-art methods on M-F1 score (0.874), using concatenation of transcript, audio, video, on-screen text, and CMA for feature extraction on raw embeddings of the modalities. The code is available at https://github.com/idiap/mm-hsd
Figures
Reference graph
Works this paper leans on
-
[71]
Haitao Xiong, Wei Jiao, and Yuanyuan Cai. 2024. TCE-DBF: Textual Context Enhanced Dynamic Bimodal Fusion for Hate Video Detection. Data Technologies and Applications (2024)
work page 2024
-
[1]
Cleber Alcântara, Viviane Moreira, and Diego Feijo. 2020. Offensive Video Detection: Dataset and Baseline Results. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 4309–4319
work page 2020
-
[2]
Ahlam Alrehili. 2019. Automatic Hate Speech Detection on Social Media: A Brief Survey. In Proceedings of the 2019 IEEE/ACS 16th International Conference on Computer Systems and Applications (AICCSA) . IEEE, 1–6
work page 2019
-
[3]
Jinmyeong An, Wonjun Lee, Yejin Jeon, Jungseul Ok, Yunsu Kim, and Gary G. Lee. 2024. An Investigation into Explainable Audio Hate Speech Detection. In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. Association for Computational Linguistics, Kyoto, Japan, 533–543. doi:10.18653/v1/2024.sigdial-1.45
-
[4]
Iván Arcos and Paolo Rosso. 2024. Sexism Identification on TikTok: A Mul- timodal AI Approach with Text, Audio, and Video. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 15th International Conference of the CLEF Association, CLEF 2024, Grenoble, France, September 9–12, 2024, Pro- ceedings, Part I (Grenoble, France). Springer-Ver...
-
[5]
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 12449–12460
work page 2020
-
[6]
Kirtilekha Bhesra and Akshay Agarwal. 2025. A Multi-modal Framework to Counter Hate Speeches. In Pattern Recognition, Apostolos Antonacopoulos, Subhasis Chaudhuri, Rama Chellappa, Cheng-Lin Liu, Saumik Bhattacharya, and Umapada Pal (Eds.). Springer Nature Switzerland, Cham, 197–207
work page 2025
-
[7]
Kirtilekha Bhesra, Shivam A. Shukla, and Akshay Agarwal. 2024. Audio vs. Text: Identify a Powerful Modality for Effective Hate Speech Detection. In The Second Tiny Papers Track at ICLR 2024
work page 2024
Show all 77 references
-
[8]
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2021. HateBERT: Retraining BERT for Abusive Language Detection in English. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021) , Aida Mostafazadeh Davani, Douwe Kiela, Mathias Lambert...
2021 doi
-
[9]
Silva, Deborah L
Lu Cheng, Kai Shu, Siqi Wu, Yasin N. Silva, Deborah L. Hall, and Huan Liu. 2020. Unsupervised Cyberbullying Detection via Time-Informed Gaussian Mixture Model. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland...
2020
-
[10]
Vishwakarma
Anusha Chhabra and Dinesh K. Vishwakarma. 2023. A Literature Survey on Multimodal and Multilingual Automatic Hate Speech Identification. Multimedia Systems 29, 3 (2023), 1203–1230
2023
-
[11]
Lu Chi, Guiyu Tian, Yadong Mu, and Qi Tian. 2019. Two-Stream Video Classifica- tion with Cross-Modality Attention. In Proceedings of the IEEE/CVF international conference on computer vision workshops
2019
-
[12]
Mithun Das, Rohit Raj, Punyajoy Saha, Binny Mathew, Manish Gupta, and Animesh Mukherjee. 2023. HateMM: A Multi-Modal Dataset for Hate Video Classification. Proceedings of the International AAAI Conference on Web and Social Media 17, 1 (Jun. 2023), 1014–1023. doi:10.1609/icwsm....
2023 doi
-
[13]
Mithun Das, Rohit Raj, Punyajoy Saha, Binny Mathew, Manish Gupta, and Animesh Mukherjee. 2023. HateMM: A Multi-modal Dataset for Hate Video Classification. doi:10.5281/zenodo.7799469 MM-HSD: Multi-Modal Hate Speech Detection in Videos
2023 doi
-
[14]
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial Bias in Hate Speech and Abusive Language Detection Datasets. In Proceedings of the Third Workshop on Abusive Language Online , Sarah T. Roberts, Joel Tetreault, Vinodkumar Prabhakaran, and Zeerak Waseem (E...
2019 doi
-
[15]
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018. Hate Speech Dataset from a White Supremacy Forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2) , Darja Fišer, Ruihong Huang, Vinodkumar Prabhakaran, Rob Voigt, Zeerak Waseem, an...
2018 doi
-
[16]
Debele and Michael M
Abreham G. Debele and Michael M. Woldeyohannis. 2022. Multimodal Amharic Hate Speech Detection Using Deep Learning. In 2022 International Conference on Information and Communication Technology for Development for Africa (ICT4DA) . IEEE, 102–107
2022
-
[17]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...
2021
-
[18]
Duong, Remi Lebret, and Karl Aberer
Chi T. Duong, Remi Lebret, and Karl Aberer. 2017. Multimodal Classification for Analysing Social Media. arXiv:1708.02099 [cs.CL]
2017 arXiv
-
[19]
Eniafe Festus Ayetiran and Özlem Özgöbek. 2024. A Review of Deep Learning Techniques for Multimodal Fake News and Harmful Languages Detection. IEEE Access 12 (2024), 76133–76153. doi:10.1109/access.2024.3406258
2024
-
[20]
Jinmiao Fu, Shaoyuan Xu, Huidong Liu, Yang Liu, Ning Xie, Chien-Chih Wang, Jia Liu, Yi Sun, and Bryan Wang. 2022. CMA-CLIP: Cross-Modality Attention Clip for Text-Image Classification. In2022 IEEE International Conference on Image Processing (ICIP). 2846–2850. doi:10.1109/icip...
2022
-
[21]
Ankita Gandhi, Param Ahir, Kinjal Adhvaryu, Pooja Shah, Ritika Lohiya, Erik Cambria, Soujanya Poria, and Amir Hussain. 2024. Hate Speech Detection: A Comprehensive Review of Recent Works. Expert Systems (2024)
2024
-
[22]
Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. 2020. Ex- ploring Hate Speech Detection in Multimodal Publications. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1470–1478
2020
-
[23]
Gongane, Mousami V
Vaishali U. Gongane, Mousami V. Munot, and Alwin D. Anuse. 2022. Detection and Moderation of Detrimental Content on Social Media Platforms: Current Status and Future Directions. Social Network Analysis and Mining 12, 1 (2022), 129
2022
-
[24]
Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu
Satya K. Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. 2022. X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5006–5015
2022
-
[25]
Jonatas Grosman. 2021. Fine-tuned XLSR-53 Large Model for Speech Recognition in English. https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53- english
2021
-
[26]
Shrey Gupta, Pratyush Priyadarshi, and Manish Gupta. 2023. Hateful Comment Detection and Hate Target Type Prediction for Video Comments. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Man- agement. Acm, Birmingham United Kingdom, 3923–3927...
2023 doi
-
[27]
Hana, Said Al Faraby, and Arif Bramantoro
Karimah M. Hana, Said Al Faraby, and Arif Bramantoro. 2020. Multi-Label Clas- sification of Indonesian Hate Speech on Twitter Using Support Vector Machines. In 2020 International Conference on Data Science and Its Applications (ICoDSA) . IEEE, 1–7
2020
-
[28]
Laura Hanu and Unitary team. 2020. Detoxify. Github. https://github.com/unitaryai/detoxify
2020
-
[29]
Hatebase. 2025. Hatebase is a collaborative, regionalized repository of multilin- gual hate speech. https://hatebase.org Accessed: 2025-02-27
2025
-
[30]
Hee, Shivam Sharma, Rui Cao, Palash Nandi, Preslav Nakov, Tanmoy Chakraborty, and Roy Ka-Wei Lee
Ming S. Hee, Shivam Sharma, Rui Cao, Palash Nandi, Preslav Nakov, Tanmoy Chakraborty, and Roy Ka-Wei Lee. 2024. Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models. In Findings of the Association for Computational Linguistics: EMNLP 202...
2024 doi
-
[31]
Hoque, M
Eftekhar Hossain, Omar Sharif, Mohammed M. Hoque, M. A. Akber Dewan, Nazmul Siddique, and Md. A. Hossain. 2022. Identification of Multilingual Offense and Troll from Social Media Memes Using Weighted Ensemble of Multimodal Features. Journal of King Saud University - Computer a...
2022 doi
-
[32]
Mohd. I. Hossain Junaid, Faisal Hossain, and Rashedur M. Rahman. 2021. Bangla Hate Speech Detection in Videos Using Machine Learning. In 2021 IEEE 12th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON). IEEE, New York, NY, USA, 0347–0351. doi:...
2021
-
[33]
Olga Jubany and Malin Roiha. 2016. Backgrounds, Experiences and Responses to Online Hate Speech: A Comparative Cross-Country Analysis. Online report. Barcelona: University of Barcelona (2016)
2016
-
[34]
Rajeshwari Kandakatla. 2016. Identifying Offensive Videos on YouTube . Mas- ter’s thesis. Wright State University. http://rave.ohiolink.edu/etdc/view?acc_ num=wright1484751212961772 Available at OhioLINK Electronic Theses and Dissertations Center
2016
-
[35]
Simon Kemp. 2025. Digital 2025: The State of Social Media in 2025. (February 2025). https://datareportal.com/reports/digital-2025-sub-section-state-of-social
2025
-
[36]
Khan, and Muhammad K
Hareem Kibriya, Ayesha Siddiqa, Wazir Z. Khan, and Muhammad K. Khan. 2024. Towards Safer Online Communities: Deep Learning and Explainable AI for Hate Speech Detection and Classification. Computers and Electrical Engineering 116 (2024), 109153
2024
-
[37]
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020. The Hateful Memes Chal- lenge: Detecting Hate Speech in Multimodal Memes. Advances in neural infor- mation processing systems 33 (2020), 2611–2624
2020
-
[38]
Kiilu, George Okeyo, Richard Rimiru, and Kennedy Ogada
Kelvin K. Kiilu, George Okeyo, Richard Rimiru, and Kennedy Ogada. 2018. Using Naïve Bayes Algorithm in Detection of Hate Tweets. International Journal of Scientific and Research Publications 8, 3 (2018), 99–107
2018
-
[39]
Dirk Kindermann. 2023. Against ‘Hate Speech’. Journal of Applied Philosophy 40, 5 (2023), 813–835. doi:10.1111/japp.12648
2023 doi
-
[40]
Koushik, Diptesh Kanojia, and Helen Treharne
Girish A. Koushik, Diptesh Kanojia, and Helen Treharne. 2025. Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content. In Companion Proceedings of the ACM on Web Conference 2025 (Sydney NSW, Australia) (WWW’25). Association for Comput...
2025
-
[41]
Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, and Fan Zhou. 2025. Biting off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection. In The Web Conference 2025. https://openreview. net/forum?id=GrJYzmDzfW
2025
-
[42]
Phillip Lippe, Nithin Holla, Shantanu Chandra, Santhosh Rajamanickam, Geor- gios Antoniou, Ekaterina Shutova, and Helen Yannakoudakis. 2020. A Multimodal Framework for the Detection of Hateful Memes. arXiv:2012.12871 [cs.CL]
2020 arXiv
-
[43]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]
2019 arXiv
-
[44]
Zhiyu Ma, Shaowen Yao, Liwen Wu, Song Gao, and Yunqi Zhang. 2022. Hateful Memes Detection Based on Multi-Task Learning. Mathematics 10, 23 (2022), 4525
2022
-
[45]
Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019. Hate Speech Detection: Challenges and Solutions. PloS one 14, 8 (2019)
2019
-
[46]
Madukwe, Xiaoying Gao, and Bing Xue
Kosisochukwu J. Madukwe, Xiaoying Gao, and Bing Xue. 2022. Token Replacement-Based Data Augmentation Methods for Hate Speech Detection. World Wide Web 25, 3 (2022), 1129–1150
2022
-
[47]
Krishanu Maity, A. S. Poornash, Sriparna Saha, and Pushpak Bhattacharyya
-
[48]
Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee
Binny Mathew, Punyajoy Saha, Seid M. Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. HateXplain: A Benchmark Dataset for Explain- able Hate Speech Detection. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 14867–14875
2021
-
[49]
Sophia Koepke, and Zeynep Akata
Otniel-Bogdan Mercea, Thomas Hummel, A. Sophia Koepke, and Zeynep Akata
-
[50]
Mullah and Wan M
Nanlir S. Mullah and Wan M. N. W. Zainon. 2021. Advances in Machine Learning Algorithms for Hate Speech Detection in Social Media: A Review. IEEE Access 9 (2021), 88364–88376. doi:10.1109/access.2021.3089515
2021
-
[51]
Noever and Samantha E
David A. Noever and Samantha E. M. Noever. 2021. Reading Isn’t Believing: Adversarial Attacks On Multi-Modal Neurons. arXiv:2103.10480 [cs.LG]
2021 arXiv
-
[52]
PaddlePaddle. 2025. PaddleOCR: Multi-Language OCR System. https://github. com/PaddlePaddle/PaddleOCR. Accessed: 2025-04-11
2025
-
[53]
Konstantinos Perifanos and Dionysis Goutsos. 2021. Multimodal Hate Speech Detection in Greek Social Media. Multimodal Technologies and Interaction 5, 7 (2021). doi:10.3390/mti5070034
2021 doi
-
[54]
Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Costantino Grana, Rita Cucchiara, and Elisa Ficarra. 2025. Semantically Condi- tioned Prompts for Visual Recognition Under Missing Modality Scenarios. In 2025 IEEE/CVF Winter Conference on Applica...
2025
-
[55]
R. G. Praveen and Jahangir Alam. 2024. Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. 4803–4813. Berta Céspedes-Sarrias, Carl...
2024
-
[56]
Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever
Alec Radford, Jong W. Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models from Natural Language Supervision. In International...
2021
-
[57]
Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever
Alec Radford, Jong W. Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust Speech Recognition via Large-Scale Weak Supervision. In International conference on machine learning . PMLR, 28492–28518
2023
-
[58]
Aneri Rana and Sonali Jha. 2022. Emotion Based Hate Speech Detection using Multimodal Learning. arXiv:2202.06218 [cs.LG]
2022 arXiv
-
[59]
Anchal Rawat, Santosh Kumar, and Surender S. Samant. 2024. Hate Speech Detection in Social Media: Techniques, Recent Trends, and Future Challenges. Wiley Interdisciplinary Reviews: Computational Statistics 16, 2 (2024)
2024
-
[60]
Roy, Asis K
Pradeep K. Roy, Asis K. Tripathy, Tapan K. Das, and Xiao-Zhi Gao. 2020. A Frame- work for Hate Speech Detection Using Deep Convolutional Neural Network. IEEE Access 8 (2020), 204951–204962
2020
-
[61]
Saleem, Kelly P
Haji M. Saleem, Kelly P. Dillon, Susan Benesch, and Derek Ruths. 2017. A Web of Hate: Tackling Hateful Speech in Online Social Spaces. arXiv:1709.10159 [cs.CL]
2017 arXiv
-
[62]
Vlad Sandulescu. 2020. Detecting Hateful Memes Using a Multimodal Deep Ensemble. arXiv:2012.13235 [cs.LG]
2020 arXiv
-
[63]
Chakravarthi, Mihael Arcan, and Paul Buitelaar
Shardul Suryawanshi, Bharathi R. Chakravarthi, Mihael Arcan, and Paul Buitelaar
-
[64]
George-Alexandru Vlad, George-Eduard Zaharia, Dumitru-Clementin Cercel, and Mihai Dascalu. 2020. UPB@DANKMEMES: Italian Memes Analysis-Employing Visual Models and Graph Convolutional Networks for Meme Identification and Hate Speech Detection. EV ALITA Evaluation of NLP and Spe...
2020
-
[65]
Hongbo Wang, Junyu Lu, Yan Han, Kai Ma, Liang Yang, and Hongfei Lin. 2025. Towards Patronizing and Condescending Language in Chinese Videos: A Multi- modal Dataset and Detector. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS...
2025
-
[66]
Tan, and Roy K
Han Wang, Rui Y. Tan, and Roy K. Lee. 2025. Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection. InProceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(Www ’25). Association for Computing Machinery, New York, NY, USA, ...
2025 doi
-
[67]
Tan, Usman Naseem, and Roy K
Han Wang, Rui Y. Tan, Usman Naseem, and Roy K. Lee. 2024. MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili. In Proceedings of the 32nd ACM International Conference on Multimedia (Melbourne VIC, Australia) (MM ’24). Association...
2024
-
[68]
Wilson and Molly K
Richard A. Wilson and Molly K. Land. 2020. Hate Speech on Social Media: Content Moderation in Context. Connecticut Law Review 52 (2020), 1029
2020
-
[69]
Wu and Unnathi Bhandary
Ching S. Wu and Unnathi Bhandary. 2020. Detection of Hate Speech in Videos Using Machine Learning. In 2020 International Conference on Computational Science and Computational Intelligence (CSCI) . IEEE, Las Vegas, NV, USA, 585–
2020
-
[70]
Yang Wu, Pengwei Zhan, Yunjian Zhang, Liming Wang, and Zhen Xu. 2021. Mul- timodal Fusion with Co-Attention Networks for Fake News Detection. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli ...
2021 doi
-
[72]
Fan Yang, Xiaochang Peng, Gargi Ghosh, Reshef Shilon, Hao Ma, Eider Moore, and Goran Predovic. 2019. Exploring Deep Multimodal Fusion of Text and Photo for Hate Speech Classification. Association for Computational Linguistics, Florence, Italy. doi:10.18653/v1/W19-3502
2019 doi
-
[73]
Weibo Zhang, Guihua Liu, Zhuohua Li, and Fuqing Zhu. 2020. Hate- ful Memes Detection via Complementary Visual and Linguistic Networks. arXiv:2012.04977 [cs.CV] MM-HSD: Multi-Modal Hate Speech Detection in Videos APPENDIX A. EARLY VS. LATE FUSION Table 6 (model I), which corres...
2020 arXiv
-
[590]
doi:10.1109/csci51800.2020.00104
2020
-
[2020]
In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying
Multimodal Meme Dataset (MultiOFF) for Identifying Offensive Content in Image and Text. In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying. European Language Resources Association (ELRA), Marseille, France, 32–41. https://aclanthology.org/2020.trac-1.6
2020
-
[2022]
In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part XX (Tel Aviv, Israel)
Temporal and Cross-modal Attention for Audio-Visual Zero-Shot Learning. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part XX (Tel Aviv, Israel). Springer-Verlag, Berlin, Heidelberg, 488–505. doi:10.1007/978-3-03...
2022 doi
-
[2024]
In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.)
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code- Mixed Videos. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 11130–1114...
2024 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.