Simple embedding fusion improves F1 by 9.9 points on the HateMM video dataset but reaches only 0.628 AUROC on the Hateful Memes dataset, showing that fusion methods do not transfer across modality types.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
Simple embedding fusion improves F1 by 9.9 points on the HateMM video dataset but reaches only 0.628 AUROC on the Hateful Memes dataset, showing that fusion methods do not transfer across modality types.