Pith. sign in

REVIEW 2 cited by

MemeFier: Dual-stage Modality Fusion for Image Meme Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02906 v2 pith:GQ3JPQJS submitted 2023-04-06 cs.CV

classification cs.CV
keywords imagefusionhatememefiermemesmodalityclassificationcontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hate speech is a societal problem that has significantly grown through the Internet. New forms of digital content such as image memes have given rise to spread of hate using multimodal means, being far more difficult to analyse and detect compared to the unimodal case. Accurate automatic processing, analysis and understanding of this kind of content will facilitate the endeavor of hindering hate speech proliferation through the digital world. To this end, we propose MemeFier, a deep learning-based architecture for fine-grained classification of Internet image memes, utilizing a dual-stage modality fusion module. The first fusion stage produces feature vectors containing modality alignment information that captures non-trivial connections between the text and image of a meme. The second fusion stage leverages the power of a Transformer encoder to learn inter-modality correlations at the token level and yield an informative representation. Additionally, we consider external knowledge as an additional input, and background image caption supervision as a regularizing component. Extensive experiments on three widely adopted benchmarks, i.e., Facebook Hateful Memes, Memotion7k and MultiOFF, indicate that our approach competes and in some cases surpasses state-of-the-art. Our code is available on https://github.com/ckoutlis/memefier.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revealing Temporal Label Noise in Multimodal Hateful Video Classification

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Temporal label noise is systemic in video-level hate annotations: trimming hateful videos to timestamped hate segments raises macro-F1 by 19.34% and 30.45% on HateMM and MultiHateClip-English.

  2. Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A method that decomposes CLIP image and text features into a shared low-rank component and modality-specific sparse components, then uses an attention-weighted soft prompt to guide an LLM for sentiment, emotion, and h...

Pith tools