Pith. sign in

REVIEW 1 cited by

Multimodal and Explainable Internet Meme Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.05612 v3 pith:P2QNBV5T submitted 2022-12-11 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords classificationexplainablemememethodsmodelsinternetmemesmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the current context where online platforms have been effectively weaponized in a variety of geo-political events and social issues, Internet memes make fair content moderation at scale even more difficult. Existing work on meme classification and tracking has focused on black-box methods that do not explicitly consider the semantics of the memes or the context of their creation. In this paper, we pursue a modular and explainable architecture for Internet meme understanding. We design and implement multimodal classification methods that perform example- and prototype-based reasoning over training cases, while leveraging both textual and visual SOTA models to represent the individual cases. We study the relevance of our modular and explainable models in detecting harmful memes on two existing tasks: Hate Speech Detection and Misogyny Classification. We compare the performance between example- and prototype-based methods, and between text, vision, and multimodal models, across different categories of harmfulness (e.g., stereotype and objectification). We devise a user-friendly interface that facilitates the comparative analysis of examples retrieved by all of our models for any given meme, informing the community about the strengths and limitations of these explainable methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

    cs.CV 2026-07 conditional novelty 3.0 of 10

    Cross-modal multi-head attention over CLIP vision and XLM-R OCR tokens beats concatenation and unimodal baselines at ~0.94 Macro-F1 on Bengali political meme detection; lexicon priors hurt.

Pith tools