Pith. sign in

REVIEW 1 cited by

SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.20541 v2 pith:XHN2C6IR submitted 2024-12-29 cs.CL cs.CY

classification cs.CLcs.CY
keywords memesframeworkhatereasoningdatasetsdetectionhatefulmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret them correctly. It makes multimodal meme moderation difficult, as existing work either lacks high-quality datasets for nuanced hate categories or relies on low-quality social media visuals. Here, we curate two novel multimodal hate speech datasets comprising English memes - MHS and MHS-Con, which capture fine-grained hateful abstractions in regular and confounding scenarios, respectively. We benchmark these datasets against several competing baselines. Furthermore, we introduce SAFE-MEME (Structured reAsoning FramEwork) with its two variants: a novel multimodal Chain-of-Thought based framework employing Q&A-style reasoning (SAFE-MEME-QA) and a hierarchical categorization (SAFE-MEME-H) to enable robust hate speech detection in memes. SAFE-MEME-QA outperforms the strongest open-source baseline model, showing an improvement of 2% on MHS and 4.7% on MHS-Con and closely follows the closed-source models, GPT-4o and Gemini 2.5. In contrast, SAFE-MEME-H surpasses SAFE-MEME-QA, marking a 3% improvement over the best open-source baseline, while achieving performance comparable to GPT-4o only on MHS. We show that fine-tuning a single-layer adapter within SAFE-MEME-H outperforms fully fine-tuned models in regular fine-grained hateful meme detection. However, the fully fine-tuning approach with a Q&A setup is more effective for handling confounding cases. We also systematically examine the error cases, offering valuable insights into the robustness and limitations of the proposed structured reasoning framework for analyzing hateful memes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    The paper reports higher harmful-output rates in three open-source VLMs from detailed image descriptions, in-context examples, and positive openings, and from a skip connection between internal layers, with memes riva...

Pith tools