Pith. sign in

REVIEW 2 cited by

Caption Enriched Samples for Improving Hateful Memes Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.10649 v1 pith:UQ6OPCXN submitted 2021-09-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelshatefulunimodalcaptionimagelanguagemememultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recently introduced hateful meme challenge demonstrates the difficulty of determining whether a meme is hateful or not. Specifically, both unimodal language models and multimodal vision-language models cannot reach the human level of performance. Motivated by the need to model the contrast between the image content and the overlayed text, we suggest applying an off-the-shelf image captioning tool in order to capture the first. We demonstrate that the incorporation of such automatic captions during fine-tuning improves the results for various unimodal and multimodal models. Moreover, in the unimodal case, continuing the pre-training of language models on augmented and original caption pairs, is highly beneficial to the classification accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MAR-12 improves humor and hate detection in memes by prompting a VLM through twelve reasoning perspectives, attention-weighting them, and generating explanations from the weighted evidence.

  2. On VLMs for Diverse Tasks in Multimodal Meme Classification

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A VLM-exclamation-to-LLM distillation pipeline (CoVExFiL) improves meme classification over prompting and LoRA fine-tuning, especially for sentiment.

Pith tools