Pith. sign in

REVIEW 9 cited by

GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01523 v4 pith:D4SR4I6K submitted 2024-01-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords abusegoat-benchlmmsmemesimplicitmodelsmultimodalsocial
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The exponential growth of social media has profoundly transformed how information is created, disseminated, and absorbed, exceeding any precedent in the digital age. Regrettably, this explosion has also spawned a significant increase in the online abuse of memes. Evaluating the negative impact of memes is notably challenging, owing to their often subtle and implicit meanings, which are not directly conveyed through the overt text and image. In light of this, large multimodal models (LMMs) have emerged as a focal point of interest due to their remarkable capabilities in handling diverse multimodal tasks. In response to this development, our paper aims to thoroughly examine the capacity of various LMMs (e.g., GPT-4o) to discern and respond to the nuanced aspects of social abuse manifested in memes. We introduce the comprehensive meme benchmark, GOAT-Bench, comprising over 6K varied memes encapsulating themes such as implicit hate speech, sexism, and cyberbullying, etc. Utilizing GOAT-Bench, we delve into the ability of LMMs to accurately assess hatefulness, misogyny, offensiveness, sarcasm, and harmful content. Our extensive experiments across a range of LMMs reveal that current models still exhibit a deficiency in safety awareness, showing insensitivity to various forms of implicit abuse. We posit that this shortfall represents a critical impediment to the realization of safe artificial intelligence. The GOAT-Bench and accompanying resources are publicly accessible at https://goatlmm.github.io/, contributing to ongoing research in this vital field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models

    cs.CL 2025-05 conditional novelty 7.0 of 10

    The authors release a large multimodal benchmark showing that current LMMs struggle to detect toxicity that emerges only from combining image and text, and that many-shot toxic demonstrations further reduce their accuracy.

  2. MSTS: A Multimodal Safety Test Suite for Vision-Language Models

    cs.CL 2025-01 conditional novelty 7.0 of 10

    MSTS is a 400-prompt multimodal safety benchmark showing that open vision-language models give unsafe responses to up to 14% of prompts, and that image-plus-text inputs trigger more unsafe answers than text alone.

  3. Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers

    cs.CR 2024-12 conditional novelty 7.0 of 10

    Injecting two small dots at the end of meme text can plant a backdoor in hateful meme detectors, making them classify triggered hateful memes as benign under automatic OCR pipelines.

  4. Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Vision-language models consistently recognize unsafe content better from text than from images, and a simplified reinforcement learning fine-tune narrows that gap.

  5. MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A new 1,565-item benchmark shows large vision-language models classify meme-context relationships far better than they infer the poster's intent.

  6. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

    cs.MM 2025-05 conditional novelty 6.0 of 10

    ShieldVLM detects multimodal implicit toxicity through deliberate cross-modal reasoning, outperforming existing moderation APIs and models on the new MMIT benchmark.

  7. ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A new 305-question benchmark uses Scratch block programs to test multimodal AI visual programming reasoning, and all ten tested models fall below 70% accuracy.

  8. STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A new dataset and an entropy-guided multi-task model show that collaborative tags improve fine-grained toxic meme detection.

  9. CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A multi-stage multimodal RAG pipeline with intent refinement, API selection, and relevance/safety filtering reports higher win rates than baseline MLLMs and near-oracle accuracy on Encyclopedic VQA.

Pith tools