Pith. sign in

REVIEW 6 cited by

MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05344 v1 pith:7NR5ZJ24 submitted 2024-06-08 cs.CL

classification cs.CL
keywords textitcontentmemeguardmemestextbfinterventionknowledgememe
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive solutions such as intervention are still limited, with current research focusing mostly on text-based content, neglecting the widespread influence of multimodal content like memes. Addressing this gap, we present \textit{MemeGuard}, a comprehensive framework leveraging Large Language Models (LLMs) and Visual Language Models (VLMs) for meme intervention. \textit{MemeGuard} harnesses a specially fine-tuned VLM, \textit{VLMeme}, for meme interpretation, and a multimodal knowledge selection and ranking mechanism (\textit{MKS}) for distilling relevant knowledge. This knowledge is then employed by a general-purpose LLM to generate contextually appropriate interventions. Another key contribution of this work is the \textit{\textbf{I}ntervening} \textit{\textbf{C}yberbullying in \textbf{M}ultimodal \textbf{M}emes (ICMM)} dataset, a high-quality, labeled dataset featuring toxic memes and their corresponding human-annotated interventions. We leverage \textit{ICMM} to test \textit{MemeGuard}, demonstrating its proficiency in generating relevant and effective responses to toxic memes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Political podcast toxic segments tend to be longer, more repetitive, more figurative, and angrier than neighboring talk, but some of that pattern is a side effect of how segments were defined.

  2. ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A rule-decomposition data pipeline and 246K question-answer pairs let instruction-tuned multimodal LLMs classify and explain image content moderation more accurately than their untuned versions.

  3. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A systematic survey and cross-benchmark evaluation showing that multimodal LLMs can recognize humor artifacts but still struggle to interpret the intended meaning and mechanisms of visual humor.

  4. MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A sticker emotion recognizer that uses an LLM's multi-view text descriptions to guide a pyramid vision transformer beats previous methods on SER30K and MET-MEME.

  5. On VLMs for Diverse Tasks in Multimodal Meme Classification

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A VLM-exclamation-to-LLM distillation pipeline (CoVExFiL) improves meme classification over prompting and LoRA fine-tuning, especially for sentiment.

  6. Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.

Pith tools