REVIEW 6 cited by
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive solutions such as intervention are still limited, with current research focusing mostly on text-based content, neglecting the widespread influence of multimodal content like memes. Addressing this gap, we present \textit{MemeGuard}, a comprehensive framework leveraging Large Language Models (LLMs) and Visual Language Models (VLMs) for meme intervention. \textit{MemeGuard} harnesses a specially fine-tuned VLM, \textit{VLMeme}, for meme interpretation, and a multimodal knowledge selection and ranking mechanism (\textit{MKS}) for distilling relevant knowledge. This knowledge is then employed by a general-purpose LLM to generate contextually appropriate interventions. Another key contribution of this work is the \textit{\textbf{I}ntervening} \textit{\textbf{C}yberbullying in \textbf{M}ultimodal \textbf{M}emes (ICMM)} dataset, a high-quality, labeled dataset featuring toxic memes and their corresponding human-annotated interventions. We leverage \textit{ICMM} to test \textit{MemeGuard}, demonstrating its proficiency in generating relevant and effective responses to toxic memes.
Forward citations
Cited by 6 Pith papers
-
Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts
Political podcast toxic segments tend to be longer, more repetitive, more figurative, and angrier than neighboring talk, but some of that pattern is a side effect of how segments were defined.
-
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
A rule-decomposition data pipeline and 246K question-answer pairs let instruction-tuned multimodal LLMs classify and explain image content moderation more accurately than their untuned versions.
-
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
A systematic survey and cross-benchmark evaluation showing that multimodal LLMs can recognize humor artifacts but still struggle to interpret the intended meaning and mechanisms of visual humor.
-
MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
A sticker emotion recognizer that uses an LLM's multi-view text descriptions to guide a pyramid vision transformer beats previous methods on SER30K and MET-MEME.
-
On VLMs for Diverse Tasks in Multimodal Meme Classification
A VLM-exclamation-to-LLM distillation pipeline (CoVExFiL) improves meme classification over prompting and LoRA fine-tuning, especially for sentiment.
-
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.
Discussion (0). Continue with ORCID to comment.