Pith. sign in

REVIEW 3 cited by

MMIDR: Teaching Large Language Model to Interpret Multimodal Misinformation via Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14171 v3 pith:NNUHOM45 submitted 2024-03-21 cs.CL

classification cs.CL
keywords misinformationmultimodalllmsdetectionmmidrdistillationinstruction-followinginterpret
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic detection of multimodal misinformation has gained a widespread attention recently. However, the potential of powerful Large Language Models (LLMs) for multimodal misinformation detection remains underexplored. Besides, how to teach LLMs to interpret multimodal misinformation in cost-effective and accessible way is still an open question. To address that, we propose MMIDR, a framework designed to teach LLMs in providing fluent and high-quality textual explanations for their decision-making process of multimodal misinformation. To convert multimodal misinformation into an appropriate instruction-following format, we present a data augmentation perspective and pipeline. This pipeline consists of a visual information processing module and an evidence retrieval module. Subsequently, we prompt the proprietary LLMs with processed contents to extract rationales for interpreting the authenticity of multimodal misinformation. Furthermore, we design an efficient knowledge distillation approach to distill the capability of proprietary LLMs in explaining multimodal misinformation into open-source LLMs. To explore several research questions regarding the performance of LLMs in multimodal misinformation detection tasks, we construct an instruction-following multimodal misinformation dataset and conduct comprehensive experiments. The experimental findings reveal that our MMIDR exhibits sufficient detection performance and possesses the capacity to provide compelling rationales to support its assessments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Personalized Large Language Models Can Increase the Belief Accuracy of Social Networks

    cs.SI 2025-06 conditional novelty 7.0 of 10

    A pre-registered experiment finds that adding a personalized, factually grounded LLM to an online discussion moves individuals' beliefs toward the truth and makes them build more accurate social networks.

  2. E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    E2LVLM improves out-of-context misinformation detection by having a vision-language model rerank and rewrite retrieved evidence, then fine-tune on LVLM-generated explanations.

  3. Multi-agent Systems for Misinformation Lifecycle : Detection, Correction And Source Identification

    cs.MA 2025-05 reject novelty 3.0 of 10

    A conceptual multi-agent architecture for classifying, detecting, correcting, and sourcing misinformation is proposed but not implemented or evaluated.

Pith tools