Pith. sign in

REVIEW 2 cited by

Multimodal Large Language Models to Support Real-World Fact-Checking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03627 v2 pith:WOSJYFFB submitted 2024-03-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsmultimodalfact-checkingmllmsreal-worldinformationknowledgelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) carry the potential to support humans in processing vast amounts of information. While MLLMs are already being used as a fact-checking tool, their abilities and limitations in this regard are understudied. Here is aim to bridge this gap. In particular, we propose a framework for systematically assessing the capacity of current multimodal models to facilitate real-world fact-checking. Our methodology is evidence-free, leveraging only these models' intrinsic knowledge and reasoning capabilities. By designing prompts that extract models' predictions, explanations, and confidence levels, we delve into research questions concerning model accuracy, robustness, and reasons for failure. We empirically find that (1) GPT-4V exhibits superior performance in identifying malicious and misleading multimodal claims, with the ability to explain the unreasonable aspects and underlying motives, and (2) existing open-source models exhibit strong biases and are highly sensitive to the prompt. Our study offers insights into combating false multimodal information and building secure, trustworthy multimodal models. To the best of our knowledge, we are the first to evaluate MLLMs for real-world fact-checking.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A zero-shot multimodal fact-checking pipeline with dynamic web, image, and geolocation tool use reports state-of-the-art accuracy on AVeriTeC, MOCHEG, and VERITE, plus a new post-cutoff benchmark where it beats GPT-4o...

  2. Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A probing classifier trained on VLM, text-encoder, and image-encoder embeddings shows that separate text and image embeddings usually outperform intrinsically fused VLM embeddings for multimodal fact-checking.

Pith tools