Pith. sign in

REVIEW 1 cited by

Generating Faithful and Salient Text from Multimodal Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03961 v1 pith:OJWWXOWS submitted 2024-09-06 cs.CV

classification cs.CV
keywords datasalientfeaturesmultimodaltextcriticfaithfulframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large multimodal models (LMMs) have obtained strong performance on many multimodal tasks, they may still hallucinate while generating text. Their performance on detecting salient features from visual data is also unclear. In this paper, we develop a framework to generate faithful and salient text from mixed-modal data, which includes images and structured data ( represented in knowledge graphs or tables). Specifically, we train a small vision critic model to identify hallucinated and non-salient features from the image modality. The critic model also generates a list of salient image features. This information is used in the post editing step to improve the generation quality. Experiments on two datasets show that our framework improves LMMs' generation quality on both faithfulness and saliency, outperforming recent techniques aimed at reducing hallucination.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MEGL: Multimodal Explanation-Guided Learning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multimodal explanation-guided learning framework that jointly uses visual saliency maps and textual rationales to train image classifiers, improving accuracy, visual explanation overlap, and text explanation scores ...

Pith tools