Pith. sign in

REVIEW 4 cited by

FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01477 v2 pith:5QYXBWAJ submitted 2023-11-02 cs.CV

classification cs.CV
keywords faithfulnessfaithscorelvlmsatomicfactsfine-grainedhallucinationsimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce FaithScore (Faithfulness to Atomic Image Facts Score), a reference-free and fine-grained evaluation metric that measures the faithfulness of the generated free-form answers from large vision-language models (LVLMs). The FaithScore evaluation first identifies sub-sentences containing descriptive statements that need to be verified, then extracts a comprehensive list of atomic facts from these sub-sentences, and finally conducts consistency verification between fine-grained atomic facts and the input image. Meta-evaluation demonstrates that our metric highly correlates with human judgments of faithfulness. We collect two benchmark datasets (i.e. LLaVA-1k and MSCOCO-Cap) for evaluating LVLMs instruction-following hallucinations. We measure hallucinations in state-of-the-art LVLMs with FaithScore on the datasets. Results reveal that current systems are prone to generate hallucinated content unfaithful to the image, which leaves room for future improvements. We hope our metric FaithScore can help evaluate future LVLMs in terms of faithfulness and provide insightful advice for enhancing LVLMs' faithfulness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARGUS: Hallucination and Omission Evaluation in Video-LLMs

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ARGUS measures hallucination and omission in free-form video captions using LLM-based entailment and temporal alignment, finding that even the best video-LLM still produces roughly 40% hallucinated content.

  2. SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new reference-free metric, SPECS, fine-tunes LongCLIP with a specificity objective and reaches LLM-level human correlation on long captions at a fraction of the computational cost.

  3. Improving Alignment in LVLMs with Debiased Self-Judgment

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A contrastive self-judgment score that subtracts a model's image-free confidence from its visual confidence is used to guide decoding, safety moderation, and DPO training, improving hallucination and safety metrics ac...

  4. Vid2Coach: Transforming How-To Videos into Task Assistants

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Vid2Coach converts how-to videos into a real-time, wearable task assistant that helps blind and low vision people cook with fewer errors.

Pith tools