Pith. sign in

REVIEW 1 cited by

Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11317 v1 pith:W245RMRE submitted 2024-03-17 cs.CL cs.CV

classification cs.CLcs.CV
keywords approachesfew-shotembeddingsimagellmsbettercaptionsdirectly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Two approaches have emerged to input images into large language models (LLMs). The first is to caption images into natural language. The second is to map image feature embeddings into the domain of the LLM and pass the mapped embeddings directly to the LLM. The majority of recent few-shot multimodal work reports performance using architectures that employ variations of one of these two approaches. But they overlook an important comparison between them. We design a controlled and focused experiment to compare these two approaches to few-shot visual question answering (VQA) with LLMs. Our findings indicate that for Flan-T5 XL, a 3B parameter LLM, connecting visual embeddings directly to the LLM embedding space does not guarantee improved performance over using image captions. In the zero-shot regime, we find using textual image captions is better. In the few-shot regimes, how the in-context examples are selected determines which is better.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free, confidence-gated iterative in-context learning framework substantially improves out-of-distribution video understanding in QA, classification, and captioning.

Pith tools