Pith. sign in

REVIEW 3 cited by

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.03354 v1 pith:NOQTIA3H submitted 2023-11-06 cs.CV

classification cs.CV
keywords visualcommunicationentitiescompositionaldetectiongeneratedlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundation models (VLMs) fall short of such compositional abilities due to their "bag-of-words" behaviors and inability to construct words that correctly represent visual entities and the relations among the entities. To this end, we propose CoVLM, which can guide the LLM to explicitly compose visual entities and relationships among the text and dynamically communicate with the vision encoder and detection network to achieve vision-language communicative decoding. Specifically, we first devise a set of novel communication tokens for the LLM, for dynamic communication between the visual detection system and the language system. A communication token is generated by the LLM following a visual entity or a relation, to inform the detection network to propose regions that are relevant to the sentence generated so far. The proposed regions-of-interests (ROIs) are then fed back into the LLM for better language generation contingent on the relevant regions. The LLM is thus able to compose the visual entities and relationships through the communication tokens. The vision-to-language and language-to-vision communication are iteratively performed until the entire sentence is generated. Our framework seamlessly bridges the gap between visual perception and LLMs and outperforms previous VLMs by a large margin on compositional reasoning benchmarks (e.g., ~20% in HICO-DET mAP, ~14% in Cola top-1 accuracy, and ~3% on ARO top-1 accuracy). We also achieve state-of-the-art performances on traditional vision-language tasks such as referring expression comprehension and visual question answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training and inference framework that decomposes visual queries into nested phrases and grounds them progressively outperforms baseline LVLMs on compositional grounding and reasoning benchmarks.

  2. Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new GPT-4-generated dataset of distractors and corrective feedback for visual commonsense reasoning, plus a compact LMM (PEIFG) that produces explainable corrections and beats existing baselines in automatic and hum...

  3. EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A multimodal LLM with a SAM-based mask decoder localizes diffusion-edited regions better than traditional forensic methods on MagicBrush, CocoGLIDE, and a new BrushNet dataset.

Pith tools