Pith. sign in

REVIEW 1 cited by

RelationVLM: Making Large Vision-Language Models Understand Visual Relations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12801 v1 pith:HQHSO3FC submitted 2024-03-19 cs.CV

classification cs.CV
keywords relationslargelvlmsrelationvlmmodelsvision-languagevisualcapability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual contents and ground text to them. Nonetheless, current LVLMs still struggle to precisely understand visual relations due to the lack of relevant data. In this work, we present RelationVLM, a large vision-language model capable of comprehending various levels and types of relations whether across multiple images or within a video. Specifically, we devise a multi-stage relation-aware training scheme and a series of corresponding data configuration strategies to bestow RelationVLM with the capabilities of understanding semantic relations, temporal associations and geometric transforms. Extensive case studies and quantitative evaluations show RelationVLM has strong capability in understanding such relations and emerges impressive in-context capability of reasoning from few-shot examples by comparison. This work fosters the advancements of LVLMs by enabling them to support a wider range of downstream applications toward artificial general intelligence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues

    cs.HC 2025-06 conditional novelty 5.0 of 10

    ViStruct automatically breaks visualization questions into ordered subtasks tied to highlighted chart regions, imitating expert analysis strategies for chart reading.

Pith tools