Pith. sign in

REVIEW 2 cited by

Learning Visual Grounding from Generative Vision and Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14563 v1 pith:JTRKMG4X submitted 2024-07-18 cs.CV

classification cs.CV
keywords groundingvisualdatagenerativeobjectreferringtaskscapture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text annotation of visual grounding data. We find that grounding knowledge already exists in generative VLM and can be elicited by proper prompting. We thus prompt a VLM to generate object-level descriptions by feeding it object regions from existing object detection datasets. We further propose attribute modeling to explicitly capture the important object attributes, and spatial relation modeling to capture inter-object relationship, both of which are common linguistic pattern in referring expression. Our constructed dataset (500K images, 1M objects, 16M referring expressions) is one of the largest grounding datasets to date, and the first grounding dataset with purely model-generated queries and human-annotated objects. To verify the quality of this data, we conduct zero-shot transfer experiments to the popular RefCOCO benchmarks for both referring expression comprehension (REC) and segmentation (RES) tasks. On both tasks, our model significantly outperform the state-of-the-art approaches without using human annotated visual grounding data. Our results demonstrate the promise of generative VLM to scale up visual grounding in the real world. Code and models will be released.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A text-conditioned U-Net can generate input-aware triggers that backdoor VLM visual grounding, forcing the model to output the attacker-chosen object's bounding box regardless of the user query.

  2. Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A disease-aware prompting method that reweights chest X-ray features using the model's own explainability map improves weakly-supervised visual grounding on three benchmarks.

Pith tools