Pith. sign in

REVIEW 6 cited by

GLIPv2: Unifying Localization and Vision-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.05836 v2 pith:GXNI7QLD submitted 2022-06-12 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords tasksunderstandinglocalizationglipv2modeldetectionpre-trainingvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies localization pre-training and Vision-Language Pre-training (VLP) with three pre-training tasks: phrase grounding as a VL reformulation of the detection task, region-word contrastive learning as a novel region-word level contrastive learning task, and the masked language modeling. This unification not only simplifies the previous multi-stage VLP procedure but also achieves mutual benefits between localization and understanding tasks. Experimental results show that a single GLIPv2 model (all model weights are shared) achieves near SoTA performance on various localization and understanding tasks. The model also shows (1) strong zero-shot and few-shot adaption performance on open-vocabulary object detection tasks and (2) superior grounding capability on VL understanding tasks. Code will be released at https://github.com/microsoft/GLIP.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Vision Harnessing Agent for Open Ad-hoc Segmentation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.

  2. Hi-TOPS: Hierarchical Topology-aware Scoring Prior for 3D Part Decomposition

    cs.GR 2026-08 conditional novelty 6.0 of 10

    Hi-TOPS, a training-free meso-scale Flow-Freeze prior with TSDF-guided superquadric fitting, achieves competitive 3D part decomposition and the best mIoU on PartNet (55.87).

  3. LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    LaV-CoT introduces a multi-stage visual CoT pipeline and GRPO training with language-consistency rewards, delivering up to 9.5% accuracy gains on multilingual VQA benchmarks over similar-sized open models.

  4. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

    cs.CV 2024-01 unverdicted novelty 6.0 of 10

    Grounded SAM integrates Grounding DINO and SAM to support text-prompted open-world detection and segmentation, achieving 48.7 mean AP on SegInW zero-shot with the base detector and huge segmenter.

  5. FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A two-stage framework with visual support constraints and GRPO reinforcement learning enables a 7B model to outperform up to 72B-parameter models on category-agnostic in-context object localization.

  6. Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Textual Inversion is applied to open-vocabulary object detectors to learn a few new tokens while keeping the VLM frozen and preserving zero-shot abilities.

Pith tools