Pith. sign in

REVIEW 3 cited by

Adapting CLIP For Phrase Localization Without Further Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03647 v1 pith:6HXEEPZ2 submitted 2022-04-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords cliplocalizationphrasesupervisedannotationsmethodsembeddingfeature
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Supervised or weakly supervised methods for phrase localization (textual grounding) either rely on human annotations or some other supervised models, e.g., object detectors. Obtaining these annotations is labor-intensive and may be difficult to scale in practice. We propose to leverage recent advances in contrastive language-vision models, CLIP, pre-trained on image and caption pairs collected from the internet. In its original form, CLIP only outputs an image-level embedding without any spatial resolution. We adapt CLIP to generate high-resolution spatial feature maps. Importantly, we can extract feature maps from both ViT and ResNet CLIP model while maintaining the semantic properties of an image embedding. This provides a natural framework for phrase localization. Our method for phrase localization requires no human annotations or additional training. Extensive experiments show that our method outperforms existing no-training methods in zero-shot phrase localization, and in some cases, it even outperforms supervised methods. Code is available at https://github.com/pals-ttic/adapting-CLIP .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero-shot 2D Grounding with Novel Affordance Types

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new benchmark and pipeline show that grounding an action word like 'cut' on a familiar object can be done without ever training on that action word, outperforming prior affordance grounding methods by a large margin.

  2. FOR: Finetuning for Object Level Open Vocabulary Image Retrieval

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FOR fine-tunes a CLIP encoder with a learnable-query decoder head and a pseudo-label loss, improving open-vocabulary object-level retrieval by up to 8 mAP@50 points over prior art.

  3. FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A training-free method that prompts an LLM to decompose referring expressions into formal components and fuses detector and CLIP scores, boosting zero-shot referring object detection and segmentation.

Pith tools