Pith. sign in

REVIEW 3 cited by

IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.00936 v1 pith:IBX6D2YR submitted 2025-03-02 cs.CV

classification cs.CV
keywords grad-camprimaryemphasisimageiterativeiterprimemodelreferring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot Referring Image Segmentation (RIS) identifies the instance mask that best aligns with a specified referring expression without training and fine-tuning, significantly reducing the labor-intensive annotation process. Despite achieving commendable results, previous CLIP-based models have a critical drawback: the models exhibit a notable reduction in their capacity to discern relative spatial relationships of objects. This is because they generate all possible masks on an image and evaluate each masked region for similarity to the given expression, often resulting in decreased sensitivity to direct positional clues in text inputs. Moreover, most methods have weak abilities to manage relationships between primary words and their contexts, causing confusion and reduced accuracy in identifying the correct target region. To address these challenges, we propose IteRPrimE (Iterative Grad-CAM Refinement and Primary word Emphasis), which leverages a saliency heatmap through Grad-CAM from a Vision-Language Pre-trained (VLP) model for image-text matching. An iterative Grad-CAM refinement strategy is introduced to progressively enhance the model's focus on the target region and overcome positional insensitivity, creating a self-correcting effect. Additionally, we design the Primary Word Emphasis module to help the model handle complex semantic relations, enhancing its ability to attend to the intended object. Extensive experiments conducted on the RefCOCO/+/g, and PhraseCut benchmarks demonstrate that IteRPrimE outperforms previous state-of-the-art zero-shot methods, particularly excelling in out-of-domain scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark

    cs.CV 2025-06 conditional novelty 6.0 of 10

    NWPU-Refer is a bilingual, high-resolution remote sensing segmentation dataset with multi-object and no-target queries, and MRSNet is a multi-scale network that achieves the best reported scores on it.

  2. SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A SAM2-based framework that uses a fused text-audio-visual token to prompt video segmentation achieves 58.5 J&F on Ref-AVS, outperforming the previous state of the art by 8.5 points.

  3. Image Embedding Sampling Method for Diverse Captioning

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A training-free hierarchical embedding sampling method (HBoP) lets a small BLIP model generate captions as diverse as human ones, beating much larger VLMs on diversity metrics.

Pith tools