Pith. sign in

Sclip: Rethinking self- attention for dense vision-language inference

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it
abstract

Recent advances in contrastive language-image pretraining (CLIP) have demonstrated strong capabilities in zero-shot classification by aligning visual representations with target text embeddings in an image level. However, in dense prediction tasks, CLIP often struggles to localize visual features within an image and fails to give accurate pixel-level predictions, which prevents it from functioning as a generalized visual foundation model. In this work, we aim to enhance CLIP's potential for semantic segmentation with minimal modifications to its pretrained models. By rethinking self-attention, we surprisingly find that CLIP can adapt to dense prediction tasks by simply introducing a novel Correlative Self-Attention (CSA) mechanism. Specifically, we replace the traditional self-attention block of CLIP vision encoder's last layer by our CSA module and reuse its pretrained projection matrices of query, key, and value, leading to a training-free adaptation approach for CLIP's zero-shot semantic segmentation. Extensive experiments show the advantage of CSA: we obtain a 38.2% average zero-shot mIoU across eight semantic segmentation benchmarks highlighted in this paper, significantly outperforming the existing SoTA's 33.9% and the vanilla CLIP's 14.1%.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 8

roles

background 1

polarities

support 1

representative citing papers

Vision Harnessing Agent for Open Ad-hoc Segmentation

cs.CV · 2026-05-19 · unverdicted · novelty 7.0

VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.

Best Segmentation Buddies for Image-Shape Correspondence

cs.CV · 2026-05-18 · unverdicted · novelty 7.0

The work defines Best Segmentation Buddies as vertices on a 3D shape whose nearest image pixel under distilled features falls inside a given 2D segment, then uses the same features to segment the shape in 3D.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

cs.CV · 2026-07-08 · conditional · novelty 5.5

Inference-time α-entmax sparsification of CLIP’s final self-attention denoises diffuse mass and lifts dense open-vocabulary segmentation and region retrieval in proportion to baseline off-class spread.

LARE: Low-Attention Region Encoding for Text-Image Retrieval

cs.CV · 2026-06-17 · unverdicted · novelty 5.0

LARE uses parallel encoding of full images and low-attention regions to improve text-image retrieval, shown on a new Dense-Set subset of COCO and Flickr30K with re-captioned overlooked areas.

TeD-Loc: Text Distillation for Weakly Supervised Object Localization

cs.CV · 2025-01-22 · unverdicted · novelty 5.0

TeD-Loc improves weakly supervised object localization by distilling CLIP text embeddings to patch embeddings through contrastive alignment plus a localization-guided classifier and QR orthogonalization of text embeddings.

citing papers explorer

Showing 8 of 8 citing papers.