{"work":{"id":"7f8eb5a6-be08-4e4f-ada8-50f9cfe9f314","openalex_id":null,"doi":null,"arxiv_id":"2310.01403","raw_key":null,"title":"CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction","authors":null,"authors_text":"Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Xiangtai Li, Wentao Liu, and Chen Change Loy","year":2023,"venue":"cs.CV","abstract":"Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP models, particularly those incorporating vision transformers (ViTs), have exhibited remarkable generalization ability in zero-shot image classification. However, when transferring the vision-language alignment of CLIP from global image representation to local region representation for the open-vocabulary dense prediction tasks, CLIP ViTs suffer from the domain shift from full images to local image regions. In this paper, we embark on an in-depth analysis of the region-language alignment in CLIP models, which is essential for downstream open-vocabulary dense prediction tasks. Subsequently, we propose an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs. CLIPSelf empowers ViTs to distill itself by aligning a region representation extracted from its dense feature map with the image-level representation of the corresponding image crop. With the enhanced CLIP ViTs, we achieve new state-of-the-art performance on open-vocabulary object detection, semantic segmentation, and panoptic segmentation across various benchmarks. Models and code are released at https://github.com/wusize/CLIPSelf.","external_url":"https://arxiv.org/abs/2310.01403","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-09T19:36:29.291166+00:00","pith_arxiv_id":"2310.01403","created_at":"2026-05-11T05:51:08.489546+00:00","updated_at":"2026-07-09T19:36:29.291166+00:00","title_quality_ok":true,"display_title":"Clipself: Vision transformer distills itself for open-vocabulary dense prediction","render_title":"Clipself: Vision transformer distills itself for open-vocabulary dense prediction"},"hub":{"state":{"work_id":"7f8eb5a6-be08-4e4f-ada8-50f9cfe9f314","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":13,"external_cited_by_count":null,"distinct_field_count":2,"first_pith_cited_at":"2025-02-26T04:50:20+00:00","last_pith_cited_at":"2026-06-29T14:19:20+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-07-03T12:24:15.736555+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2}],"polarity_counts":[{"context_polarity":"background","n":2}],"runs":{},"summary":{},"graph":{},"authors":[]}}