{"work":{"id":"32f5cddf-1446-4477-986d-22ac06fd151f","openalex_id":null,"doi":null,"arxiv_id":"2112.05814","raw_key":null,"title":"Deep ViT Features as Dense Visual Descriptors","authors":null,"authors_text":"S","year":2021,"venue":"cs.CV","abstract":"We study the use of deep features extracted from a pretrained Vision Transformer (ViT) as dense visual descriptors. We observe and empirically demonstrate that such features, when extractedfrom a self-supervised ViT model (DINO-ViT), exhibit several striking properties, including: (i) the features encode powerful, well-localized semantic information, at high spatial granularity, such as object parts; (ii) the encoded semantic information is shared across related, yet different object categories, and (iii) positional bias changes gradually throughout the layers. These properties allow us to design simple methods for a variety of applications, including co-segmentation, part co-segmentation and semantic correspondences. To distill the power of ViT features from convoluted design choices, we restrict ourselves to lightweight zero-shot methodologies (e.g., binning and clustering) applied directly to the features. Since our methods require no additional training nor data, they are readily applicable across a variety of domains. We show by extensive qualitative and quantitative evaluation that our simple methodologies achieve competitive results with recent state-of-the-art supervised methods, and outperform previous unsupervised methods by a large margin. Code is available in dino-vit-features.github.io.","external_url":"https://arxiv.org/abs/2112.05814","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-09T16:56:21.644793+00:00","pith_arxiv_id":"2112.05814","created_at":"2026-05-11T21:11:18.690691+00:00","updated_at":"2026-07-09T16:56:21.644793+00:00","title_quality_ok":true,"display_title":"Deep vit features as dense visual descriptors.arXiv preprint arXiv:2112.05814, 2(3):4","render_title":"Deep vit features as dense visual descriptors.arXiv preprint arXiv:2112.05814, 2(3):4"},"hub":{"state":{"work_id":"32f5cddf-1446-4477-986d-22ac06fd151f","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":14,"external_cited_by_count":null,"distinct_field_count":4,"first_pith_cited_at":"2024-09-03T06:45:22+00:00","last_pith_cited_at":"2026-07-08T10:13:20+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-05T02:34:31.280105+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2}],"polarity_counts":[{"context_polarity":"background","n":1},{"context_polarity":"support","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}