Pith. sign in

hub

Deep vit features as dense visual descriptors.arXiv preprint arXiv:2112.05814, 2(3):4

14 Pith papers cite this work. Polarity classification is still indexing.

14 Pith papers citing it
abstract

We study the use of deep features extracted from a pretrained Vision Transformer (ViT) as dense visual descriptors. We observe and empirically demonstrate that such features, when extractedfrom a self-supervised ViT model (DINO-ViT), exhibit several striking properties, including: (i) the features encode powerful, well-localized semantic information, at high spatial granularity, such as object parts; (ii) the encoded semantic information is shared across related, yet different object categories, and (iii) positional bias changes gradually throughout the layers. These properties allow us to design simple methods for a variety of applications, including co-segmentation, part co-segmentation and semantic correspondences. To distill the power of ViT features from convoluted design choices, we restrict ourselves to lightweight zero-shot methodologies (e.g., binning and clustering) applied directly to the features. Since our methods require no additional training nor data, they are readily applicable across a variety of domains. We show by extensive qualitative and quantitative evaluation that our simple methodologies achieve competitive results with recent state-of-the-art supervised methods, and outperform previous unsupervised methods by a large margin. Code is available in dino-vit-features.github.io.

hub tools

citation-role summary

background 2

citation-polarity summary

years

2026 13 2024 1

roles

background 2

polarities

background 1 support 1

representative citing papers

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization

cs.CV · 2026-05-11 · unverdicted · novelty 7.0 · 2 refs

DRoRAE adaptively fuses multi-layer features from vision encoders via energy-constrained routing to enrich visual tokens, cutting rFID from 0.57 to 0.29 and generation FID from 1.74 to 1.65 on ImageNet-256 while revealing a log-linear scaling law with fusion capacity.

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution

cs.CV · 2026-06-08 · unverdicted · novelty 6.0

TUDSR applies a twice-upsampling diffusion strategy with chunk-based training to achieve state-of-the-art super-resolution at 1024^2 and 2048^2 resolutions using a one-step GAN on SD2.1-base.

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation

cs.CV · 2026-05-20 · unverdicted · novelty 6.0

PaintCopilot models painting as an open-ended autoregressive process that predicts coherent brushstrokes from partial canvas observations using a ViT target predictor, flow-matching stroke generator, and VAE region sampler.

Unsupervised Semantic Segmentation Facilitates Model Understanding

cs.CV · 2026-05-28 · unverdicted · novelty 5.0 · 2 refs

A visualization protocol using unsupervised semantic segmentation outputs reveals positional biases, scaling behaviors, and boundary artifacts in self-supervised ViTs and distinguishes them from locality bias.

citing papers explorer

Showing 14 of 14 citing papers.