The subspace intervention framework reveals that pre-training objectives shape how ViTs encode geometric information in compressible low-rank subspaces, with peak precision at intermediate layers.
hub
Deep vit features as dense visual descriptors.arXiv preprint arXiv:2112.05814, 2(3):4
14 Pith papers cite this work. Polarity classification is still indexing.
abstract
We study the use of deep features extracted from a pretrained Vision Transformer (ViT) as dense visual descriptors. We observe and empirically demonstrate that such features, when extractedfrom a self-supervised ViT model (DINO-ViT), exhibit several striking properties, including: (i) the features encode powerful, well-localized semantic information, at high spatial granularity, such as object parts; (ii) the encoded semantic information is shared across related, yet different object categories, and (iii) positional bias changes gradually throughout the layers. These properties allow us to design simple methods for a variety of applications, including co-segmentation, part co-segmentation and semantic correspondences. To distill the power of ViT features from convoluted design choices, we restrict ourselves to lightweight zero-shot methodologies (e.g., binning and clustering) applied directly to the features. Since our methods require no additional training nor data, they are readily applicable across a variety of domains. We show by extensive qualitative and quantitative evaluation that our simple methodologies achieve competitive results with recent state-of-the-art supervised methods, and outperform previous unsupervised methods by a large margin. Code is available in dino-vit-features.github.io.
hub tools
citation-role summary
citation-polarity summary
roles
background 2representative citing papers
DRoRAE adaptively fuses multi-layer features from vision encoders via energy-constrained routing to enrich visual tokens, cutting rFID from 0.57 to 0.29 and generation FID from 1.74 to 1.65 on ImageNet-256 while revealing a log-linear scaling law with fusion capacity.
ReKep encodes robotic tasks as optimizable Python functions over 3D keypoints that are generated automatically from language and RGB-D input, enabling real-time hierarchical planning on single- and dual-arm platforms without task-specific data.
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
FROST performs training-free few-shot segmentation on remote-sensing imagery by nonparametric density-ratio classification on frozen DINOv3 features and reports 5.6 mIoU gains from one example across 17 benchmarks.
TUDSR applies a twice-upsampling diffusion strategy with chunk-based training to achieve state-of-the-art super-resolution at 1024^2 and 2048^2 resolutions using a one-step GAN on SD2.1-base.
Matching in semantic SSL feature space via Sinkhorn divergence enables effective one-step generation on ImageNet by inducing compact geometry for distribution matching, with training and evaluation features best kept distinct.
PaintCopilot models painting as an open-ended autoregressive process that predicts coherent brushstrokes from partial canvas observations using a ViT target predictor, flow-matching stroke generator, and VAE region sampler.
VISION-SLS learns visual features with state-dependent error bounds and optimizes causal affine output-feedback policies via system level synthesis to achieve safe nonlinear control from RGB images.
Realiz3D decouples visual domain from 3D controls in diffusion models via domain-aware residual adapters to enable photorealistic controllable generation.
Casper3D is a backbone-agnostic variational model that infers stable 3D semantics from noisy 2D embeddings by treating them as observations of a latent 3D state and training via held-out viewpoint prediction.
A visualization protocol using unsupervised semantic segmentation outputs reveals positional biases, scaling behaviors, and boundary artifacts in self-supervised ViTs and distinguishes them from locality bias.
Register tokens improve pixel-space Diffusion Transformers by cleaning high-noise feature maps, and Register Guidance amplifies that effect.
Zero-shot DINOv3 features, many-to-many candidate matching, and Harmonic Consensus Maximization give out-of-domain camera pose accuracy comparable to supervised matchers.
citing papers explorer
-
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
The subspace intervention framework reveals that pre-training objectives shape how ViTs encode geometric information in compressible low-rank subspaces, with peak precision at intermediate layers.
-
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
DRoRAE adaptively fuses multi-layer features from vision encoders via energy-constrained routing to enrich visual tokens, cutting rFID from 0.57 to 0.29 and generation FID from 1.74 to 1.65 on ImageNet-256 while revealing a log-linear scaling law with fusion capacity.
-
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
ReKep encodes robotic tasks as optimizable Python functions over 3D keypoints that are generated automatically from language and RGB-D input, enabling real-time hierarchical planning on single- and dual-arm platforms without task-specific data.
-
`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
-
FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics
FROST performs training-free few-shot segmentation on remote-sensing imagery by nonparametric density-ratio classification on frozen DINOv3 features and reports 5.6 mIoU gains from one example across 17 benchmarks.
-
TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution
TUDSR applies a twice-upsampling diffusion strategy with chunk-based training to achieve state-of-the-art super-resolution at 1024^2 and 2048^2 resolutions using a one-step GAN on SD2.1-base.
-
Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation
Matching in semantic SSL feature space via Sinkhorn divergence enables effective one-step generation on ImageNet by inducing compact geometry for distribution matching, with training and evaluation features best kept distinct.
-
PaintCopilot: Modeling Painting as Autonomous Artistic Continuation
PaintCopilot models painting as an open-ended autoregressive process that predicts coherent brushstrokes from partial canvas observations using a ViT target predictor, flow-matching stroke generator, and VAE region sampler.
-
VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis
VISION-SLS learns visual features with state-dependent error bounds and optimizes causal affine output-feedback policies via system level synthesis to achieve safe nonlinear control from RGB images.
-
Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
Realiz3D decouples visual domain from 3D controls in diffusion models via domain-aware residual adapters to enable photorealistic controllable generation.
-
Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models
Casper3D is a backbone-agnostic variational model that infers stable 3D semantics from noisy 2D embeddings by treating them as observations of a latent 3D state and training via held-out viewpoint prediction.
-
Unsupervised Semantic Segmentation Facilitates Model Understanding
A visualization protocol using unsupervised semantic segmentation outputs reveals positional biases, scaling behaviors, and boundary artifacts in self-supervised ViTs and distinguishes them from locality bias.
-
Registers Matter for Pixel-Space Diffusion Transformers
Register tokens improve pixel-space Diffusion Transformers by cleaning high-noise feature maps, and Register Guidance amplifies that effect.
-
Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association
Zero-shot DINOv3 features, many-to-many candidate matching, and Harmonic Consensus Maximization give out-of-domain camera pose accuracy comparable to supervised matchers.