REVIEW 39 cited by
Deep ViT Features as Dense Visual Descriptors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We study the use of deep features extracted from a pretrained Vision Transformer (ViT) as dense visual descriptors. We observe and empirically demonstrate that such features, when extractedfrom a self-supervised ViT model (DINO-ViT), exhibit several striking properties, including: (i) the features encode powerful, well-localized semantic information, at high spatial granularity, such as object parts; (ii) the encoded semantic information is shared across related, yet different object categories, and (iii) positional bias changes gradually throughout the layers. These properties allow us to design simple methods for a variety of applications, including co-segmentation, part co-segmentation and semantic correspondences. To distill the power of ViT features from convoluted design choices, we restrict ourselves to lightweight zero-shot methodologies (e.g., binning and clustering) applied directly to the features. Since our methods require no additional training nor data, they are readily applicable across a variety of domains. We show by extensive qualitative and quantitative evaluation that our simple methodologies achieve competitive results with recent state-of-the-art supervised methods, and outperform previous unsupervised methods by a large margin. Code is available in dino-vit-features.github.io.
Forward citations
Cited by 39 Pith papers
-
Weakly-Supervised Learning of Dense Functional Correspondences
A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...
-
Object-level Self-Distillation for Vision Pretraining
ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.
-
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
Adding point-tracking supervision to video diffusion features reduces appearance drift in generated videos while preserving generation quality.
-
Human-like Object Grouping in Self-supervised Vision Transformers
DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.
-
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
DDMS distills multi-view geometric knowledge from a frozen Depth Anything 3 model into a single-view DINOv2 backbone, improving cross-view feature consistency and discriminability without sacrificing semantic transfer.
-
Feature-Guided Diffusion for Non-Differentiable Inverse Rendering
A feature-conditioned diffusion model teamed with CMA-ES solves black-box inverse rendering tasks without gradients, beating scalar-loss and gradient-based baselines.
-
SeeSE3: Emergence of 3D Space in Vision Features
Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.
-
`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
-
Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling
One-shot DINOv2 detections of target corals and their co-occurring habitat let a greedy AUV planner sample up to 75% of sparse targets in half the time of exhaustive coverage.
-
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
UPLiFT shows that iterative 2× feature upsampling with a locally-defined attention operator beats cross-attention-based upsamplers on dense prediction while scaling linearly with token count.
-
Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness
A robotic framework using semantic keypoints plus box-edge contact packs shoe pairs from arbitrary initial states into a standard side-by-side configuration.
-
Generalizable Object Re-Identification via Visual In-Context Prompting
VICP uses an LLM to generate per-category visual prompts for a frozen DINOv2, enabling few-shot generalization to unseen object categories in re-identification without parameter updates.
-
Structure-Preserving Medical Image Generation from a Latent Graph Representation
A latent graph representation of chest X-rays, with a learned topology, is used to generate structure-preserving synthetic images that improve data augmentation for classification and segmentation.
-
Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
A unified RGB-only model jointly performs category-level object detection and 6D pose estimation using neural mesh prototypes and multi-model RANSAC, reporting a 22.9% average improvement on REAL275.
-
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
MotionShot transfers motion from a reference video to an unseen target object in text-to-video generation by combining semantic and morphological alignment in a training-free pipeline.
-
ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations
ADAM combines LLM-generated contextual labels, CLIP embeddings, and nearest-neighbor voting to label novel objects without a predefined class list.
-
Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models
Zero-shot robot skill transfer is achieved by grounding task-axis controllers in semantic keypoints matched with SD-DINO across object instances.
-
Robotic Task Ambiguity Resolution via Natural Language Interaction
A fine-tuned vision-language model can detect ambiguous robot commands, ask for clarification, and resolve the ambiguity enough to boost real-robot manipulation success from 69.6% to 97.1%.
-
Exploring Temporally-Aware Features for Point Tracking
A DINOv2 backbone augmented with temporal adapters tracks video points accurately using only soft-argmax matching, without iterative refinement.
-
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
A neural surface representation with foreground and background modules performs self-supervised object segmentation from multi-view images, producing finer masks than NeRF-based counterparts.
-
Few-Shot Adaptation of Training-Free Foundation Model for 3D Medical Image Segmentation
Using a handful of labeled slices as memory in SAM2's video-segmentation pipeline enables prompt-free, fine-tuning-free segmentation of 3D medical volumes.
-
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.
-
DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo
DenseMatcher combines 2D image features with a 3D neural network and functional maps to compute dense semantic correspondences between textured 3D objects, enabling single-demo cross-category robot manipulation.
-
Distillation of Diffusion Features for Semantic Correspondence
A DINOv2 student trained with LoRA to imitate DINOv2-plus-SDXL-Turbo similarity maps, then fine-tuned on 3D-derived correspondences, sets new state-of-the-art on three semantic correspondence benchmarks.
-
DIVE: Taming DINO for Subject-Driven Video Editing
DIVE uses DINOv2 feature maps as automatic video correspondences to carry source motion, while LoRA adapters carry the target identity.
-
MAGMA: Manifold Regularization for MAEs
Adding a manifold regularization loss between intermediate and final transformer layers improves MAE linear probing accuracy, e.g., from 58.0 to 69.0 on ImageNet-100.
-
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
Finetuning ViT features with SmoothAP on Objaverse multiview correspondences improves 3D correspondence tasks, with meaningful gains even from a single object and a single iteration.
-
UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation
UrbanCAD retrieves a matching CAD model from a single car image, optimizes its materials, and inserts it into reconstructed urban scenes, showing that perception models degrade when the cars are edited into out-of-dis...
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
Preserve, Then Resolve: Many-to-Many Association and Robust Estimation with General-Purpose Visual Features
Zero-shot DINOv3 features, many-to-many candidate matching, and Harmonic Consensus Maximization give out-of-domain camera pose accuracy comparable to supervised matchers.
-
Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment
A DINO-based coarse-to-fine matcher localizes monocular street photos in OpenStreetMap, reporting 3° orientation recall above the prior method's 5° recall.
-
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
CLI injects many ViT layers into many LLM layers via LoRA projectors and gated fusion, yielding modest and inconsistent benchmark gains over LLaVA baselines.
-
Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
A lightweight CNN upsampler, distilled from FeatUp features, makes frozen DINOv2 patch features sharp enough for interactive segmentation of micrographs with sparse labels, and its workflow beats fine-tuning a U-Net i...
-
PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
PCR-GS stabilizes pose-free 3D Gaussian Splatting on fast-moving video by aligning DINO semantic features and wavelet high-frequency details between neighboring frames.
-
Automated Measurement of Eczema Severity with Self-Supervised Learning
Frozen DINO features plus a 128-unit MLP on SegGPT-segmented eczema regions achieve weighted F1 0.67 in 4-class severity prediction, beating finetuned ResNet-18 and ViT-B on a 528-image in-the-wild dataset.
-
No Masks Needed: Explainable AI for Deriving Segmentation from Classification
ExplainSeg obtains segmentation masks from classification-only training by converting integrated-gradient heatmaps of a fine-tuned DINO model into masks via NCut or morphology and DenseCRF.
-
Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation
Deep Geometric Moments used as diffusion guidance achieve a middle-ground fidelity-diversity trade-off, but the evaluation lacks error bars and a principled balance criterion.
-
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
A diffusion-based keypoint propagation plus categorical keypoint positional embedding is reported to improve animal ReID accuracy, but the evidence is limited to a single baseline comparison.
-
Online Long-term Point Tracking in the Foundation Model Era
A frame-by-frame point tracker with spatial and context memory reaches accuracy comparable to offline trackers on seven video benchmarks, making online long-term point tracking feasible.
Discussion (0). Continue with ORCID to comment.