Pith. sign in

REVIEW 39 cited by

Deep ViT Features as Dense Visual Descriptors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.05814 v3 pith:TT3RICPK submitted 2021-12-10 cs.CV

classification cs.CV
keywords featuresmethodssemanticacrossco-segmentationdeepdensedescriptors
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study the use of deep features extracted from a pretrained Vision Transformer (ViT) as dense visual descriptors. We observe and empirically demonstrate that such features, when extractedfrom a self-supervised ViT model (DINO-ViT), exhibit several striking properties, including: (i) the features encode powerful, well-localized semantic information, at high spatial granularity, such as object parts; (ii) the encoded semantic information is shared across related, yet different object categories, and (iii) positional bias changes gradually throughout the layers. These properties allow us to design simple methods for a variety of applications, including co-segmentation, part co-segmentation and semantic correspondences. To distill the power of ViT features from convoluted design choices, we restrict ourselves to lightweight zero-shot methodologies (e.g., binning and clustering) applied directly to the features. Since our methods require no additional training nor data, they are readily applicable across a variety of domains. We show by extensive qualitative and quantitative evaluation that our simple methodologies achieve competitive results with recent state-of-the-art supervised methods, and outperform previous unsupervised methods by a large margin. Code is available in dino-vit-features.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 39 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weakly-Supervised Learning of Dense Functional Correspondences

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...

  2. Object-level Self-Distillation for Vision Pretraining

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.

  3. Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Adding point-tracking supervision to video diffusion features reduces appearance drift in generated videos while preserving generation quality.

  4. Human-like Object Grouping in Self-supervised Vision Transformers

    cs.CV 2026-03 conditional novelty 6.5 of 10

    DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.

  5. DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    DDMS distills multi-view geometric knowledge from a frozen Depth Anything 3 model into a single-view DINOv2 backbone, improving cross-view feature consistency and discriminability without sacrificing semantic transfer.

  6. Feature-Guided Diffusion for Non-Differentiable Inverse Rendering

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feature-conditioned diffusion model teamed with CMA-ES solves black-box inverse rendering tasks without gradients, beating scalar-loss and gradient-based baselines.

  7. SeeSE3: Emergence of 3D Space in Vision Features

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.

  8. `Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.

  9. Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling

    cs.RO 2026-03 conditional novelty 6.0 of 10

    One-shot DINOv2 detections of target corals and their co-occurring habitat let a greedy AUV planner sample up to 75% of sparse targets in half the time of exhaustive coverage.

  10. UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders

    cs.CV 2026-01 conditional novelty 6.0 of 10

    UPLiFT shows that iterative 2× feature upsampling with a locally-defined attention operator beats cross-attention-based upsamplers on dense prediction while scaling linearly with token count.

  11. Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robotic framework using semantic keypoints plus box-edge contact packs shoe pairs from arbitrary initial states into a standard side-by-side configuration.

  12. Generalizable Object Re-Identification via Visual In-Context Prompting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VICP uses an LLM to generate per-category visual prompts for a frozen DINOv2, enabling few-shot generalization to unseen object categories in re-identification without parameter updates.

  13. Structure-Preserving Medical Image Generation from a Latent Graph Representation

    eess.IV 2025-08 conditional novelty 6.0 of 10

    A latent graph representation of chest X-rays, with a learned topology, is used to generate structure-preserving synthetic images that improve data augmentation for classification and segmentation.

  14. Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A unified RGB-only model jointly performs category-level object detection and 6D pose estimation using neural mesh prototypes and multi-model RANSAC, reporting a 22.9% average improvement on REAL275.

  15. MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MotionShot transfers motion from a reference video to an unseen target object in text-to-video generation by combining semantic and morphological alignment in a training-free pipeline.

  16. ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ADAM combines LLM-generated contextual labels, CLIP embeddings, and nearest-neighbor voting to label novel objects without a predefined class list.

  17. Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Zero-shot robot skill transfer is achieved by grounding task-axis controllers in semantic keypoints matched with SD-DINO across object instances.

  18. Robotic Task Ambiguity Resolution via Natural Language Interaction

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A fine-tuned vision-language model can detect ambiguous robot commands, ask for clarification, and resolve the ambiguity enough to boost real-robot manipulation success from 69.6% to 97.1%.

  19. Exploring Temporally-Aware Features for Point Tracking

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A DINOv2 backbone augmented with temporal adapters tracks video points accurately using only soft-argmax matching, without iterative refinement.

  20. Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A neural surface representation with foreground and background modules performs self-supervised object segmentation from multi-view images, producing finer masks than NeRF-based counterparts.

  21. Few-Shot Adaptation of Training-Free Foundation Model for 3D Medical Image Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Using a handful of labeled slices as memory in SAM2's video-segmentation pipeline enables prompt-free, fine-tuning-free segmentation of 3D medical volumes.

  22. DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.

  23. DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo

    cs.RO 2024-12 conditional novelty 6.0 of 10

    DenseMatcher combines 2D image features with a 3D neural network and functional maps to compute dense semantic correspondences between textured 3D objects, enabling single-demo cross-category robot manipulation.

  24. Distillation of Diffusion Features for Semantic Correspondence

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A DINOv2 student trained with LoRA to imitate DINOv2-plus-SDXL-Turbo similarity maps, then fine-tuned on 3D-derived correspondences, sets new state-of-the-art on three semantic correspondence benchmarks.

  25. DIVE: Taming DINO for Subject-Driven Video Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DIVE uses DINOv2 feature maps as automatic video correspondences to carry source motion, while LoRA adapters carry the target identity.

  26. MAGMA: Manifold Regularization for MAEs

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Adding a manifold regularization loss between intermediate and final transformer layers improves MAE linear probing accuracy, e.g., from 58.0 to 69.0 on ImageNet-100.

  27. Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Finetuning ViT features with SmoothAP on Objaverse multiview correspondences improves 3D correspondence tasks, with meaningful gains even from a single object and a single iteration.

  28. UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    UrbanCAD retrieves a matching CAD model from a single car image, optimizes its materials, and inserts it into reconstructed urban scenes, showing that perception models degrade when the cars are edited into out-of-dis...

  29. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

  30. Preserve, Then Resolve: Many-to-Many Association and Robust Estimation with General-Purpose Visual Features

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    Zero-shot DINOv3 features, many-to-many candidate matching, and Harmonic Consensus Maximization give out-of-domain camera pose accuracy comparable to supervised matchers.

  31. Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A DINO-based coarse-to-fine matcher localizes monocular street photos in OpenStreetMap, reporting 3° orientation recall above the prior method's 5° recall.

  32. From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion

    cs.CV 2026-01 conditional novelty 5.0 of 10

    CLI injects many ViT layers into many LLM layers via LoRA projectors and gated fusion, yielding modest and inconsistent benchmark gains over LLaVA baselines.

  33. Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A lightweight CNN upsampler, distilled from FeatUp features, makes frozen DINOv2 patch features sharp enough for interactive segmentation of micrographs with sparse labels, and its workflow beats fine-tuning a U-Net i...

  34. PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations

    cs.CV 2025-07 conditional novelty 5.0 of 10

    PCR-GS stabilizes pose-free 3D Gaussian Splatting on fast-moving video by aligning DINO semantic features and wavelet high-frequency details between neighboring frames.

  35. Automated Measurement of Eczema Severity with Self-Supervised Learning

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Frozen DINO features plus a 128-unit MLP on SegGPT-segmented eczema regions achieve weighted F1 0.67 in 4-class severity prediction, beating finetuned ResNet-18 and ViT-B on a 528-image in-the-wild dataset.

  36. No Masks Needed: Explainable AI for Deriving Segmentation from Classification

    cs.CV 2025-08 reject novelty 4.0 of 10

    ExplainSeg obtains segmentation masks from classification-only training by converting integrated-gradient heatmaps of a fine-tuned DINO model into masks via NCut or morphology and DenseCRF.

  37. Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Deep Geometric Moments used as diffusion guidance achieve a middle-ground fidelity-diversity trade-off, but the evaluation lacks error bars and a principled balance criterion.

  38. Categorical Keypoint Positional Embedding for Robust Animal Re-Identification

    cs.CV 2024-12 reject novelty 4.0 of 10

    A diffusion-based keypoint propagation plus categorical keypoint positional embedding is reported to improve animal ReID accuracy, but the evidence is limited to a single baseline comparison.

  39. Online Long-term Point Tracking in the Foundation Model Era

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A frame-by-frame point tracker with spatial and context memory reaches accuracy comparable to offline trackers on seven video benchmarks, making online long-term point tracking feasible.

Pith tools