Pith. sign in

REVIEW 3 cited by

Upsampling DINOv2 features for unsupervised vision tasks and weakly supervised materials segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19836 v2 pith:UYXDYTRO submitted 2024-10-20 cs.CV cond-mat.mtrl-scieess.IV

classification cs.CVcond-mat.mtrl-scieess.IV
keywords featuressegmentationmaterialssupervisedweaklyclusteringdinov2like
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The features of self-supervised vision transformers (ViTs) contain strong semantic and positional information relevant to downstream tasks like object localization and segmentation. Recent works combine these features with traditional methods like clustering, graph partitioning or region correlations to achieve impressive baselines without finetuning or training additional networks. We leverage upsampled features from ViT networks (e.g DINOv2) in two workflows: in a clustering based approach for object localization and segmentation, and paired with standard classifiers in weakly supervised materials segmentation. Both show strong performance on benchmarks, especially in weakly supervised segmentation where the ViT features capture complex relationships inaccessible to classical approaches. We expect the flexibility and generalizability of these features will both speed up and strengthen materials characterization, from segmentation to property-prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling

    cs.RO 2026-03 conditional novelty 6.0 of 10

    One-shot DINOv2 detections of target corals and their co-occurring habitat let a greedy AUV planner sample up to 75% of sparse targets in half the time of exhaustive coverage.

  2. Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark

    cs.MM 2024-12 conditional novelty 6.0 of 10

    A public benchmark and unified test conditions for compressing intermediate features of large models, with two image-codec baselines evaluated.

  3. Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A lightweight CNN upsampler, distilled from FeatUp features, makes frozen DINOv2 patch features sharp enough for interactive segmentation of micrographs with sparse labels, and its workflow beats fine-tuning a U-Net i...

Pith tools