REVIEW 14 cited by
Unsupervised Semantic Segmentation by Distilling Feature Correspondences
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clusters. Unlike previous works which achieve this with a single end-to-end framework, we propose to separate feature learning from cluster compactification. Empirically, we show that current unsupervised feature learning frameworks already generate dense features whose correlations are semantically consistent. This observation motivates us to design STEGO ($\textbf{S}$elf-supervised $\textbf{T}$ransformer with $\textbf{E}$nergy-based $\textbf{G}$raph $\textbf{O}$ptimization), a novel framework that distills unsupervised features into high-quality discrete semantic labels. At the core of STEGO is a novel contrastive loss function that encourages features to form compact clusters while preserving their relationships across the corpora. STEGO yields a significant improvement over the prior state of the art, on both the CocoStuff ($\textbf{+14 mIoU}$) and Cityscapes ($\textbf{+9 mIoU}$) semantic segmentation challenges.
Forward citations
Cited by 14 Pith papers
-
Object-level Self-Distillation for Vision Pretraining
ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
LEGO: Leveled Language Gaussian Splatting
LEGO builds view-consistent, multi-level 3D semantic hierarchies from multi-view SAM masks by clustering their physical 3D scales, and grounds them with CLIP for open-vocabulary segmentation and LLM-driven spatial grounding.
-
Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling
One-shot DINOv2 detections of target corals and their co-occurring habitat let a greedy AUV planner sample up to 75% of sparse targets in half the time of exhaustive coverage.
-
Discovering and using Spelke segments
SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.
-
DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF
A NeRF-based method jointly performs unsupervised semantic clustering and CLIP-guided relevancy to discover and segment query-relevant sub-concepts in 3D scenes, with a new Replica benchmark.
-
FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
FlowCut creates pseudo-labels from real videos using DINO features plus optical flow, filters them by IoU matching, and trains a video segmentation model that reaches state-of-the-art on YouTubeVIS and DAVIS.
-
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
A coordinate-based cross-attention transformer, trained with mask-refined and self-distilled pseudo-groundtruth, upsamples VFM features to full resolution and improves downstream task performance.
-
Exploring Temporally-Aware Features for Point Tracking
A DINOv2 backbone augmented with temporal adapters tracks video points accurately using only soft-argmax matching, without iterative refinement.
-
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
CASS achieves 44.4 average mIoU across eight open-vocabulary segmentation benchmarks by injecting DINO's low-rank spectral attention structure into CLIP and adjusting text embeddings with an object presence prior.
-
CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance Segmentation
CutS3D improves unsupervised instance segmentation by cutting semantic masks along 3D depth boundaries rather than 2D image boundaries, gaining about +1 AP over prior state of the art.
-
TACoS: Weakly Supervised Learning of Two-Dimensional Materials from Scribble Annotations to Precise Segmentation
TACoS achieves over 96% of fully supervised segmentation performance on 2D material flakes using less than 0.6% annotated pixels via a unified framework of consistency learning, tree energy regularization, and asymmet...
-
SyncMapV2: Robust and Adaptive Unsupervised Segmentation
SyncMapV2 reports unsupervised segmentation that stays nearly unchanged under common corruptions and across changing input images, using self-organizing dynamics with an untrained echo state network.
-
Temporally Consistent Unsupervised Segmentation for Mobile Robot Perception
Frontier-Seg clusters DINOv2 superpixel features locally per window and globally across a whole video to produce unsupervised, temporally stable terrain segmentations.
Discussion (0). Continue with ORCID to comment.