REVIEW 10 cited by
Localizing Objects with Self-Supervised Transformers and no Labels
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Localizing objects in image collections without supervision can help to avoid expensive annotation campaigns. We propose a simple approach to this problem, that leverages the activation features of a vision transformer pre-trained in a self-supervised manner. Our method, LOST, does not require any external object proposal nor any exploration of the image collection; it operates on a single image. Yet, we outperform state-of-the-art object discovery methods by up to 8 CorLoc points on PASCAL VOC 2012. We also show that training a class-agnostic detector on the discovered objects boosts results by another 7 points. Moreover, we show promising results on the unsupervised object discovery task. The code to reproduce our results can be found at https://github.com/valeoai/LOST.
Forward citations
Cited by 10 Pith papers
-
Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive Segmentation
By counting Markov-chain hitting times in Stable Diffusion attention, M2N2 performs training-free interactive segmentation that outperforms trained unsupervised methods on three of four benchmarks.
-
Human-like Object Grouping in Self-supervised Vision Transformers
DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.
-
`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
-
Discovering and using Spelke segments
SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.
-
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
A dual-branch CLIP training pipeline with regional prompts and hierarchical feature alignment reaches state-of-the-art on long- and short-text retrieval.
-
Cut out and Replay: A Simple yet Versatile Strategy for Multi-Label Online Continual Learning
CUTER replays cropped label-specific object regions instead of whole multi-label images, and regularizes patch-feature graphs to keep the cropping ability alive, improving multi-label online continual learning across ...
-
Few-Shot Adaptation of Training-Free Foundation Model for 3D Medical Image Segmentation
Using a handful of labeled slices as memory in SAM2's video-segmentation pipeline enables prompt-free, fine-tuning-free segmentation of 3D medical volumes.
-
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
A 10-step ODE text-to-vision flow with tangent-space projection outperforms single-step MLP projection for open-vocabulary semantic segmentation, reaching 49.5 average mIoU without mask refinement.
-
Mixing Configurations for Downstream Prediction
Mixing multiple resolution clusterings of an embedding, aligned between train and test, and fused by attention, improves downstream regression and classification over single-resolution baselines.
-
Object-Centric Cropping for Visual Few-Shot Classification
The supplied full text (arXiv:2508.00225) is a different paper from the claimed metadata (arXiv:2508.00218), so no claim about few-shot classification is backed by the manuscript.
Discussion (0). Continue with ORCID to comment.