REVIEW 3 cited by
DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but neglected by the existing self-supervised transformers, we introduce point-level supervision across views in a novel token-based way. Specifically, DenseDINO introduces some extra input tokens called reference tokens to match the point-level features with the position prior. With the reference token, the model could maintain spatial consistency and deal with multi-object complex scene images, thus generalizing better on dense prediction tasks. Compared with the vanilla DINO, our approach obtains competitive performance when evaluated on classification in ImageNet and achieves a large margin (+7.2% mIoU) improvement in semantic segmentation on PascalVOC under the linear probing protocol for segmentation.
Forward citations
Cited by 3 Pith papers
-
ACE: Anatomically Consistent Embeddings in Composition and Decomposition
A self-supervised pretraining method that aligns global and local patch embeddings via composition and decomposition improves transfer to medical imaging tasks.
-
Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing
A token-level shuffling and mixing training framework that reduces position and content bias improves cross-dataset deepfake detection accuracy.
-
Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
Location-aware training improves seafloor image classification for CNN-style self-supervised models, but a pretrained vision transformer matches the best location-regularised result without any fine-tuning.
Discussion (0). Continue with ORCID to comment.