Pith. sign in

REVIEW 3 cited by

DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04654 v1 pith:GKCSPZPA submitted 2023-06-06 cs.CV

classification cs.CV
keywords densedensedinopoint-levelself-supervisedcalledconsistencylearningprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but neglected by the existing self-supervised transformers, we introduce point-level supervision across views in a novel token-based way. Specifically, DenseDINO introduces some extra input tokens called reference tokens to match the point-level features with the position prior. With the reference token, the model could maintain spatial consistency and deal with multi-object complex scene images, thus generalizing better on dense prediction tasks. Compared with the vanilla DINO, our approach obtains competitive performance when evaluated on classification in ImageNet and achieves a large margin (+7.2% mIoU) improvement in semantic segmentation on PascalVOC under the linear probing protocol for segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ACE: Anatomically Consistent Embeddings in Composition and Decomposition

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A self-supervised pretraining method that aligns global and local patch embeddings via composition and decomposition improves transfer to medical imaging tasks.

  2. Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A token-level shuffling and mixing training framework that reduces position and content bias improves cross-dataset deepfake detection accuracy.

  3. Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Location-aware training improves seafloor image classification for CNN-style self-supervised models, but a pretrained vision transformer matches the best location-regularised result without any fine-tuning.

Pith tools