Pith. sign in

REVIEW 1 cited by

Towards Demystifying Representation Learning with Non-contrastive Self-supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.04947 v2 pith:GH6MP27K submitted 2021-10-11 cs.LG stat.ML

classification cs.LGstat.ML
keywords non-contrastivealgorithmdirectcopydirectpredfeatureslearnlearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Non-contrastive methods of self-supervised learning (such as BYOL and SimSiam) learn representations by minimizing the distance between two views of the same image. These approaches have achieved remarkable performance in practice, but the theoretical understanding lags behind. Tian et al. 2021 explained why the representation does not collapse to zero, however, how the feature is learned still remains mysterious. In our work, we prove in a linear network, non-contrastive methods learn a desirable projection matrix and also reduce the sample complexity on downstream tasks. Our analysis suggests that weight decay acts as an implicit threshold that discards the features with high variance under data augmentations, and keeps the features with low variance. Inspired by our theory, we design a simpler and more computationally efficient algorithm DirectCopy by removing the eigen-decomposition step in the original DirectPred algorithm in Tian et al. 2021. Our experiments show that DirectCopy rivals or even outperforms DirectPred on STL-10, CIFAR-10, CIFAR-100, and ImageNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Clustering via Self-Supervised Diffusion

    cs.AI 2025-07 conditional novelty 6.0 of 10

    CLUDI trains a student to imitate stochastic diffusion-generated cluster assignments on pre-trained DINO image features and averages multiple assignments to cluster images.

Pith tools