Pith. sign in

REVIEW 16 cited by

A Cookbook of Self-Supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.12210 v2 pith:LRW2ZKR4 submitted 2023-04-24 cs.LG cs.CV

A Cookbook of Self-Supervised Learning

classification cs.LG cs.CV
keywords learningtrainingbarriercookbookentrymethodsself-supervisedadvance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Self-supervised learning, dubbed the dark matter of intelligence, is a promising path to advance machine learning. Yet, much like cooking, training SSL methods is a delicate art with a high barrier to entry. While many components are familiar, successfully training a SSL method involves a dizzying set of choices from the pretext tasks to training hyper-parameters. Our goal is to lower the barrier to entry into SSL research by laying the foundations and latest SSL recipes in the style of a cookbook. We hope to empower the curious researcher to navigate the terrain of methods, understand the role of the various knobs, and gain the know-how required to explore how delicious SSL can be.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learn from your own latents and not from tokens: A sample-complexity theory

    cs.LG 2026-05 unverdicted novelty 7.0

    Latent prediction SSL recovers latent trees from PCFG data with sample complexity constant in hierarchy depth L (up to logs), unlike exponential for token-level or supervised methods.

  2. Optimal Representations for Generalized Contrastive Learning with Imbalanced Datasets

    cs.LG 2026-05 unverdicted novelty 7.0

    In generalized contrastive learning with imbalanced classes, optimal representations collapse to class means whose angular geometry is determined by class proportions via convex optimization, and extreme imbalance cau...

  3. Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach

    cs.CV 2026-06 conditional novelty 6.0

    A two-stage self-supervised protocol using SimCLR on generic then target unlabeled data, followed by generic-label fine-tuning, reaches 97.8% average accuracy for parking occupancy across three public datasets without...

  4. RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems

    cs.IR 2026-04 unverdicted novelty 6.0

    RankUp raises effective rank of representations in deep MetaFormer recommenders via randomized splitting and multi-embeddings, delivering 2-5% GMV gains in production deployments at Weixin.

  5. Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment

    cs.RO 2026-04 unverdicted novelty 6.0

    A contrastive alignment model plus offline preference learning explicitly grounds hierarchical VLA language descriptions to actions and visuals on LanguageTable, achieving performance comparable to fully supervised fi...

  6. Rapidly deploying on-device eye tracking by distilling visual foundation models

    cs.CV 2026-04 unverdicted novelty 6.0

    DistillGaze reduces median gaze error by 58.62% on a 2000+ participant dataset by distilling foundation models into a 256K-parameter on-device model using synthetic labeled data and unlabeled real data.

  7. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

    cs.LG 2025-11 conditional novelty 6.0

    LeJEPA derives an optimal isotropic Gaussian target for embeddings and enforces it via sketched regularization to deliver scalable, heuristics-free self-supervised pretraining with 79% ImageNet linear accuracy on ViT-H/14.

  8. Next-Latent Prediction Transformers Learn Compact World Models

    cs.LG 2025-11 unverdicted novelty 6.0

    NextLat augments next-token prediction with latent next-state prediction, theoretically converging latents to belief states and showing empirical gains in world modeling, reasoning, planning, and faster inference via ...

  9. Self-supervised neural operator for solving partial differential equations

    physics.comp-ph 2025-08 unverdicted novelty 6.0

    Self-supervised neural operator uses Bayesian PINNs to generate training data and a Transformer to learn PDE operators, achieving high accuracy on 1D/2D reaction-diffusion and fluid vibration problems with optional li...

  10. Statistical learnability of smooth boundaries via pairwise binary classification with deep ReLU networks

    math.ST 2025-01 unverdicted novelty 6.0

    Proves learnability of ordered multiple smooth boundaries in pairwise binary classification via localized deep ReLU networks.

  11. Learning General Representation of 12-Lead Electrocardiogram with a Joint-Embedding Predictive Architecture

    cs.LG 2024-10 unverdicted novelty 6.0

    ECG-JEPA applies a joint-embedding predictive architecture with Cross-Pattern Attention to learn semantic representations from unlabeled 12-lead ECG data and reports state-of-the-art results on diagnostic classificati...

  12. Self-Supervised Learning of Plant Image Representations

    cs.CV 2026-04 unverdicted novelty 5.0

    Domain-specific augmentations and plant-only training data produce stronger self-supervised representations for fine-grained plant recognition than standard SSL pipelines or ImageNet pretraining.

  13. Self-Supervised Learning of Plant Image Representations

    cs.CV 2026-04 unverdicted novelty 5.0

    Domain-adapted augmentations and plant-specific training data improve self-supervised representations for fine-grained plant species recognition over standard SSL pipelines.

  14. RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems

    cs.IR 2026-04 unverdicted novelty 5.0

    RankUp enhances representation capacity in deep MetaFormer recommenders via permutation splitting and multi-embeddings, achieving GMV improvements of 2-5% in Weixin production systems.

  15. Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

    cs.CV 2026-05 unverdicted novelty 4.0

    Target-informed self-supervised pretraining via masked image modeling and contrastive learning, plus a confidence-aware infusion head, yields over 6% Dice improvement on unlabeled target-domain POCUS images for pediat...

  16. There Will Be a Scientific Theory of Deep Learning

    stat.ML 2026-04 unverdicted novelty 2.0

    A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universa...