Pith. sign in

hub Mixed citations

A Simple Framework for Contrastive Learning of Visual Representations

Mixed citation behavior. Most common role is background (57%).

63 Pith papers citing it
7,335 external citations · Pith
Background 57% of classified citations
abstract

This paper presents SimCLR: a simple framework for contrastive learning of visual representations. We simplify recently proposed contrastive self-supervised learning algorithms without requiring specialized architectures or a memory bank. In order to understand what enables the contrastive prediction tasks to learn useful representations, we systematically study the major components of our framework. We show that (1) composition of data augmentations plays a critical role in defining effective predictive tasks, (2) introducing a learnable nonlinear transformation between the representation and the contrastive loss substantially improves the quality of the learned representations, and (3) contrastive learning benefits from larger batch sizes and more training steps compared to supervised learning. By combining these findings, we are able to considerably outperform previous methods for self-supervised and semi-supervised learning on ImageNet. A linear classifier trained on self-supervised representations learned by SimCLR achieves 76.5% top-1 accuracy, which is a 7% relative improvement over previous state-of-the-art, matching the performance of a supervised ResNet-50. When fine-tuned on only 1% of the labels, we achieve 85.8% top-5 accuracy, outperforming AlexNet with 100X fewer labels.

hub tools

citation-role summary

background 4 method 3

citation-polarity summary

representative citing papers

LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

cs.CV · 2026-07-01 · unverdicted · novelty 7.0

LeVLJEPA is the first non-contrastive vision-language pretraining method that learns via cross-modal prediction without negatives, producing stronger dense features than contrastive baselines on VQA and segmentation tasks.

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning

cs.LG · 2026-05-13 · unverdicted · novelty 7.0

SMA uses a submodular mutual information objective on data sets to deliver competitive zero-shot classification and retrieval performance on CLIP benchmarks with only tens of thousands of samples, orders of magnitude fewer than standard approaches.

Self-Directed Task Identification

cs.LG · 2026-04-02 · unverdicted · novelty 7.0

SDTI lets models identify the correct target variable in datasets in a zero-shot setting using standard neural networks, beating baselines by 14% F1 on synthetic benchmarks.

BEiT: BERT Pre-Training of Image Transformers

cs.CV · 2021-06-15 · conditional · novelty 7.0

BEiT pre-trains vision transformers via masked image modeling on visual tokens and reaches 83.2% ImageNet top-1 accuracy for the base model and 86.3% for the large model using only ImageNet-1K data.

Mastering Atari with Discrete World Models

cs.LG · 2020-10-05 · accept · novelty 7.0

DreamerV2 reaches human-level performance on 55 Atari games by learning behaviors inside a separately trained discrete-latent world model.

TactX: Learning Shared Tactile Representations Across Diverse Sensors

cs.RO · 2026-06-30 · unverdicted · novelty 6.0

TactX learns a shared latent representation across three tactile sensor modalities via joint training on paired contacts, enabling zero-shot policy transfer and higher success on pick-and-place, insertion, wiping, and reorientation tasks.

Interpretable Neural Marked Statistics for Cosmological Inference

astro-ph.CO · 2026-06-09 · unverdicted · novelty 6.0

A neural marking scheme trained with contrastive learning tightens constraints on σ8 by 2.9× and Ωm by 1.8× over classical marks at k_max=0.2 h/Mpc while breaking their degeneracy at the Fisher level.

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

cs.CV · 2026-06-05 · unverdicted · novelty 6.0

DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.

Robust Multi-view Clustering against Imperfect Information

cs.CV · 2026-06-03 · unverdicted · novelty 6.0

PLCI infers posterior distributions over latent cross-view counterparts using instance reliability and prototype transport to jointly address incomplete views and noisy correspondences in multi-view clustering.

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

cs.RO · 2026-05-18 · unverdicted · novelty 6.0

DexHoldem is a new benchmark providing 1,470 teleoperated demonstrations across 14 manipulation primitives, plus standardized tests for dexterous policy execution and agentic perception in a physical Texas Hold'em setting.

citing papers explorer

Showing 50 of 63 citing papers.