Pith. sign in

hub Canonical reference

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

Canonical reference. 86% of citing Pith papers cite this work as background.

55 Pith papers citing it
88 external citations · Pith
Background 86% of classified citations
abstract

Recent self-supervised methods for image representation learning are based on maximizing the agreement between embedding vectors from different views of the same image. A trivial solution is obtained when the encoder outputs constant vectors. This collapse problem is often avoided through implicit biases in the learning architecture, that often lack a clear justification or interpretation. In this paper, we introduce VICReg (Variance-Invariance-Covariance Regularization), a method that explicitly avoids the collapse problem with a simple regularization term on the variance of the embeddings along each dimension individually. VICReg combines the variance term with a decorrelation mechanism based on redundancy reduction and covariance regularization, and achieves results on par with the state of the art on several downstream tasks. In addition, we show that incorporating our new variance term into other methods helps stabilize the training and leads to performance improvements.

hub tools

citation-role summary

background 6 method 1

citation-polarity summary

representative citing papers

When Does LeJEPA Learn a World Model?

stat.ML · 2026-05-25 · unverdicted · novelty 8.0

LeJEPA achieves linear identifiability of latent variables uniquely when the latents are Gaussian in worlds with stationary additive-noise transitions.

Spectral Guidance for Flexible and Efficient Control of Diffusion Models

cs.LG · 2026-05-27 · unverdicted · novelty 7.0

Spectral Guidance learns singular functions via self-supervised objective to project guidance signals onto diffusion sampling trajectories, enabling stable control without retraining or backpropagation and improving CIFAR-10 accuracy by 37 points with 4x faster sampling.

Normalizing Trajectory Models

cs.CV · 2026-05-08 · unverdicted · novelty 7.0 · 2 refs

NTM models each generative reverse step as a conditional normalizing flow with a hybrid shallow-deep architecture, enabling exact-likelihood training and strong four-step sampling performance on text-to-image tasks.

Coevolving Representations in Joint Image-Feature Diffusion

cs.CV · 2026-04-19 · unverdicted · novelty 7.0

CoReDi coevolves semantic representations with the diffusion model via a jointly learned linear projection stabilized by stop-gradient, normalization, and regularization, yielding faster convergence and higher sample quality than fixed-representation baselines.

Pattern-Calibrated Multimodal Prediction under Blockwise Missingness

stat.ME · 2026-07-02 · unverdicted · novelty 6.0

MOSAIC learns overlap-aware shared-specific representations, fits a first-stage predictor on overlapping data, and calibrates the gap using target-pattern samples, with non-asymptotic error bounds decomposing overlap size, calibration gap, and representation error.

Group-Equivariant Poincar\'e Convolutional Networks

cs.LG · 2026-07-01 · unverdicted · novelty 6.0

Equivariant Poincaré ResNets combine hyperbolic geometry with C4 and D4 group symmetries via specialized reshaping, permutations, and batch norm to reduce optimization space and speed convergence while staying inside the Poincaré ball.

Real-Time Source-Free Object Detection

cs.CV · 2026-06-30 · unverdicted · novelty 6.0

RT-SFOD adapts dual-head detectors like YOLOv10 for source-free object detection via DHF pseudo-label fusion and MARD loss, delivering 1.4-3.5% mAP gains with 1.3x higher throughput and ~2x fewer parameters than prior SFOD methods.

Overcoming Rank Collapse in Feedback Alignment

cs.LG · 2026-06-09 · unverdicted · novelty 6.0

Feedback alignment in deep networks is limited by low-rank error signals; orthogonal weight updates and activity normalization raise effective rank and boost performance.

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

cs.CV · 2026-06-05 · unverdicted · novelty 6.0

DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.

BRo-JEPA: Learning Modular Arithmetic in Latent Space

cs.LG · 2026-05-31 · unverdicted · novelty 6.0

A block-rotation predictor inside a JEPA model imposes the circular geometry of modular arithmetic on latent representations of MNIST digits and yields strong zero-shot generalization to unseen operations.

Geometry-First Generative Spatial Single-Cell Reconstruction

cs.LG · 2026-05-27 · unverdicted · novelty 6.0

GEARS is a geometry-first generative framework that learns domain-invariant encoders and permutation-equivariant diffusion generators to reconstruct intrinsic 2D cell coordinates and distance matrices from unpaired scRNA-seq guided by ST.

Uncovering the Latent Potential of Deep Intermediate Representations

cs.LG · 2026-05-21 · unverdicted · novelty 6.0

Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains increasing with model depth across modalities.

Predictive but Not Plannable: RC-aux for Latent World Models

cs.LG · 2026-05-08 · unverdicted · novelty 6.0

RC-aux corrects spatiotemporal mismatch in reconstruction-free latent world models by adding multi-horizon prediction and reachability supervision, improving planning performance on goal-conditioned pixel-control tasks.

citing papers explorer

Showing 50 of 55 citing papers.