Pith. sign in

REVIEW 3 cited by

Contrastive Multiview Coding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.05849 v5 pith:ZCXIWH2Z submitted 2019-06-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords viewsapproachcontrastiverepresentationchannelfactorsheardhypothesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans view the world through many sensory channels, e.g., the long-wavelength light channel, viewed by the left eye, or the high-frequency vibrations channel, heard by the right ear. Each view is noisy and incomplete, but important factors, such as physics, geometry, and semantics, tend to be shared between all views (e.g., a "dog" can be seen, heard, and felt). We investigate the classic hypothesis that a powerful representation is one that models view-invariant factors. We study this hypothesis under the framework of multiview contrastive learning, where we learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact. Our approach scales to any number of views, and is view-agnostic. We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the resulting representation captures underlying scene semantics. Our approach achieves state-of-the-art results on image and video unsupervised learning benchmarks. Code is released at: http://github.com/HobbitLong/CMC/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field

    q-bio.NC 2026-07 conditional novelty 6.0 of 10

    Self-supervised networks trained on fovea-only vs periphery-only egocentric videos develop different task and neural-alignment profiles, with periphery-trained models better matching scene-selective cortex.

  2. AIM: Amending Inherent Interpretability via Self-Supervised Masking

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    AIM uses multi-stage feature guidance for self-supervised masking to improve both interpretability (EPG) and accuracy on vision benchmarks.

  3. A Generalized Learning Framework for Self-Supervised Contrastive Learning

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A single framework unifies BYOL, Barlow Twins, and SwAV, plus a plug-in calibration method, ADC, that improves learned representations by preserving input-space distances.

Pith tools