Pith. sign in

REVIEW 5 cited by

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.14548 v2 pith:32HILPGW submitted 2021-04-29 cs.CV

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

classification cs.CV
keywords learningcontrastiveimagenetmethodmethodsnearest-neighboronlypositives
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other instances in the dataset. Our method, Nearest-Neighbor Contrastive Learning of visual Representations (NNCLR), samples the nearest neighbors from the dataset in the latent space, and treats them as positives. This provides more semantic variations than pre-defined transformations. We find that using the nearest-neighbor as positive in contrastive losses improves performance significantly on ImageNet classification, from 71.7% to 75.6%, outperforming previous state-of-the-art methods. On semi-supervised learning benchmarks we improve performance significantly when only 1% ImageNet labels are available, from 53.8% to 56.5%. On transfer learning benchmarks our method outperforms state-of-the-art methods (including supervised learning with ImageNet) on 8 out of 12 downstream datasets. Furthermore, we demonstrate empirically that our method is less reliant on complex data augmentations. We see a relative reduction of only 2.1% ImageNet Top-1 accuracy when we train using only random crops.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    cs.CL 2023-09 unverdicted novelty 8.0

    Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.

  2. Star-forming clump detection in nearby galaxies using Faster R-CNN and $ugrizy$ imaging data from CLAUDS and HSC-SSP

    astro-ph.IM 2026-07 conditional novelty 6.5

    A multi-band Faster R-CNN with Zoobot backbone detects star-forming clumps in low-z galaxies at ≥0.9 completeness and ≥0.8 purity on simulated injections, yielding ~1.5M candidates.

  3. Star-forming clump detection in nearby galaxies using Faster R-CNN and $ugrizy$ imaging data from CLAUDS and HSC-SSP

    astro-ph.IM 2026-07 conditional novelty 6.0

    A six-band Faster R-CNN with the Zoobot backbone detects star-forming clump candidates in ~700,000 local galaxies, claiming ~90% completeness and ~80% purity for clumps brighter than the surveys' detection limits.

  4. Contrastive learning of extragalactic stellar streams. Sculpting a latent space of representations with DES DR2 photometry

    astro-ph.GA 2026-01 conditional novelty 6.0

    Applying NNCLR contrastive learning to DES DR2 galaxy cutouts yields embeddings that cluster major-merger galaxies but not stellar streams; a tiered sigmoid scaling redirects the network's saliency toward low-surface-...

  5. Vision Transformers Need Registers

    cs.CV 2023-09 unverdicted novelty 6.0

    Adding register tokens to Vision Transformers eliminates high-norm background artifacts and raises state-of-the-art performance on dense visual prediction tasks.