Pith. sign in

REVIEW 1 cited by

On the Pros and Cons of Momentum Encoder in Self-Supervised Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.05744 v1 pith:T3UTY43K submitted 2022-08-11 cs.CV

classification cs.CV
keywords encodermomentumperformancebenefitfinalframeworkslayerslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exponential Moving Average (EMA or momentum) is widely used in modern self-supervised learning (SSL) approaches, such as MoCo, for enhancing performance. We demonstrate that such momentum can also be plugged into momentum-free SSL frameworks, such as SimCLR, for a performance boost. Despite its wide use as a fundamental component in modern SSL frameworks, the benefit caused by momentum is not well understood. We find that its success can be at least partly attributed to the stability effect. In the first attempt, we analyze how EMA affects each part of the encoder and reveal that the portion near the encoder's input plays an insignificant role while the latter parts have much more influence. By monitoring the gradient of the overall loss with respect to the output of each block in the encoder, we observe that the final layers tend to fluctuate much more than other layers during backpropagation, i.e. less stability. Interestingly, we show that using EMA to the final part of the SSL encoder, i.e. projector, instead of the whole deep network encoder can give comparable or preferable performance. Our proposed projector-only momentum helps maintain the benefit of EMA but avoids the double forward computation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A masked diffusion transformer with a compact condition collector beats the heavier AnyDoor baseline on VITON-HD quality metrics while using a quarter of the parameters and less compute.

Pith tools