Pith. sign in

REVIEW 1 cited by

On the Role of Discrete Tokenization in Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.09087 v1 pith:TUFVU2Q3 submitted 2024-07-12 cs.LG cs.CV

classification cs.LGcs.CV
keywords discretelearningtokensclustermimcontrastivemaskedmetricnamed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the realm of self-supervised learning (SSL), masked image modeling (MIM) has gained popularity alongside contrastive learning methods. MIM involves reconstructing masked regions of input images using their unmasked portions. A notable subset of MIM methodologies employs discrete tokens as the reconstruction target, but the theoretical underpinnings of this choice remain underexplored. In this paper, we explore the role of these discrete tokens, aiming to unravel their benefits and limitations. Building upon the connection between MIM and contrastive learning, we provide a comprehensive theoretical understanding on how discrete tokenization affects the model's generalization capabilities. Furthermore, we propose a novel metric named TCAS, which is specifically designed to assess the effectiveness of discrete tokens within the MIM framework. Inspired by this metric, we contribute an innovative tokenizer design and propose a corresponding MIM method named ClusterMIM. It demonstrates superior performance on a variety of benchmark datasets and ViT backbones. Code is available at https://github.com/PKU-ML/ClusterMIM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Identifying and Understanding Cross-Class Features in Adversarial Training

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Models trained against adversarial attacks first learn class-shared features, then forget them as robust overfitting sets in, and preserving these features explains why soft-label training helps.

Pith tools