Pith. sign in

REVIEW 19 cited by

BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.06366 v2 pith:LFZ2O5J2 submitted 2022-08-12 cs.CV

classification cs.CV
keywords imagebeitmaskedsemanticaccuracypatchesrepresentationtop-1
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked image modeling (MIM) has demonstrated impressive results in self-supervised representation learning by recovering corrupted image patches. However, most existing studies operate on low-level image pixels, which hinders the exploitation of high-level semantics for representation models. In this work, we propose to use a semantic-rich visual tokenizer as the reconstruction target for masked prediction, providing a systematic way to promote MIM from pixel-level to semantic-level. Specifically, we propose vector-quantized knowledge distillation to train the tokenizer, which discretizes a continuous semantic space to compact codes. We then pretrain vision Transformers by predicting the original visual tokens for the masked image patches. Furthermore, we introduce a patch aggregation strategy which associates discrete image patches to enhance global semantic representation. Experiments on image classification and semantic segmentation show that BEiT v2 outperforms all compared MIM methods. On ImageNet-1K (224 size), the base-size BEiT v2 achieves 85.5% top-1 accuracy for fine-tuning and 80.1% top-1 accuracy for linear probing. The large-size BEiT v2 obtains 87.3% top-1 accuracy for ImageNet-1K (224 size) fine-tuning, and 56.7% mIoU on ADE20K for semantic segmentation. The code and pretrained models are available at https://aka.ms/beitv2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

    cs.CV 2025-06 accept novelty 8.0 of 10

    LAION-C supplies six novel corruptions that stay OOD for web-scale training sets and demonstrates that leading models now rival or exceed human robustness on them.

  2. Multiplayer Interactive World Models with Representation Autoencoders

    cs.CV 2026-07 accept novelty 7.0 of 10

    A 5B-parameter latent diffusion model generates real-time four-player Rocket League matches conditioned on all players' actions, staying stable far beyond its training horizon.

  3. Structure over Pixels: Learning Variable-Length Visual Programs

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    STROP learns variable-length discrete visual programs for images by training a length head against frozen DINOv3 features in a four-phase curriculum while bypassing pixel reconstruction.

  4. Attention Transfer Is Not Universally Effective for Vision Transformers

    cs.CV 2026-05 accept novelty 7.0 of 10

    Attention transfer from ViT teachers succeeds for only 7 of 11 families and fails for the rest because of architectural mismatch between teacher and student.

  5. Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

    cs.CV 2024-10 unverdicted novelty 7.0 of 10

    Janus decouples visual encoding into task-specific pathways inside a single autoregressive transformer to unify multimodal understanding and generation while outperforming earlier unified models.

  6. Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Treating time as an entity-level modality with median-K timestamp selection, attention pooling, and three-stage temporal injection yields large link-prediction gains on the hardest multi-modal ambiguity cases.

  7. ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ExPLoRe turns MoE dispatch weights into per-patch loss coefficients for multi-objective masked image modeling, reporting gains on ImageNet-1K and ADE20K transfer.

  8. Morphology-Aware Multimodal Representation Learning for Insect Phylogenetic Reconstruction

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A multimodal alignment framework derives morphology-aware image embeddings as continuous traits for Bayesian phylogenetic reconstruction, improving topological agreement on the Rove-Tree-11 dataset.

  9. Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Introduces a structural score on token composition in discrete visual token space that correlates with higher validation performance in distilled datasets and guides diffusion-based distillation.

  10. CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    CLIP-RD adds VRD for cross-modality distillation consistency and XRD for bidirectional cross-modal symmetry to align student embedding geometry more closely with the teacher, yielding a 0.8 percentage point gain over ...

  11. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A self-supervised diffusion framework with a low-bitrate vector-quantization bottleneck learns disentangled motion and content latents supporting motion transfer and auto-regressive generation.

  12. BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A self-supervision method makes multimodal LLMs align their input image embeddings with the model's own refined internal representations, improving visual QA scores over LLaVA baselines.

  13. Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection

    cs.CV 2024-11 unverdicted novelty 6.0 of 10

    Orthogonal subspace decomposition via SVD on vision foundation model features preserves high-rank pre-trained knowledge by freezing principal components and adapting residuals, reducing overfitting for better generali...

  14. Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    The work establishes an evaluation framework for personality induction and switching in MLLMs, reporting improved captioning but impaired VQA performance plus balancing and residual effects during multi-trait and dyna...

  15. BAT: Better Audio Transformer Guided by Convex Gated Probing

    cs.SD 2026-02 conditional novelty 5.0 of 10

    CGP probing—layer-gating plus prototypes—closes much of the gap between frozen and fine-tuned audio SSL evaluation, and guides a re-engineered audio transformer (BAT) that improves on the authors' reproduced baselines.

  16. Robustifying Diffusion-Denoised Smoothing Against Covariate Shift

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Adversarially perturbing the noise term of a diffusion denoiser during training improves the certified l2 robustness of denoised randomized smoothing on MNIST, CIFAR-10, and ImageNet, with the largest gains at large p...

  17. TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers

    cs.CV 2025-09 conditional novelty 4.0 of 10

    TinyDrop uses a lightweight model's confidence and attention map to early-exit easy samples and drop uninformative tokens in frozen ViTs, cutting FLOPs by up to 87%.

  18. PaliGemma: A versatile 3B VLM for transfer

    cs.CV 2024-07 unverdicted novelty 4.0 of 10

    PaliGemma is an open 3B VLM based on SigLIP and Gemma that achieves strong performance on nearly 40 diverse open-world tasks including benchmarks, remote-sensing, and segmentation.

  19. Multilingual Vision-Language Models, A Survey

    cs.CL 2025-09 accept novelty 3.0 of 10

    The survey identifies a key tension in multilingual vision-language models between language neutrality via contrastive learning and cultural awareness via diverse data, with most benchmarks relying on translation-base...

Pith tools