Pith. sign in

REVIEW 8 cited by

Masked Vision and Language Modeling for Multi-modal Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02131 v2 pith:GDYZIHN5 submitted 2022-08-03 cs.CV cs.CLcs.LG

Masked Vision and Language Modeling for Multi-modal Representation Learning

classification cs.CV cs.CLcs.LG
keywords maskedlanguagemodelingmodalitydataimagesignalvision
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this paper, we study how to use masked signal modeling in vision and language (V+L) representation learning. Instead of developing masked language modeling (MLM) and masked image modeling (MIM) independently, we propose to build joint masked vision and language modeling, where the masked signal of one modality is reconstructed with the help from another modality. This is motivated by the nature of image-text paired data that both of the image and the text convey almost the same information but in different formats. The masked signal reconstruction of one modality conditioned on another modality can also implicitly learn cross-modal alignment between language tokens and image patches. Our experiments on various V+L tasks show that the proposed method, along with common V+L alignment losses, achieves state-of-the-art performance in the regime of millions of pre-training data. Also, we outperforms the other competitors by a significant margin in limited data scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

    cs.AI 2026-04 unverdicted novelty 8.0

    FeynmanBench is the first benchmark for evaluating multimodal LLMs on diagrammatic reasoning with Feynman diagrams, revealing systematic failures in enforcing physical constraints and global topology.

  2. FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

    cs.AI 2026-04 conditional novelty 7.0

    Across 19 multimodal LLMs, local Feynman-diagram recognition stays high while topological reconstruction and full amplitude derivation collapse.

  3. Neural Scene Designer: Self-Styled Semantic Image Manipulation

    cs.CV 2025-09 conditional novelty 6.0

    NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.

  4. Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

    cs.CV 2026-06 unverdicted novelty 5.0

    MACCO applies cross-modal masked reconstruction of compositional concepts with inter- and intra-modal auxiliary objectives to improve visio-linguistic compositionality in VLMs.

  5. Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

    cs.CV 2026-05 unverdicted novelty 5.0

    Strong generalist vision foundation models match or outperform electro-optical specific models in remote sensing retrieval with better cross-scene stability.

  6. Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation

    cs.CV 2026-04 unverdicted novelty 5.0

    PCMECL improves speech-preserving facial expression manipulation by learning personalized prompts from individual visuals and using feature differencing to align visual and semantic changes from VLMs.

  7. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

    cs.CV 2024-08 unverdicted novelty 5.0

    mPLUG-Owl3 introduces hyper attention blocks to integrate vision and language for long image-sequence understanding and reports SOTA results on single-image, multi-image, and video benchmarks.

  8. mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

    cs.CL 2023-11 unverdicted novelty 5.0

    mPLUG-Owl2 presents a modular MLLM architecture that enables modality collaboration via shared functional modules and modality-adaptive components, achieving SOTA on both text and multi-modal tasks with one generic model.