Pith. sign in

REVIEW 3 cited by

Multimodal Learning with Transformers: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.06488 v2 pith:QHMSR3A3 submitted 2022-06-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords multimodaltransformerlearningapplicationsdatasurveyresearchreview
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer is a promising neural network learner, and has achieved great success in various machine learning tasks. Thanks to the recent prevalence of multimodal applications and big data, Transformer-based multimodal learning has become a hot topic in AI research. This paper presents a comprehensive survey of Transformer techniques oriented at multimodal data. The main contents of this survey include: (1) a background of multimodal learning, Transformer ecosystem, and the multimodal big data era, (2) a theoretical review of Vanilla Transformer, Vision Transformer, and multimodal Transformers, from a geometrically topological perspective, (3) a review of multimodal Transformer applications, via two important paradigms, i.e., for multimodal pretraining and for specific multimodal tasks, (4) a summary of the common challenges and designs shared by the multimodal Transformer models and applications, and (5) a discussion of open problems and potential research directions for the community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) reduces memory and boosts throughput for multi-channel vision foundation models by spreading tokenization and channel fusion across GPUs with only a small qu...

  2. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

  3. Farm-Level, In-Season Crop Identification for India

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A Google DeepMind team built a transformer-based system that maps 12 crops across India at farm level, in-season, with state-level area agreement of 94% (winter) and 75% (monsoon) against the 2023-24 census.

Pith tools