Pith. sign in

REVIEW 11 cited by

A Survey of Vision-Language Pre-Trained Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.10936 v2 pith:LIBXGDPW submitted 2022-02-18 cs.CV cs.CLcs.LG

A Survey of Vision-Language Pre-Trained Models

classification cs.CV cs.CLcs.LG
keywords modelspre-trainedpre-trainingdownstreamintroducelearningmainstreamrecent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As transformer evolves, pre-trained models have advanced at a breakneck pace in recent years. They have dominated the mainstream techniques in natural language processing (NLP) and computer vision (CV). How to adapt pre-training to the field of Vision-and-Language (V-L) learning and improve downstream task performance becomes a focus of multimodal learning. In this paper, we review the recent progress in Vision-Language Pre-Trained Models (VL-PTMs). As the core content, we first briefly introduce several ways to encode raw images and texts to single-modal embeddings before pre-training. Then, we dive into the mainstream architectures of VL-PTMs in modeling the interaction between text and image representations. We further present widely-used pre-training tasks, and then we introduce some common downstream tasks. We finally conclude this paper and present some promising research directions. Our survey aims to provide researchers with synthesis and pointer to related research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Temporal Inversion for Learning Interval Change in Chest X-Rays

    cs.CV 2026-04 unverdicted novelty 7.0

    TILA uses temporal inversion of image pairs as a supervisory signal to make existing temporal vision-language models more sensitive to directional interval changes in chest X-rays.

  2. Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

    cs.CV 2025-08 unverdicted novelty 7.0

    The paper offers a comprehensive survey and proposes a new taxonomy for continual learning strategies in VLMs and MLLMs to combat catastrophic forgetting beyond traditional methods.

  3. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 conditional novelty 6.0

    Large multi-source tactile data plus question-guided Gaussian temporal MoE yields ~7-point gains over VTV-LLM on tactile property and commonsense reasoning tasks, with improved unseen-sensor generalization.

  4. Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization

    cs.CV 2026-04 unverdicted novelty 6.0

    RCSR is a personalization-friendly federated framework that improves cross-modal retrieval accuracy and stability under missing modalities via semantic routing and adapters.

  5. Towards Long-horizon Agentic Multimodal Search

    cs.CV 2026-04 unverdicted novelty 6.0

    LMM-Searcher uses file-based visual UIDs and a fetch tool plus 12K synthesized trajectories to fine-tune a multimodal agent that scales to 100-turn horizons and reaches SOTA among open-source models on MM-BrowseComp a...

  6. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 6.0

    LLM Orchestration integrates modality experts via an LLM controller, cross-modal memory, and interaction layer to enable multimodal input-output without gradient-based training.

  7. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 unverdicted novelty 5.0

    TouchThinker introduces a 1M-scale multi-source tactile dataset and action-aware modeling to scale commonsense reasoning from tactile observations, reporting competitive performance on new and existing benchmarks.

  8. Horizontal and Longitudinal Comparisons Among AI Subfields: A Bibliometric Perspective

    cs.DL 2026-05 unverdicted novelty 5.0

    Bibliometric analysis of AI subfields from 2000-2024 shows high-intensity knowledge diffusion with structural differentiation: CV leads in impact, ML has concentrated international ties, Web&IR is industry-driven, whi...

  9. Efficient Prompt Learning for Traffic Forecasting

    cs.LG 2026-05 unverdicted novelty 5.0

    SimpleST is a model-agnostic prompt tuning framework that lets pre-trained spatio-temporal GNNs adapt to distribution shifts in traffic data while keeping all original model weights fixed.

  10. Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid

    cs.AI 2026-05 unverdicted novelty 5.0

    A formalized Minimal Cognitive Grid ranks computational models of analogy and metaphor by alignment with cognitive theories using Functional/Structural Ratio, Generality, and Performance Match dimensions.

  11. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 5.0

    A training-free orchestration framework integrates off-the-shelf modality experts via an LLM controller, text-centric cross-modal memory, and unified interaction layer to enable multimodal input-output without joint training.