Pith. sign in

REVIEW 4 cited by

Vision-and-Language Pretrained Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07356 v5 pith:EQ54JXV6 submitted 2022-04-15 cs.CV cs.CL

classification cs.CVcs.CL
keywords languagevisionmodelspretrainedvlpmsjointpretrainingrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision and language pretraining by feeding visual and linguistic contents into a multi-layer transformer, Visual-Language Pretrained Models (VLPMs). In this paper, we present an overview of the major advances achieved in VLPMs for producing joint representations of vision and language. As the preliminaries, we briefly describe the general task definition and genetic architecture of VLPMs. We first discuss the language and vision data encoding methods and then present the mainstream VLPM structure as the core content. We further summarise several essential pretraining and fine-tuning strategies. Finally, we highlight three future directions for both CV and NLP researchers to provide insightful guidance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 7 citations worldwide. Full citation record

  1. Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Using maps and metadata as extra context for GPT-4o caption generation yields a richer remote sensing dataset, fMoW-mm, with claimed lower hallucination rates and better few-shot detection than prior datasets.

  2. One Fits All: General Mobility Trajectory Modeling via Masked Conditional Diffusion

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A masked conditional diffusion model with historical user embeddings simultaneously performs trajectory generation, recovery, and prediction, beating task-specific baselines on two datasets.

  3. CF-VLM:CounterFactual Vision-Language Fine-tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.

  4. One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning

    cs.CV 2025-01 reject novelty 2.0 of 10

    Block-LoRA partitions and shares the down-projection of LoRA, but this reduces mathematically to a standard LoRA with a lower rank.

Pith tools