REVIEW 4 cited by
Vision-and-Language Pretrained Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision and language pretraining by feeding visual and linguistic contents into a multi-layer transformer, Visual-Language Pretrained Models (VLPMs). In this paper, we present an overview of the major advances achieved in VLPMs for producing joint representations of vision and language. As the preliminaries, we briefly describe the general task definition and genetic architecture of VLPMs. We first discuss the language and vision data encoding methods and then present the mainstream VLPM structure as the core content. We further summarise several essential pretraining and fine-tuning strategies. Finally, we highlight three future directions for both CV and NLP researchers to provide insightful guidance.
Forward citations
Cited by 4 Pith papers
-
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing
Using maps and metadata as extra context for GPT-4o caption generation yields a richer remote sensing dataset, fMoW-mm, with claimed lower hallucination rates and better few-shot detection than prior datasets.
-
One Fits All: General Mobility Trajectory Modeling via Masked Conditional Diffusion
A masked conditional diffusion model with historical user embeddings simultaneously performs trajectory generation, recovery, and prediction, beating task-specific baselines on two datasets.
-
CF-VLM:CounterFactual Vision-Language Fine-tuning
CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.
-
One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning
Block-LoRA partitions and shares the down-projection of LoRA, but this reduces mathematically to a standard LoRA with a lower rank.
Discussion (0). Continue with ORCID to comment.