Pith. sign in

REVIEW 1 cited by

Large Language Models Meet Computer Vision: A Brief Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16673 v1 pith:GN5HQRQY submitted 2023-11-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords llmssurveymodelsvisionlanguagepotentialtransformerscomputer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, the intersection of Large Language Models (LLMs) and Computer Vision (CV) has emerged as a pivotal area of research, driving significant advancements in the field of Artificial Intelligence (AI). As transformers have become the backbone of many state-of-the-art models in both Natural Language Processing (NLP) and CV, understanding their evolution and potential enhancements is crucial. This survey paper delves into the latest progressions in the domain of transformers and their subsequent successors, emphasizing their potential to revolutionize Vision Transformers (ViTs) and LLMs. This survey also presents a comparative analysis, juxtaposing the performance metrics of several leading paid and open-source LLMs, shedding light on their strengths and areas of improvement as well as a literature review on how LLMs are being used to tackle vision related tasks. Furthermore, the survey presents a comprehensive collection of datasets employed to train LLMs, offering insights into the diverse data available to achieve high performance in various pre-training and downstream tasks of LLMs. The survey is concluded by highlighting open directions in the field, suggesting potential venues for future research and development. This survey aims to underscores the profound intersection of LLMs on CV, leading to a new era of integrated and advanced AI models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

    cs.CV 2025-05 reject novelty 2.0 of 10

    The paper reports lower hallucination scores when selecting the best of three filtered image variants, but the selection uses the ground truth, so the improvement is an artifact of choosing the minimum.

Pith tools