Pith. sign in

REVIEW 5 cited by

TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16862 v3 pith:P3CC5WS6 submitted 2023-12-28 cs.CV cs.CL

classification cs.CVcs.CL
keywords languagetinygpt-vmodelstrainingbackbonescomputationalefficientinference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, multimodal large language models (MLLMs) such as GPT-4V have demonstrated remarkable advancements, excelling in a variety of vision-language tasks. Despite their prowess, the closed-source nature and computational demands of such models limit their accessibility and applicability. This study introduces TinyGPT-V, a novel open-source MLLM, designed for efficient training and inference across various vision-language tasks, including image captioning (IC) and visual question answering (VQA). Leveraging a compact yet powerful architecture, TinyGPT-V integrates the Phi-2 language model with pre-trained vision encoders, utilizing a unique mapping module for visual and linguistic information fusion. With a training regimen optimized for small backbones and employing a diverse dataset amalgam, TinyGPT-V requires significantly lower computational resources 24GB for training and as little as 8GB for inference without compromising on performance. Our experiments demonstrate that TinyGPT-V, with its language model 2.8 billion parameters, achieves comparable results in VQA and image inference tasks to its larger counterparts while being uniquely suited for deployment on resource-constrained devices through innovative quantization techniques. This work not only paves the way for more accessible and efficient MLLMs but also underscores the potential of smaller, optimized models in bridging the gap between high performance and computational efficiency in real-world applications. Additionally, this paper introduces a new approach to multimodal large language models using smaller backbones. Our code and training weights are available in the supplementary material.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Jointly optimizing visual-token retention and LLM depth/width under a shared FLOPs budget yields better accuracy-efficiency Pareto frontiers than optimizing either dimension alone.

  2. Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

    cs.MM 2025-01 conditional novelty 6.0 of 10

    A tuning-free pipeline using frozen LLaMA-3, MiniGPT-v2, and Video-ChatGPT reports state-of-the-art zero-shot video moment retrieval on three benchmarks.

  3. NanoVLMs: How small can we go and still make coherent Vision Language Models?

    cs.CV 2025-02 reject novelty 5.0 of 10

    NanoVLMs, 5M to 25M parameter vision-language models trained on simplified GPT-4o captions, are judged by GPT-4o as nearly as coherent as the 50x larger Kosmos-2 on a 25-sample test.

  4. Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A 1.8B-parameter vision-language model, Eve, uses elastic visual experts and type-aware token routing to reach a 68.87% average on six VLM benchmarks while preserving language performance.

  5. EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    EvoMoE creates MoE experts as decaying averages of a single trained FFN and routes tokens with hypernetwork-generated weights, yielding small benchmark gains over MoE-LLaVA.

Pith tools