Pith. sign in

REVIEW 4 cited by

Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16602 v1 pith:TVTN2SYN submitted 2023-12-27 cs.CV

classification cs.CV
keywords instructionmodeltasktuningvisualtasksvisioninstructions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional computer vision generally solves each single task independently by a dedicated model with the task instruction implicitly designed in the model architecture, arising two limitations: (1) it leads to task-specific models, which require multiple models for different tasks and restrict the potential synergies from diverse tasks; (2) it leads to a pre-defined and fixed model interface that has limited interactivity and adaptability in following user' task instructions. To address them, Visual Instruction Tuning (VIT) has been intensively studied recently, which finetunes a large vision model with language as task instructions, aiming to learn from a wide range of vision tasks described by language instructions a general-purpose multimodal model that can follow arbitrary instructions and thus solve arbitrary tasks specified by the user. This work aims to provide a systematic review of visual instruction tuning, covering (1) the background that presents computer vision task paradigms and the development of VIT; (2) the foundations of VIT that introduce commonly used network architectures, visual instruction tuning frameworks and objectives, and evaluation setups and tasks; (3) the commonly used datasets in visual instruction tuning and evaluation; (4) the review of existing VIT methods that categorizes them with a taxonomy according to both the studied vision task and the method design and highlights the major contributions, strengths, and shortcomings of them; (5) the comparison and discussion of VIT methods over various instruction-following benchmarks; (6) several challenges, open directions and possible future works in visual instruction tuning research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning in Deep Networks under Dale's Constraint

    cs.AI 2026-08 reject novelty 7.0 of 10

    An on-off two-channel network with fixed-sign synapses and local Hebbian learning is claimed to recover backpropagation exactly under symmetric weights and to beat comparable vanilla networks on Tiny ImageNet.

  2. Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

    cs.CV 2024-12 reject novelty 6.0 of 10

    A new 40K-pair vision-language dataset for social navigation and a fine-tuned VLM that reportedly beats GPT-4V and Gemini in human-judged scene reasoning.

  3. PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

    cs.CV 2025-01 reject novelty 4.0 of 10

    PAINT reduces hallucination in LLaVA-1.5 by selectively amplifying attention to ViT-defined local and summary tokens, but the gains are selected on the test set and lack error bars.

  4. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools