Pith. sign in

REVIEW 9 cited by

MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10773 v3 pith:X3GTCZB4 submitted 2022-12-21 cs.CL

classification cs.CL
keywords tasksinstructionsinstructionmultimodallearningtuningzero-shotdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction tuning, a new learning paradigm that fine-tunes pre-trained language models on tasks specified through instructions, has shown promising zero-shot performance on various natural language processing tasks. However, it has yet to be explored for vision and multimodal tasks. In this work, we introduce MUL-TIINSTRUCT, the first multimodal instruction tuning benchmark dataset that consists of 62 diverse multimodal tasks in a unified seq-to-seq format covering 10 broad categories. The tasks are derived from 21 existing open-source datasets and each task is equipped with 5 expert-written instructions. We take OFA as the base pre-trained model for multimodal instruction tuning, and to further improve its zero-shot performance, we explore multiple transfer learning strategies to leverage the large-scale NATURAL INSTRUCTIONS dataset. Experimental results demonstrate strong zero-shot performance on various unseen multimodal tasks and the benefit of transfer learning from a text-only instruction dataset. We also design a new evaluation metric - Sensitivity, to evaluate how sensitive the model is to the variety of instructions. Our results indicate that fine-tuning the model on a diverse set of tasks and instructions leads to a reduced sensitivity to variations in instructions for each task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  2. Pilot: Building the Federated Multimodal Instruction Tuning Framework

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Pilot is a federated multimodal instruction tuning framework that combines task-specific and client-specific adapters with a cross-task mixture-of-adapters module and Euclidean-distance-based text adapter aggregation.

  3. Error-driven Data-efficient Large Multimodal Model Tuning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An error-driven teacher-student pipeline extracts a student LMM's missing skills from validation mistakes and retrieves targeted samples from a task-agnostic dataset to fine-tune it.

  4. RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    RSUniVLM is a 1-billion-parameter remote sensing vision-language model that unifies image-, region-, and pixel-level tasks plus multi-image change analysis, achieving state-of-the-art visual grounding on VRSBench and ...

  5. Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By fine-tuning LLaVA on a knowledge-infused agricultural dataset, Agri-LLaVA improves agricultural conversation and VQA over general LMMs, with gains of about 5 points over LLaVA on the new benchmark.

  6. Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A frozen LLM with a small EMG adaptor converts unvoiced EMG to text at 0.49 average word error rate on a 67-word closed vocabulary without any voiced audio.

  7. Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Instructify converts image metadata into visual instruction-tuning conversations with open LLMs, matching or exceeding GPT-4-generated data quality on LMM benchmarks.

  8. LinVT: Empower Your Image-level Large Language Model to Understand Videos

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A plug-and-play linear video tokenizer converts existing image-based LLMs into video-understanding LLMs by condensing frames into weighted-average tokens while preserving image capabilities.

  9. From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Adding an L2 loss that pushes the language model's image hidden states back toward the input image embeddings improves LLaVA-style models on several VQA benchmarks, with some benchmarks unaffected or slightly worse.

Pith tools