Pith. sign in

REVIEW 4 cited by

CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08350 v2 pith:KMADEEAM submitted 2024-03-13 cs.CV

classification cs.CV
keywords instructioncoinknowledgemllmstuningalignmentforgettingassess
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Instruction tuning represents a prevalent strategy employed by Multimodal Large Language Models (MLLMs) to align with human instructions and adapt to new tasks. Nevertheless, MLLMs encounter the challenge of adapting to users' evolving knowledge and demands. Therefore, how to retain existing skills while acquiring new knowledge needs to be investigated. In this paper, we present a comprehensive benchmark, namely Continual Instruction tuNing (CoIN), to assess existing MLLMs in the sequential instruction tuning paradigm. CoIN comprises 10 commonly used datasets spanning 8 task categories, ensuring a diverse range of instructions and tasks. Besides, the trained model is evaluated from two aspects: Instruction Following and General Knowledge, which assess the alignment with human intention and knowledge preserved for reasoning, respectively. Experiments on CoIN demonstrate that current powerful MLLMs still suffer catastrophic forgetting, and the failure in intention alignment assumes the main responsibility, instead of the knowledge forgetting. To this end, we introduce MoELoRA to MLLMs which is effective to retain the previous instruction alignment. Experimental results consistently illustrate the forgetting decreased from this method on CoIN.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Context Collapse in Vision-Language Models and How to Mitigate it?

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Many-shot in-context learning in vision-language models can collapse accuracy as demonstrations accumulate, and the failure is causally localized to the vision-language integration pathway, where a small adapter repairs it.

  2. SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SMoLoRA uses two separately routed LoRA expert groups, one for visual understanding and one for instruction following, to reduce dual catastrophic forgetting in continual visual instruction tuning.

  3. MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Training a multimodal LLM on mostly text-only instructions with a small vision-language tail matches or beats vision-heavy instruction tuning on held-out text and vision tasks at about half the token cost.

  4. How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.

Pith tools