Pith. sign in

REVIEW 9 cited by

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05849 v2 pith:2N3BK6JS submitted 2024-10-08 cs.CV

classification cs.CV
keywords mcitpromptinstructionlearninglmmsmultimodalabilitycomplexity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Large Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets. However, novel tasks would be encountered sequentially in dynamic world, which urges for equipping LMMs with multimodal continual instruction learning (MCIT) ability especially for diverse and challenging generative tasks. Existing MCIT methods do not fully exploit the unique attribute of LMMs and often gain performance at the expense of efficiency. In this paper, we propose a novel prompt learning framework for MCIT to effectively alleviate forgetting of previous knowledge while managing computational complexity with natural image-text supervision. Concretely, we learn prompts for each task and exploit efficient prompt fusion for knowledge transfer and prompt selection for complexity management with dual-modality guidance. Extensive experiments demonstrate that our approach achieves substantial +14.26% performance gain on MCIT benchmarks with remarkable $\times$ 1.42 inference speed free from growing computation. Code is available at https://github.com/AuroraZengfh/ModalPrompt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Context Collapse in Vision-Language Models and How to Mitigate it?

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Many-shot in-context learning in vision-language models can collapse accuracy as demonstrations accumulate, and the failure is causally localized to the vision-language integration pathway, where a small adapter repairs it.

  2. LLaVA-c: Continual Improved Visual Instruction Tuning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    With spectral-aware consolidation and unsupervised inquiry regularization, LLaVA-1.5 can be trained task-by-task with performance matching or surpassing joint multitask training.

  3. EventVAD: Training-Free Event-Aware Video Anomaly Detection

    cs.CV 2025-04 conditional novelty 6.0 of 10

    EventVAD improves training-free video anomaly detection by detecting event boundaries from CLIP and RAFT features and feeding coherent event segments to a 7B multimodal LLM, achieving 82.03 AUC on UCF-Crime and 64.04 ...

  4. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

    cs.AI 2025-01 conditional novelty 6.0 of 10

    ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.

  5. Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Reinforcement fine-tuning largely prevents catastrophic forgetting during continual post-training of a multimodal LLM, while supervised fine-tuning degrades both task and general performance.

  6. Creating User-steerable Projections with Interactive Semantic Mapping

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fusing data embeddings with text-label embeddings generated from user prompts changes dimensionality-reduction scatterplots so that clusters follow the user's semantic questions.

  7. SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    SEFE reduces both answer-format drift and knowledge forgetting in multimodal continual instruction tuning by mixing question styles across tasks and penalizing changes at high-importance LoRA weight positions.

  8. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

  9. Continually Evolved Multimodal Foundation Models for Cancer Prognosis

    cs.LG 2025-01 reject novelty 4.0 of 10

    A cancer prognosis model that grows with new data modalities via LoRA adapters and gated query fusion is claimed to beat fusion baselines, but the supporting table is incomplete and internally inconsistent.

Pith tools