REVIEW 9 cited by
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Large Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets. However, novel tasks would be encountered sequentially in dynamic world, which urges for equipping LMMs with multimodal continual instruction learning (MCIT) ability especially for diverse and challenging generative tasks. Existing MCIT methods do not fully exploit the unique attribute of LMMs and often gain performance at the expense of efficiency. In this paper, we propose a novel prompt learning framework for MCIT to effectively alleviate forgetting of previous knowledge while managing computational complexity with natural image-text supervision. Concretely, we learn prompts for each task and exploit efficient prompt fusion for knowledge transfer and prompt selection for complexity management with dual-modality guidance. Extensive experiments demonstrate that our approach achieves substantial +14.26% performance gain on MCIT benchmarks with remarkable $\times$ 1.42 inference speed free from growing computation. Code is available at https://github.com/AuroraZengfh/ModalPrompt.
Forward citations
Cited by 9 Pith papers
-
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Many-shot in-context learning in vision-language models can collapse accuracy as demonstrations accumulate, and the failure is causally localized to the vision-language integration pathway, where a small adapter repairs it.
-
LLaVA-c: Continual Improved Visual Instruction Tuning
With spectral-aware consolidation and unsupervised inquiry regularization, LLaVA-1.5 can be trained task-by-task with performance matching or surpassing joint multitask training.
-
EventVAD: Training-Free Event-Aware Video Anomaly Detection
EventVAD improves training-free video anomaly detection by detecting event boundaries from CLIP and RAFT features and feeding coherent event segments to a 7B multimodal LLM, achieving 82.03 AUC on UCF-Crime and 64.04 ...
-
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.
-
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
Reinforcement fine-tuning largely prevents catastrophic forgetting during continual post-training of a multimodal LLM, while supervised fine-tuning degrades both task and general performance.
-
Creating User-steerable Projections with Interactive Semantic Mapping
Fusing data embeddings with text-label embeddings generated from user prompts changes dimensionality-reduction scatterplots so that clusters follow the user's semantic questions.
-
SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
SEFE reduces both answer-format drift and knowledge forgetting in multimodal continual instruction tuning by mixing question styles across tasks and penalizing changes at high-importance LoRA weight positions.
-
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.
-
Continually Evolved Multimodal Foundation Models for Cancer Prognosis
A cancer prognosis model that grows with new data modalities via LoRA adapters and gated query fusion is claimed to beat fusion baselines, but the supporting table is incomplete and internally inconsistent.
Discussion (0). Continue with ORCID to comment.