RCL preserves evidence-reliance in continual multimodal learning to reduce hidden forgetting beyond standard accuracy metrics.
Continual llava: Continual instruction tuning in large vision-language models
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 9verdicts
UNVERDICTED 9representative citing papers
DRAPE generates query-image conditioned prompts on the fly for multimodal continual instruction tuning and reports SOTA results on MCIT benchmarks.
Fine-tuning VLMs for driving erodes pre-trained world knowledge, but shifting adaptation to prompt space via the Drive Expert Adapter preserves generalization while improving task performance.
CGM derives optimal soft and hard mixing strategies for MLLM parameters via curvature-aware second-order analysis to improve the specialization versus forgetting trade-off.
ProtoAda uses format-aware prototypes for better task routing and geometry-aware consolidation to reduce interference in multimodal continual instruction tuning.
Octopus introduces history-free gradient orthogonalization in a two-stage finetuning framework to achieve state-of-the-art continual learning results for multimodal LLMs on the UCIT benchmark.
ECA introduces continual alignment with MoQ, FeDEx, and DR for exemplar-free incremental learning in open-ended image-to-text generation, evaluated on four new benchmarks showing reduced forgetting.
CRAM uses adaptive MoE with centroid routing and orthogonality constraints to enable parameter-efficient multimodal continual instruction tuning while mitigating forgetting.
Integrates SAM-Audio dense representations with guided attention and dual distillation for audio-visual class-incremental learning, reporting consistent outperformance on benchmarks.
citing papers explorer
-
Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails
RCL preserves evidence-reliance in continual multimodal learning to reduce hidden forgetting beyond standard accuracy metrics.
-
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
DRAPE generates query-image conditioned prompts on the fly for multimodal continual instruction tuning and reports SOTA results on MCIT benchmarks.
-
The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
Fine-tuning VLMs for driving erodes pre-trained world knowledge, but shifting adaptation to prompt space via the Drive Expert Adapter preserves generalization while improving task performance.
-
Curvature-Guided Mixing for MLLM Adaptation
CGM derives optimal soft and hard mixing strategies for MLLM parameters via curvature-aware second-order analysis to improve the specialization versus forgetting trade-off.
-
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
ProtoAda uses format-aware prototypes for better task routing and geometry-aware consolidation to reduce interference in multimodal continual instruction tuning.
-
Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models
Octopus introduces history-free gradient orthogonalization in a two-stage finetuning framework to achieve state-of-the-art continual learning results for multimodal LLMs on the UCIT benchmark.
-
ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation
ECA introduces continual alignment with MoQ, FeDEx, and DR for exemplar-free incremental learning in open-ended image-to-text generation, evaluated on four new benchmarks showing reduced forgetting.
-
CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
CRAM uses adaptive MoE with centroid routing and orthogonality constraints to enable parameter-efficient multimodal continual instruction tuning while mitigating forgetting.
-
Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio
Integrates SAM-Audio dense representations with guided attention and dual distillation for audio-visual class-incremental learning, reporting consistent outperformance on benchmarks.