Pith. sign in

REVIEW 3 cited by

Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.02684 v3 pith:SEKRBCZZ submitted 2023-11-05 cs.CV cs.CL

classification cs.CVcs.CL
keywords multimodallearningmllmscalleddownstreaminterferencelanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies have demonstrated Large Language Models (LLMs) can extend their zero-shot generalization capabilities to multimodal learning through instruction tuning. As more modalities and downstream tasks are introduced, negative conflicts and interference may have a worse impact on performance. While this phenomenon has been overlooked in previous work, we propose a novel and extensible framework, called Octavius, for comprehensive studies and experimentation on multimodal learning with Multimodal Large Language Models (MLLMs). Specifically, we combine the well-known Mixture-of-Experts (MoE) and one of the representative PEFT techniques, i.e., LoRA, designing a novel LLM-based decoder, called LoRA-MoE, for multimodal learning. To the best of our knowledge, we are one of the pioneering efforts to introduce MoE into MLLMs to address this problem. The experimental results (about 20% improvement) have shown the effectiveness and versatility of our design in various 2D and 3D downstream tasks. Code and datasets are available at https://openlamm.github.io/tutorial/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A two-stage TPC→ADS training schedule yields the best balance across Arabic ASR, speech summarization, dialect ID, and emotion recognition in a low-resource audio LLM.

  2. VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A CoT-annotated dataset (VaccineRAG) plus segment-level GRPO (Partial-GRPO) improves multimodal large language models' ability to ignore harmful retrieved samples in retrieval-augmented generation tasks.

  3. ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

    cs.LG 2026-04 unverdicted novelty 4.0 of 10

    Standard Conditional Flow Matching loss is a misleading early plateau; physics-informed metrics keep improving, so ScatterPrism and multi-metric diagnostics are needed for kinematic fidelity.

Pith tools