Pith. sign in

REVIEW 7 cited by

Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10237 v3 pith:R4N2FOWY submitted 2024-04-16 cs.CV cs.CL

Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

classification cs.CV cs.CL
keywords medicalmodelmultimodaldomain-specifictasksacrossmed-moetuning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advancements in general-purpose or domain-specific multimodal large language models (LLMs) have witnessed remarkable progress for medical decision-making. However, they are designated for specific classification or generative tasks, and require model training or finetuning on large-scale datasets with sizeable parameters and tremendous computing, hindering their clinical utility across diverse resource-constrained scenarios in practice. In this paper, we propose a novel and lightweight framework Med-MoE (Mixture-of-Experts) that tackles both discriminative and generative multimodal medical tasks. The learning of Med-MoE consists of three steps: multimodal medical alignment, instruction tuning and routing, and domain-specific MoE tuning. After aligning multimodal medical images with LLM tokens, we then enable the model for different multimodal medical tasks with instruction tuning, together with a trainable router tailored for expert selection across input modalities. Finally, the model is tuned by integrating the router with multiple domain-specific experts, which are selectively activated and further empowered by meta expert. Comprehensive experiments on both open- and close-end medical question answering (Med-VQA) and image classification tasks across datasets such as VQA-RAD, SLAKE and Path-VQA demonstrate that our model can achieve performance superior to or on par with state-of-the-art baselines, while only requiring approximately 30\%-50\% of activated model parameters. Extensive analysis and ablations corroborate the effectiveness and practical utility of our method.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BLEG: LLM Functions as Powerful fMRI Graph-Enhancer for Brain Network Analysis

    cs.LG 2026-04 unverdicted novelty 7.0

    BLEG enhances GNNs for fMRI brain network analysis by prompting LLMs for text augmentation, using cost-effective instruction tuning, and applying alignment losses during joint training.

  2. Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis

    cs.CV 2026-04 unverdicted novelty 6.0

    Transferring a 2D MLLM to 3D CT inputs via parameter reuse, a Text-Guided Hierarchical MoE framework, and two-stage training yields better performance than prior 3D medical MLLMs on medical report generation and visua...

  3. Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

    cs.CV 2026-03 conditional novelty 6.0

    SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.

  4. HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding

    cs.LG 2025-06 unverdicted novelty 6.0

    HeartcareGPT proposes Dual Stream Projection Alignment (DSPA) on a structure-aware tokenizer for unified ECG signal-image modeling, supported by Heartcare-400K dataset and Heartcare-Bench.

  5. When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis

    cs.MA 2026-07 reject novelty 5.0

    BrainAgent, a training-free agentic LLM framework with graph understanding, knowledge retrieval, case retrieval, and reflection, claims improved but still moderate connectome classification and interpretability.

  6. Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

    cs.LG 2026-05 accept novelty 5.0

    A literature survey that categorizes how Mixture-of-Experts architectures address multimodal learning challenges and identifies open research gaps.

  7. Cross-Stage Attention Multi-Expert Network for Radiologist-Inspired Breast Ultrasound Diagnosis

    cs.CV 2026-05 conditional novelty 4.0

    CSA-MoE-Net achieves 96.33% accuracy, 94.09% precision, 98.53% recall, 96.25% F1-score and 99.50% AUC on 2,129 balanced breast ultrasound images, improving over ResNet-18 by 3.01 to 5.42 percentage points.