Pith. sign in

REVIEW 7 cited by

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.19835 v1 pith:67BC3JCR submitted 2025-06-24 cs.CL

classification cs.CL
keywords medicalllmsdiagnosticframeworkknowledgemodulardiagnosismodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in medical Large Language Models (LLMs) have showcased their powerful reasoning and diagnostic capabilities. Despite their success, current unified multimodal medical LLMs face limitations in knowledge update costs, comprehensiveness, and flexibility. To address these challenges, we introduce the Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis (MAM). Inspired by our empirical findings highlighting the benefits of role assignment and diagnostic discernment in LLMs, MAM decomposes the medical diagnostic process into specialized roles: a General Practitioner, Specialist Team, Radiologist, Medical Assistant, and Director, each embodied by an LLM-based agent. This modular and collaborative framework enables efficient knowledge updates and leverages existing medical LLMs and knowledge bases. Extensive experimental evaluations conducted on a wide range of publicly accessible multimodal medical datasets, incorporating text, image, audio, and video modalities, demonstrate that MAM consistently surpasses the performance of modality-specific LLMs. Notably, MAM achieves significant performance improvements ranging from 18% to 365% compared to baseline models. Our code is released at https://github.com/yczhou001/MAM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  2. Context-Adaptive Synthesis and Compression for Enhanced Retrieval-Augmented Generation in Complex Domains

    cs.CL 2025-08 conditional novelty 4.0 of 10

    CASC uses a fine-tuned Llama-2-7B to extract, de-conflict, and structure retrieved contexts, reporting higher F1 and lower hallucination than RAG baselines on the new SciDocs-QA benchmark.

  3. LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

    cs.LG 2025-08 reject novelty 4.0 of 10

    An LVLM-driven iterative text-to-image framework whose claimed performance scores are explicitly labeled fictitious, so no empirical result is established.

  4. CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

    cs.LG 2025-07 reject novelty 4.0 of 10

    CIMR, an iterative reasoning wrapper around LLaVA-1.5-7B, reports 91.5% task completion on a newly constructed but unreleased synthetic MAP dataset, above GPT-4V at 89.2%.

  5. LVLM-Composer's Explicit Planning for Image Generation

    cs.CV 2025-07 reject novelty 4.0 of 10

    An image generation model that explicitly plans objects, attributes, locations, and relations before synthesizing the image, with reported gains on LongBench-T2I that cannot be verified from the paper.

  6. Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

    cs.CV 2025-07 reject novelty 3.0 of 10

    Hi-SSLVLM combines hierarchical self-captioning, internal sub-prompt planning, and a CLIP-based consistency loss, and reports judged compositional fidelity gains of roughly 0.04 to 0.09 points that no significance tes...

  7. Large Language Models for Zero-Shot Multicultural Name Recognition

    cs.CL 2025-07 reject novelty 3.0 of 10

    A prompt-tuned LLM with data augmentation and cultural context prompts reportedly recognizes multicultural names at 93.1% accuracy and unseen names at 89.5%, but the evidence is not reproducible.

Pith tools