Pith. sign in

REVIEW 12 cited by

MMedAgent: Learning to Use Medical Tools with Multi-modal Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02483 v2 pith:FHYIQY6N submitted 2024-07-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords medicaltoolsagentmmedagentmodelstextbfacrossbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by selecting appropriate specialized models as tools based on user inputs. However, such advancements have not been extensively explored within the medical domain. To bridge this gap, this paper introduces the first agent explicitly designed for the medical field, named \textbf{M}ulti-modal \textbf{Med}ical \textbf{Agent} (MMedAgent). We curate an instruction-tuning dataset comprising six medical tools solving seven tasks across five modalities, enabling the agent to choose the most suitable tools for a given task. Comprehensive experiments demonstrate that MMedAgent achieves superior performance across a variety of medical tasks compared to state-of-the-art open-source methods and even the closed-source model, GPT-4o. Furthermore, MMedAgent exhibits efficiency in updating and integrating new medical tools. Codes and models are all available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

    cs.CV 2026-01 conditional novelty 7.0 of 10

    IBISAgent enables MLLMs to perform iterative pixel-level visual reasoning for biomedical object referring and segmentation via text-based clicks and agentic RL, outperforming prior SOTA methods without model modifications.

  2. Evaluating Agentic Harness Systems for Autonomous Computational Pathology

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Across 369 agent trajectories on 41 pathology tasks, formal end-to-end workflow completion was rare (10/369), with tool execution, result binding, and reflection far weaker than planning and reporting.

  3. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  4. Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

    cs.AI 2025-10 unverdicted novelty 6.0 of 10

    TRACE is a reference-free multi-dimensional evaluation framework for tool-augmented LLM reasoning trajectories that uses an evidence bank and is validated on a new meta-evaluation dataset of flawed trajectories.

  5. Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

    cs.AI 2025-10 conditional novelty 6.0 of 10

    TRACE uses an evidence bank to score tool-augmented LLM agents on efficiency, hallucination, and adaptivity without ground-truth trajectories.

  6. Improving Alignment in LVLMs with Debiased Self-Judgment

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A contrastive self-judgment score that subtracts a model's image-free confidence from its visual confidence is used to guide decoding, safety moderation, and DPO training, improving hallucination and safety metrics ac...

  7. WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.

  8. AT-CXR: Uncertainty-Aware Agentic Triage for Chest X-rays

    eess.IV 2025-08 conditional novelty 5.0 of 10

    An uncertainty-aware agentic router with abstention improves pulmonary edema triage on a small balanced chest x-ray subset, but the selective-prediction gains rest on under-specified evaluation.

  9. A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    A comprehensive review of self-evolving AI agents that improve themselves over time, organized via a framework of inputs, agent system, environment, and optimizers, with domain-specific and safety discussions.

  10. AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    AURA is an agentic system that orchestrates chest X-ray tools to produce self-evaluated visual and textual explanations via counterfactual image generation.

  11. Agentic Reasoning for Large Language Models

    cs.AI 2026-01 unverdicted novelty 4.0 of 10

    The survey structures agentic reasoning for LLMs into foundational, self-evolving, and collective multi-agent layers while distinguishing in-context orchestration from post-training optimization and reviewing applicat...

  12. A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

    cs.LG 2025-07 reject novelty 4.0 of 10

    A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.

Pith tools