Pith. sign in

REVIEW 47 cited by

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00578 v1 pith:AEHJQOE6 submitted 2024-03-31 cs.CV

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

classification cs.CV
keywords medicalanalysisimagemulti-modallanguagelargemodelsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images under-explored, despite their richer spatial information. This paper aims to advance 3D medical image analysis with MLLMs. To this end, we present a large-scale 3D multi-modal medical dataset, M3D-Data, comprising 120K image-text pairs and 662K instruction-response pairs specifically tailored for various 3D medical tasks, such as image-text retrieval, report generation, visual question answering, positioning, and segmentation. Additionally, we propose M3D-LaMed, a versatile multi-modal large language model for 3D medical image analysis. Furthermore, we introduce a new 3D multi-modal medical benchmark, M3D-Bench, which facilitates automatic evaluation across eight tasks. Through comprehensive evaluation, our method proves to be a robust model for 3D medical image analysis, outperforming existing solutions. All code, data, and models are publicly available at: https://github.com/BAAI-DCAI/M3D.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 47 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding

    cs.CV 2026-05 accept novelty 8.0

    NeuroQA is a large-scale 3D brain MRI visual question answering benchmark with verified image-grounded QA pairs, multi-domain coverage, and baseline evaluations showing current models lag behind text-only performance.

  2. DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

    cs.CV 2026-05 accept novelty 8.0

    DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.

  3. MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI

    cs.CV 2026-06 unverdicted novelty 7.0

    MRI2Rep generates LI-RADS structured reports from 3D liver MRI via autoregressive modeling on 3929 real-world pairs, reporting 76% case-level sensitivity and 70-75% clinical acceptability in reader study.

  4. CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

    cs.CV 2026-05 unverdicted novelty 7.0

    CardioLens is a leakage-resistant CMR testbed of 473k slices and 13k QA pairs showing current MLLMs exhibit a large clinical reality gap with category-collapse failures on real workflows.

  5. SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    SliceWorld introduces a world-state model for CT report generation that uses predictive and factor-aware objectives on axial slice sequences.

  6. Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 7.0

    CT-SpatialVQA benchmark shows 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks in CT volumes, often below random chance.

  7. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  8. Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis

    cs.CV 2026-04 unverdicted novelty 7.0

    Agentic LLMs autonomously execute complex neuro-radiological workflows like glioma segmentation and multi-timepoint response assessment by directing off-the-shelf tools, without any model training.

  9. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

    cs.CV 2026-01 conditional novelty 7.0

    IBISAgent enables MLLMs to perform iterative pixel-level visual reasoning for biomedical object referring and segmentation via text-based clicks and agentic RL, outperforming prior SOTA methods without model modifications.

  10. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    cs.CV 2026-07 conditional novelty 6.5

    A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.

  11. Astra: a generalizable report generation foundation model for 3D computed tomography

    cs.CV 2026-05 conditional novelty 6.5

    Astra generates style-consistent multi-organ CT reports that generalize across institutions via report harmonization plus GRPO reinforcement learning, improving clinical drafting and enabling synthetic pretraining.

  12. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

    cs.CV 2026-07 conditional novelty 6.0

    ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.

  13. Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining

    cs.CV 2026-07 conditional novelty 6.0

    Cross-patient report-based pair mining plus burden-direction alignment improves CT vision-language pretraining, reaching 85.6 AUROC on CT-RATE zero-shot diagnosis.

  14. MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Training-free multi-cue token compression for 3D medical VLMs that retains and merges tokens using attention, text similarity, and VFM saliency, maintaining diagnostic performance at 50–80% token retention.

  15. Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Disease-probe AUROC on frozen 3D-CT tokens predicts report-generation clinical micro-F1 across encoder×compression cells at r=0.95, ρ=0.89 (six cells, preliminary, one dataset).

  16. TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

    cs.CV 2026-06 unverdicted novelty 6.0

    TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class glioblastoma response classification on longitudinal 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.

  17. TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

    cs.CV 2026-06 unverdicted novelty 6.0

    TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.

  18. Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation

    cs.CV 2026-06 unverdicted novelty 6.0

    Venice-H1 improves failure-case mIoU by 0.89-1.40 points in referring image segmentation via multi-scale grid signatures and a failure-aware re-ranker, with positive CIs on all tested pairs and low harmful-switch rates.

  19. ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

    cs.CV 2026-05 unverdicted novelty 6.0

    ASAP introduces an anatomy-aware semantically-adaptive pre-training method for medical volumetric vision-language models and reports state-of-the-art results on a new benchmark spanning 15 datasets and 22 tasks.

  20. Astra: a generalizable report generation foundation model for 3D computed tomography

    cs.CV 2026-05 unverdicted novelty 6.0

    Astra is a 3D CT vision-language foundation model trained on 90,678 thoracoabdominal scans that claims 44.1% better diagnostic metrics on internal and six external cohorts plus 29.6% faster chest reporting in real workflows.

  21. MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0

    MedVol-R1 is an RL framework that decouples 2D evidence grounding from 3D mask generation for volumetric reasoning segmentation and reports SOTA results on M3D-Seg benchmarks.

  22. Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0

    A unified autoregressive vision-language framework integrates segmentation, detection, and appearance reasoning for CT images via task-routing tokens and progressive refinement, with gains on public benchmarks.

  23. CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    CA-GCL adds global contrastive separation and clinical text augmentation to fine-grained vision-language pretraining, reducing textual embedding collapse and prompt variance in 3D medical image tasks.

  24. RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

    cs.CV 2026-05 unverdicted novelty 6.0

    RadThinking releases a large longitudinal CT VQA dataset stratified into foundation perception questions, single-rule reasoning questions, and compositional multi-step chains grounded in clinical reporting standards f...

  25. Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 6.0

    CT-SpatialVQA benchmark reveals that eight 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks from CT data, frequently below random performance.

  26. MedScribe: Clinically Grounded CT Reporting through Agentic Workflows

    cs.CV 2026-05 unverdicted novelty 6.0

    MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than stan...

  27. RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

    cs.AI 2026-04 conditional novelty 6.0

    An RL-trained tool-using agent improves chest CT report generation over CT-Chat by 5.8 macro-F1 points, 24.7 robustness points, and 37% faithfulness while exposing intermediate tool traces.

  28. RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

    cs.AI 2026-04 unverdicted novelty 6.0

    RadAgent generates stepwise, tool-augmented chest CT reports with traceable decisions, improving accuracy, robustness, and adding a 37% faithfulness score absent in standard 3D VLMs.

  29. Representation geometry shapes task performance in vision-language modeling for CT enterography

    cs.CV 2026-04 unverdicted novelty 6.0

    Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.

  30. Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis

    cs.CV 2026-04 unverdicted novelty 6.0

    Transferring a 2D MLLM to 3D CT inputs via parameter reuse, a Text-Guided Hierarchical MoE framework, and two-stage training yields better performance than prior 3D medical MLLMs on medical report generation and visua...

  31. Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks

    cs.CV 2026-04 conditional novelty 6.0

    VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models...

  32. Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks

    cs.CV 2026-04 unverdicted novelty 6.0

    LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.

  33. Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

    cs.CR 2026-03 conditional novelty 6.0

    DiffVP turns scan-to-normal semantic discrepancies into learnable visual prefix tokens that guide an LLM to write more accurate, fine-grained 3D CT reports.

  34. Medical Image Spatial Grounding with Semantic Sampling

    cs.CV 2026-03 conditional novelty 6.0

    MIS-Ground stress-tests 3D medical spatial grounding in VLMs, and MIS-SemSam raises Qwen3-VL-32B accuracy on it by 13.06% via semantic-neighborhood decoding.

  35. Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

    cs.CV 2026-03 conditional novelty 6.0

    SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.

  36. Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

    cs.CV 2026-07 conditional novelty 5.0

    A 19M-parameter JEPA-style 3D-CT encoder with a routed Mamba+GQA hybrid and orthogonal hidden-state regularization gives a 4B total model the best mean accuracy on M3D-VQA closed-ended questions and the best average o...

  37. Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology

    cs.AI 2026-07 reject novelty 5.0

    A multi-LLM review pipeline produces a synthetic 3D MRI-text dataset on which a VQ-GAN-perceiver-Vicuna model reports large gains over 2D and 3D baselines in brain-tumor report generation and VQA.

  38. Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

    cs.CV 2026-07 reject novelty 5.0

    A synthetic chain-of-thought dataset generated from CT reports lets a 2D-pretrained medical MLLM improve on 3D CT spatial-reasoning benchmarks.

  39. E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

    eess.IV 2026-06 unverdicted novelty 5.0

    E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.

  40. An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation

    eess.IV 2026-05 unverdicted novelty 5.0

    A multi-center whole-body FDG PET/CT foundation model with early fusion and masked autoencoding pretraining achieves label-efficient tumor segmentation on downstream tasks.

  41. Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

    cs.CV 2026-05 unverdicted novelty 5.0

    TIF-GRPO uses integral feedback on pseudo-temporal trajectories to regulate anatomy-aware rewards in RL for clinical faithfulness in volumetric CT analysis.

  42. M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification

    cs.CV 2026-05 conditional novelty 5.0

    M3Net achieves state-of-the-art accuracies of 86.96% on LIDC-IDRI and 84.24% on USTC-FHLN for pulmonary nodule classification using a hierarchical multi-scale 3D network with cross-scale consistency.

  43. Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

    cs.CV 2026-07 conditional novelty 4.0

    Direct answer-only supervised fine-tuning is the most robust adaptation family on MedFrameQA, beating frozen baselines by ~6 points and outperforming complex variants on seed stability.

  44. UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

    cs.CV 2026-06 unverdicted novelty 4.0

    UniReason-Med introduces a unified framework for 2D and 3D medical VQA with shared grounded reasoning, trained on a 220K dataset, claiming that joint 2D+3D supervision improves 3D performance over 3D-only training.

  45. Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes

    cs.CV 2026-06 unverdicted novelty 4.0

    LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.

  46. CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

    cs.CV 2026-05 unverdicted novelty 4.0

    CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.

  47. Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

    cs.CR 2026-03 unverdicted novelty 3.0

    A unified multi-modal NIDS dataset from CIC-IDS-2017, CIC-IoT-2023, UNSW-NB15 and CIC-DDoS-2019 is used to train stable ML attack classifiers and to generate synthetic data whose fidelity and utility are assessed via ...