REVIEW 47 cited by
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
read the original abstract
Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images under-explored, despite their richer spatial information. This paper aims to advance 3D medical image analysis with MLLMs. To this end, we present a large-scale 3D multi-modal medical dataset, M3D-Data, comprising 120K image-text pairs and 662K instruction-response pairs specifically tailored for various 3D medical tasks, such as image-text retrieval, report generation, visual question answering, positioning, and segmentation. Additionally, we propose M3D-LaMed, a versatile multi-modal large language model for 3D medical image analysis. Furthermore, we introduce a new 3D multi-modal medical benchmark, M3D-Bench, which facilitates automatic evaluation across eight tasks. Through comprehensive evaluation, our method proves to be a robust model for 3D medical image analysis, outperforming existing solutions. All code, data, and models are publicly available at: https://github.com/BAAI-DCAI/M3D.
Forward citations
Cited by 47 Pith papers
-
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding
NeuroQA is a large-scale 3D brain MRI visual question answering benchmark with verified image-grounded QA pairs, multi-domain coverage, and baseline evaluations showing current models lag behind text-only performance.
-
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.
-
MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI
MRI2Rep generates LI-RADS structured reports from 3D liver MRI via autoregressive modeling on 3929 real-world pairs, reporting 76% case-level sensitivity and 70-75% clinical acceptability in reader study.
-
CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations
CardioLens is a leakage-resistant CMR testbed of 473k slices and 13k QA pairs showing current MLLMs exhibit a large clinical reality gap with category-collapse failures on real workflows.
-
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
SliceWorld introduces a world-state model for CT report generation that uses predictive and factor-aware objectives on axial slice sequences.
-
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
CT-SpatialVQA benchmark shows 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks in CT volumes, often below random chance.
-
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs
Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.
-
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
Agentic LLMs autonomously execute complex neuro-radiological workflows like glioma segmentation and multi-timepoint response assessment by directing off-the-shelf tools, without any model training.
-
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
IBISAgent enables MLLMs to perform iterative pixel-level visual reasoning for biomedical object referring and segmentation via text-based clicks and agentic RL, outperforming prior SOTA methods without model modifications.
-
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.
-
Astra: a generalizable report generation foundation model for 3D computed tomography
Astra generates style-consistent multi-organ CT reports that generalize across institutions via report harmonization plus GRPO reinforcement learning, improving clinical drafting and enabling synthetic pretraining.
-
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.
-
Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining
Cross-patient report-based pair mining plus burden-direction alignment improves CT vision-language pretraining, reaching 85.6 AUROC on CT-RATE zero-shot diagnosis.
-
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
Training-free multi-cue token compression for 3D medical VLMs that retains and merges tokens using attention, text similarity, and VFM saliency, maintaining diagnostic performance at 50–80% token retention.
-
Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models
Disease-probe AUROC on frozen 3D-CT tokens predicts report-generation clinical micro-F1 across encoder×compression cells at r=0.95, ρ=0.89 (six cells, preliminary, one dataset).
-
TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class glioblastoma response classification on longitudinal 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
-
TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
-
Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation
Venice-H1 improves failure-case mIoU by 0.89-1.40 points in referring image segmentation via multi-scale grid signatures and a failure-aware re-ranker, with positive CIs on all tested pairs and low harmful-switch rates.
-
ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training
ASAP introduces an anatomy-aware semantically-adaptive pre-training method for medical volumetric vision-language models and reports state-of-the-art results on a new benchmark spanning 15 datasets and 22 tasks.
-
Astra: a generalizable report generation foundation model for 3D computed tomography
Astra is a 3D CT vision-language foundation model trained on 90,678 thoracoabdominal scans that claims 44.1% better diagnostic metrics on internal and six external cohorts plus 29.6% faster chest reporting in real workflows.
-
MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation
MedVol-R1 is an RL framework that decouples 2D evidence grounding from 3D mask generation for volumetric reasoning segmentation and reports SOTA results on M3D-Seg benchmarks.
-
Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning
A unified autoregressive vision-language framework integrates segmentation, detection, and appearance reasoning for CT images via task-routing tokens and progressive refinement, with gains on public benchmarks.
-
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
CA-GCL adds global contrastive separation and clinical text augmentation to fine-grained vision-language pretraining, reducing textual embedding collapse and prompt variance in 3D medical image tasks.
-
RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology
RadThinking releases a large longitudinal CT VQA dataset stratified into foundation perception questions, single-rule reasoning questions, and compositional multi-step chains grounded in clinical reporting standards f...
-
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
CT-SpatialVQA benchmark reveals that eight 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks from CT data, frequently below random performance.
-
MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than stan...
-
RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
An RL-trained tool-using agent improves chest CT report generation over CT-Chat by 5.8 macro-F1 points, 24.7 robustness points, and 37% faithfulness while exposing intermediate tool traces.
-
RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
RadAgent generates stepwise, tool-augmented chest CT reports with traceable decisions, improving accuracy, robustness, and adding a 37% faithfulness score absent in standard 3D VLMs.
-
Representation geometry shapes task performance in vision-language modeling for CT enterography
Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.
-
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
Transferring a 2D MLLM to 3D CT inputs via parameter reuse, a Text-Guided Hierarchical MoE framework, and two-stage training yields better performance than prior 3D medical MLLMs on medical report generation and visua...
-
Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models...
-
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.
-
Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation
DiffVP turns scan-to-normal semantic discrepancies into learnable visual prefix tokens that guide an LLM to write more accurate, fine-grained 3D CT reports.
-
Medical Image Spatial Grounding with Semantic Sampling
MIS-Ground stress-tests 3D medical spatial grounding in VLMs, and MIS-SemSam raises Qwen3-VL-32B accuracy on it by 13.06% via semantic-neighborhood decoding.
-
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.
-
Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
A 19M-parameter JEPA-style 3D-CT encoder with a routed Mamba+GQA hybrid and orthogonal hidden-state regularization gives a 4B total model the best mean accuracy on M3D-VQA closed-ended questions and the best average o...
-
Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
A multi-LLM review pipeline produces a synthetic 3D MRI-text dataset on which a VQ-GAN-perceiver-Vicuna model reports large gains over 2D and 3D baselines in brain-tumor report generation and VQA.
-
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
A synthetic chain-of-thought dataset generated from CT reports lets a 2D-pretrained medical MLLM improve on 3D CT spatial-reasoning benchmarks.
-
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.
-
An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation
A multi-center whole-body FDG PET/CT foundation model with early fusion and masked autoencoding pretraining achieves label-efficient tumor segmentation on downstream tasks.
-
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis
TIF-GRPO uses integral feedback on pseudo-temporal trajectories to regulate anatomy-aware rewards in RL for clinical faithfulness in volumetric CT analysis.
-
M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification
M3Net achieves state-of-the-art accuracies of 86.96% on LIDC-IDRI and 84.24% on USTC-FHLN for pulmonary nodule classification using a hierarchical multi-scale 3D network with cross-scale consistency.
-
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
Direct answer-only supervised fine-tuning is the most robust adaptation family on MedFrameQA, beating frozen baselines by ~6 points and outperforming complex variants on seed stability.
-
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA
UniReason-Med introduces a unified framework for 2D and 3D medical VQA with shared grounded reasoning, trained on a 220K dataset, claiming that joint 2D+3D supervision improves 3D performance over 3D-only training.
-
Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes
LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.
-
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.
-
Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation
A unified multi-modal NIDS dataset from CIC-IDS-2017, CIC-IoT-2023, UNSW-NB15 and CIC-DDoS-2019 is used to train stable ML attack classifiers and to generate synthetic data whose fidelity and utility are assessed via ...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.