NeuroQA is a large-scale 3D brain MRI visual question answering benchmark with verified image-grounded QA pairs, multi-domain coverage, and baseline evaluations showing current models lag behind text-only performance.
hub Canonical reference
M3d: Advancing 3d medical image analysis with multi-modal large language models
Canonical reference. 80% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
years
2026 29representative citing papers
DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.
MRI2Rep generates LI-RADS structured reports from 3D liver MRI via autoregressive modeling on 3929 real-world pairs, reporting 76% case-level sensitivity and 70-75% clinical acceptability in reader study.
CardioLens is a leakage-resistant CMR testbed of 473k slices and 13k QA pairs showing current MLLMs exhibit a large clinical reality gap with category-collapse failures on real workflows.
SliceWorld introduces a world-state model for CT report generation that uses predictive and factor-aware objectives on axial slice sequences.
Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.
Agentic LLMs autonomously execute complex neuro-radiological workflows like glioma segmentation and multi-timepoint response assessment by directing off-the-shelf tools, without any model training.
IBISAgent enables MLLMs to perform iterative pixel-level visual reasoning for biomedical object referring and segmentation via text-based clicks and agentic RL, outperforming prior SOTA methods without model modifications.
Astra generates style-consistent multi-organ CT reports that generalize across institutions via report harmonization plus GRPO reinforcement learning, improving clinical drafting and enabling synthetic pretraining.
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
Venice-H1 improves failure-case mIoU by 0.89-1.40 points in referring image segmentation via multi-scale grid signatures and a failure-aware re-ranker, with positive CIs on all tested pairs and low harmful-switch rates.
ASAP introduces an anatomy-aware semantically-adaptive pre-training method for medical volumetric vision-language models and reports state-of-the-art results on a new benchmark spanning 15 datasets and 22 tasks.
MedVol-R1 is an RL framework that decouples 2D evidence grounding from 3D mask generation for volumetric reasoning segmentation and reports SOTA results on M3D-Seg benchmarks.
A unified autoregressive vision-language framework integrates segmentation, detection, and appearance reasoning for CT images via task-routing tokens and progressive refinement, with gains on public benchmarks.
RadThinking releases a large longitudinal CT VQA dataset stratified into foundation perception questions, single-rule reasoning questions, and compositional multi-step chains grounded in clinical reporting standards for cancer screening.
CT-SpatialVQA benchmark reveals that eight 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks from CT data, frequently below random performance.
MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than standard VLMs on CT-RATE and RadChestCT.
An RL-trained tool-using agent improves chest CT report generation over CT-Chat by 5.8 macro-F1 points, 24.7 robustness points, and 37% faithfulness while exposing intermediate tool traces.
Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.
Transferring a 2D MLLM to 3D CT inputs via parameter reuse, a Text-Guided Hierarchical MoE framework, and two-stage training yields better performance than prior 3D medical MLLMs on medical report generation and visual question answering.
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models even on report generation.
LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.
E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.
A multi-center whole-body FDG PET/CT foundation model with early fusion and masked autoencoding pretraining achieves label-efficient tumor segmentation on downstream tasks.
citing papers explorer
-
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding
NeuroQA is a large-scale 3D brain MRI visual question answering benchmark with verified image-grounded QA pairs, multi-domain coverage, and baseline evaluations showing current models lag behind text-only performance.
-
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.
-
MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI
MRI2Rep generates LI-RADS structured reports from 3D liver MRI via autoregressive modeling on 3929 real-world pairs, reporting 76% case-level sensitivity and 70-75% clinical acceptability in reader study.
-
CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations
CardioLens is a leakage-resistant CMR testbed of 473k slices and 13k QA pairs showing current MLLMs exhibit a large clinical reality gap with category-collapse failures on real workflows.
-
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
SliceWorld introduces a world-state model for CT report generation that uses predictive and factor-aware objectives on axial slice sequences.
-
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs
Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.
-
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
Agentic LLMs autonomously execute complex neuro-radiological workflows like glioma segmentation and multi-timepoint response assessment by directing off-the-shelf tools, without any model training.
-
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
IBISAgent enables MLLMs to perform iterative pixel-level visual reasoning for biomedical object referring and segmentation via text-based clicks and agentic RL, outperforming prior SOTA methods without model modifications.
-
Astra: a generalizable report generation foundation model for 3D computed tomography
Astra generates style-consistent multi-organ CT reports that generalize across institutions via report harmonization plus GRPO reinforcement learning, improving clinical drafting and enabling synthetic pretraining.
-
TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
-
Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation
Venice-H1 improves failure-case mIoU by 0.89-1.40 points in referring image segmentation via multi-scale grid signatures and a failure-aware re-ranker, with positive CIs on all tested pairs and low harmful-switch rates.
-
ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training
ASAP introduces an anatomy-aware semantically-adaptive pre-training method for medical volumetric vision-language models and reports state-of-the-art results on a new benchmark spanning 15 datasets and 22 tasks.
-
MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation
MedVol-R1 is an RL framework that decouples 2D evidence grounding from 3D mask generation for volumetric reasoning segmentation and reports SOTA results on M3D-Seg benchmarks.
-
Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning
A unified autoregressive vision-language framework integrates segmentation, detection, and appearance reasoning for CT images via task-routing tokens and progressive refinement, with gains on public benchmarks.
-
RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology
RadThinking releases a large longitudinal CT VQA dataset stratified into foundation perception questions, single-rule reasoning questions, and compositional multi-step chains grounded in clinical reporting standards for cancer screening.
-
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
CT-SpatialVQA benchmark reveals that eight 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks from CT data, frequently below random performance.
-
MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than standard VLMs on CT-RATE and RadChestCT.
-
RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
An RL-trained tool-using agent improves chest CT report generation over CT-Chat by 5.8 macro-F1 points, 24.7 robustness points, and 37% faithfulness while exposing intermediate tool traces.
-
Representation geometry shapes task performance in vision-language modeling for CT enterography
Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.
-
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
Transferring a 2D MLLM to 3D CT inputs via parameter reuse, a Text-Guided Hierarchical MoE framework, and two-stage training yields better performance than prior 3D medical MLLMs on medical report generation and visual question answering.
-
Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models even on report generation.
-
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.
-
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
E-MRL trains VLMs via RL on a diagnosis-localization-verification MDP with a novel cross-view consistency reward to ground 3D tumor reports in verifiable CT slices.
-
An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation
A multi-center whole-body FDG PET/CT foundation model with early fusion and masked autoencoding pretraining achieves label-efficient tumor segmentation on downstream tasks.
-
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis
TIF-GRPO uses integral feedback on pseudo-temporal trajectories to regulate anatomy-aware rewards in RL for clinical faithfulness in volumetric CT analysis.
-
M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification
M3Net achieves state-of-the-art accuracies of 86.96% on LIDC-IDRI and 84.24% on USTC-FHLN for pulmonary nodule classification using a hierarchical multi-scale 3D network with cross-scale consistency.
-
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA
UniReason-Med introduces a unified framework for 2D and 3D medical VQA with shared grounded reasoning, trained on a 220K dataset, claiming that joint 2D+3D supervision improves 3D performance over 3D-only training.
-
Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes
LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.
-
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.