DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.
hub
arXiv preprint arXiv:2403.17834 , year=
19 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Cached structured report annotations let radiology label schemas be changed with dictionary edits instead of relabeling the corpus, recovering long-tail findings at near-zero marginal cost.
Introduces CORTEX benchmark supplying 76,177 validated four-stage diagnostic reasoning traces for open/closed VQA and report generation on chest CT to enable traceable MLLM supervision and evaluation.
ConQuer augments global CLIP alignment with independent per-concept contrastive losses on anatomical regions extracted from reports, producing Jolia which outperforms CLIP baselines on classification, report generation, and transfer.
ORACLE-CT improves CT classification performance by using anatomy-specific support pooling based on multi-organ segmentation, showing gains in AUROC on internal and external datasets.
A foundation VAE pretrained on natural images and videos serves as a frozen interface for CT reconstruction, augmentation, and generation, yielding 3.9% NSD gains in segmentation and improved generation metrics across 18 diseases.
Pretrained autoencoders in medical latent diffusion encode discriminative features well for reconstruction but structure their latent spaces in ways that hinder classifier learning, a gap that persists across architectures and is not closed by domain fine-tuning.
MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than standard VLMs on CT-RATE and RadChestCT.
CheXmix combines masked autoencoder pretraining with early-fusion generative modeling to outperform prior models on chest X-ray classification by up to 8.6% AUROC, inpainting by 51%, and report generation by 45% on GREEN.
Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.
SUMI distills photon-counting CT quality into routine chest CT by learning to reverse clinically validated acquisition degradations, yielding 15-20% gains in image metrics, better radiologist utility, and up to 15% higher lesion detection sensitivity.
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models even on report generation.
Frozen CT-CLIP representations plus a lightweight DeepSurv head outperform CoxPH and match or beat other multimodal baselines for lung-cancer survival on a 242-patient real-world cohort.
TIF-GRPO uses integral feedback on pseudo-temporal trajectories to regulate anatomy-aware rewards in RL for clinical faithfulness in volumetric CT analysis.
RadGenome-Anatomy is a large-scale chest radiograph dataset with anatomy labels obtained by projecting 3D CT masks into 2D radiographic space for 210 structures in 25,692 studies.
A vision-language model pre-trained via instruction tuning on CT-report pairs improves survival prediction accuracy over baselines, especially when clinical data alone is weak, while also producing text answers to clinical questions.
LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.
MedGemma 1.5 4B reports absolute gains of 11% on 3D MRI classification, 3% on 3D CT, 47% macro F1 on pathology slides, 35% IoU on anatomical localization, and 5-22% on clinical QA tasks over MedGemma 1.
citing papers explorer
-
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
DeepTumorVQA is a new stage-wise 3D CT VQA benchmark showing that quantitative measurement is the main failure point for current medical VLMs and that tool augmentation substantially improves later reasoning stages.
-
Reconfigurable Radiology Labels Without Relabeling
Cached structured report annotations let radiology label schemas be changed with dictionary edits instead of relabeling the corpus, recovering long-tail findings at near-zero marginal cost.
-
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Introduces CORTEX benchmark supplying 76,177 validated four-stage diagnostic reasoning traces for open/closed VQA and report generation on chest CT to enable traceable MLLM supervision and evaluation.
-
Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning
ConQuer augments global CLIP alignment with independent per-concept contrastive losses on anatomical regions extracted from reports, producing Jolia which outperforms CLIP baselines on classification, report generation, and transfer.
-
ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification
ORACLE-CT improves CT classification performance by using anatomy-specific support pooling based on multi-organ segmentation, showing gains in AUROC on internal and external datasets.
-
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
A foundation VAE pretrained on natural images and videos serves as a frozen interface for CT reconstruction, augmentation, and generation, yielding 3.9% NSD gains in segmentation and improved generation metrics across 18 diseases.
-
The Learnability Gap in Medical Latent Diffusion
Pretrained autoencoders in medical latent diffusion encode discriminative features well for reconstruction but structure their latent spaces in ways that hinder classifier learning, a gap that persists across architectures and is not closed by domain fine-tuning.
-
MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
MedScribe reformulates CT radiology reporting as an agentic evidence-acquisition workflow using LLM-invoked diagnostic tools and pathology-aligned retrieval, yielding higher clinical accuracy and consistency than standard VLMs on CT-RATE and RadChestCT.
-
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
CheXmix combines masked autoencoder pretraining with early-fusion generative modeling to outperform prior models on chest X-ray classification by up to 8.6% AUROC, inpainting by 51%, and report generation by 45% on GREEN.
-
Representation geometry shapes task performance in vision-language modeling for CT enterography
Mean pooling and multi-window RGB encoding optimize vision-language performance on CT enterography, with retrieval-augmented generation substantially improving automated report severity accuracy over fine-tuning alone.
-
Distilling Photon-Counting CT into Routine Chest CT through Clinically Validated Degradation Modeling
SUMI distills photon-counting CT quality into routine chest CT by learning to reverse clinically validated acquisition degradations, yielding 15-20% gains in image metrics, better radiologist utility, and up to 15% higher lesion detection sensitivity.
-
Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models even on report generation.
-
CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction
Frozen CT-CLIP representations plus a lightweight DeepSurv head outperform CoxPH and match or beat other multimodal baselines for lung-cancer survival on a 242-patient real-world cohort.
-
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis
TIF-GRPO uses integral feedback on pseudo-temporal trajectories to regulate anatomy-aware rewards in RL for clinical faithfulness in volumetric CT analysis.
-
RadGenome-Anatomy: A Large-Scale Anatomy-Labeled Chest Radiograph Dataset via Physically Grounded Volumetric Projection
RadGenome-Anatomy is a large-scale chest radiograph dataset with anatomy labels obtained by projecting 3D CT masks into 2D radiographic space for 210 structures in 25,692 studies.
-
Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning
A vision-language model pre-trained via instruction tuning on CT-report pairs improves survival prediction accuracy over baselines, especially when clinical data alone is weak, while also producing text answers to clinical questions.
-
Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes
LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.
-
MedGemma 1.5 Technical Report
MedGemma 1.5 4B reports absolute gains of 11% on 3D MRI classification, 3% on 3D CT, 47% macro F1 on pathology slides, 35% IoU on anatomical localization, and 5-22% on clinical QA tasks over MedGemma 1.
- INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT