REVIEW 5 cited by
E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The development of 3D medical vision-language models holds significant potential for disease diagnosis and patient treatment. However, compared to 2D medical images, 3D medical images, such as CT scans, face challenges related to limited training data and high dimension, which severely restrict the progress of 3D medical vision-language models. To address these issues, we collect a large amount of unlabeled 3D CT data and utilize self-supervised learning to construct a 3D visual foundation model for extracting 3D visual features. Then, we apply 3D spatial convolutions to aggregate and project high-level image features, reducing computational complexity while preserving spatial information. We also construct two instruction-tuning datasets based on BIMCV-R and CT-RATE to fine-tune the 3D vision-language model. Our model demonstrates superior performance compared to existing methods in report generation, visual question answering, and disease diagnosis. Code and data will be made publicly available soon.
Forward citations
Cited by 5 Pith papers
-
Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
A volumetric vision-language model trained jointly on classification labels and segmentation masks from three CT datasets reaches 83% AUROC on CT-RATE and shows cross-dataset zero-shot behavior.
-
Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
A train-free method that compresses 3D medical volumes into small embeddings via a frozen 2D foundation model and random projections, outperforming several medical-volume pretrained models on benchmark tasks.
-
Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
A multi-LLM review pipeline produces a synthetic 3D MRI-text dataset on which a VQ-GAN-perceiver-Vicuna model reports large gains over 2D and 3D baselines in brain-tumor report generation and VQA.
-
MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms
Morphology-aware masking plus cross-modal ECG–SpO2 pretraining on MIMIC yields stronger transfer than MAE, contrastive, Barlow Twins, and JEPA on several clinical prediction tasks.
-
HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding
HSENet improves 3D CT vision-language understanding by combining global and local 3D encoders with a centroid-based spatial token compressor, posting state-of-the-art results on CT-RATE and RadGenome-ChestCT.
Discussion (0). Sign in to comment.