REVIEW 5 cited by
Read Like a Radiologist: Efficient Vision-Language Model for 3D Medical Imaging Interpretation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs.
Forward citations
Cited by 5 Pith papers
-
Disorder-induced stress-flow misalignment in soft glassy materials revealed using multi-directional shear
Soft glassy materials show a transient stress response orthogonal to a newly applied shear direction, which a mesoscopic elasto-plastic model attributes to local yield-stress disorder.
-
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
A cross-modal graph connecting CT slices and question tokens, aggregated by an attentive GCN, improves LLM-based CT visual question answering on M3D-VQA.
-
Aligning Proteins and Language: A Foundation Model for Protein Retrieval
A CLIP-style model aligns protein surface point clouds with GO-derived captions and achieves roughly 60% Top-5 zero-shot retrieval on PDB and 36% on EMDB.
-
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
CT-Agent combines an LLM planner, region-specific LoRA adapters, and global/local token compression to improve 3D chest CT report generation and question answering on CT-RATE and RadGenome-ChestCT.
-
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.
Discussion (0). Sign in to comment.