Pith. sign in

REVIEW 5 cited by

Read Like a Radiologist: Efficient Vision-Language Model for 3D Medical Imaging Interpretation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.13558 v1 pith:ZQFA4EX2 submitted 2024-12-18 eess.IV cs.CLcs.CVcs.LG

classification eess.IVcs.CLcs.CVcs.LG
keywords medicalms-vlmimageimagesinterpretationimagingslicesvlms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disorder-induced stress-flow misalignment in soft glassy materials revealed using multi-directional shear

    cond-mat.soft 2025-08 unverdicted novelty 6.0 of 10

    Soft glassy materials show a transient stress response orthogonal to a newly applied shear direction, which a mesoscopic elasto-plastic model attributes to local yield-stress disorder.

  2. Computed Tomography Visual Question Answering with Cross-modal Feature Graphing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A cross-modal graph connecting CT slices and question tokens, aggregated by an attentive GCN, improves LLM-based CT visual question answering on M3D-VQA.

  3. Aligning Proteins and Language: A Foundation Model for Protein Retrieval

    q-bio.BM 2025-05 conditional novelty 5.0 of 10

    A CLIP-style model aligns protein surface point clouds with GO-derived captions and achieves roughly 60% Top-5 zero-shot retrieval on PDB and 36% on EMDB.

  4. CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CT-Agent combines an LLM planner, region-specific LoRA adapters, and global/local token compression to improve 3D chest CT report generation and question answering on CT-RATE and RadGenome-ChestCT.

  5. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

Pith tools