Pith. sign in

REVIEW 2 cited by

T3D: Advancing 3D Medical Vision-Language Pre-training by Learning Multi-View Visual Consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01529 v3 pith:KKZL3HXJ submitted 2023-12-03 cs.CV cs.CLcs.LGeess.IV

classification cs.CVcs.CLcs.LGeess.IV
keywords visualmedicalacrossmedvlprepresentationsalignmentbenchmarkconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile, 3D medical vision-language pre-training (MedVLP) remains underexplored due to the lack of a large-scale, publicly available 3D medical image-report dataset. To bridge this gap, we introduce **CT-3DVLP**, the first and largest **public** 3D volume-report dataset, establishing a comprehensive benchmark for 3D MedVLP research. Meanwhile, we propose the **T3D** framework, which enhances 3D MedVLP beyond naive CLIP-style alignment that directly pairs volumes with reports but neglects local visual representations. Instead, we introduce **Text-informed Multi-view Alignment (TMA)**, a novel approach that clusters volumetric data while enforcing consistency across different views of the same volume-report pair. TMA integrates textual features into fine-grained visual representations, ensuring contextual coherence across views. We evaluate T3D across multiple downstream tasks in both unimodal and cross-modal settings, including zero-shot and fine-tuned classification, cross-modal retrieval, report generation, and semantic segmentation. Our results show that T3D consistently outperforms existing vSSL and multimodal methods, demonstrating superior zero-shot and fine-tuning capabilities and setting a new benchmark for 3D medical image understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual form Complementary Masking for Domain-Adaptive Image Segmentation

    cs.CV 2025-07 reject novelty 5.0 of 10

    The paper proposes complementary masking consistency for UDA segmentation and reports empirical gains, but its theoretical proof contains a direct internal contradiction.

  2. Conditional Latent Coding with Learnable Synthesized Reference for Deep Image Compression

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Conditional Latent Coding compresses images by synthesizing a per-image reference latent from a learned feature dictionary, improving low-bitrate rate-distortion over TCM, VTM, and BPG.

Pith tools