Pith. sign in

REVIEW 12 cited by

Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14548 v1 pith:7QO2Y7KS submitted 2025-01-24 cs.CV

Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

classification cs.CV
keywords imageinterpretationmedicalcontrastivediagnosisfine-grainedimagesmodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Artificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 7.0

    CT-SpatialVQA benchmark shows 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks in CT volumes, often below random chance.

  2. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

    cs.CV 2026-07 conditional novelty 6.0

    ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.

  3. Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Disease-probe AUROC on frozen 3D-CT tokens predicts report-generation clinical micro-F1 across encoder×compression cells at r=0.95, ρ=0.89 (six cells, preliminary, one dataset).

  4. When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models?

    cs.CV 2026-07 conditional novelty 6.0

    CARVE is a label-free, cardinality-aware test-time adaptation method that consistently improves multi-label CT diagnosis when the base model is already discriminative and input depth matches pretraining.

  5. Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

    cs.CV 2026-07 conditional novelty 6.0

    Injecting multi-expert anatomy/lesion segmentation priors into vision–language alignment and calibrating text attention with lesion masks yields broad CT diagnosis plus specialist-level tumor performance and lesion grounding.

  6. ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification

    cs.CV 2026-06 unverdicted novelty 6.0

    ORACLE-CT improves CT classification performance by using anatomy-specific support pooling based on multi-organ segmentation, showing gains in AUROC on internal and external datasets.

  7. JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift

    cs.CV 2026-05 unverdicted novelty 6.0

    JANUS conditions Vision Transformer embeddings on macro-radiomic priors via anatomically guided gating, reaching macro-AUROC 0.88 on an internal test set of 5082 cases and 0.87 on an external set of 2000 cases while i...

  8. CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    CA-GCL adds global contrastive separation and clinical text augmentation to fine-grained vision-language pretraining, reducing textual embedding collapse and prompt variance in 3D medical image tasks.

  9. Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 6.0

    CT-SpatialVQA benchmark reveals that eight 3D medical VLMs achieve only 34% average accuracy on semantic-spatial reasoning tasks from CT data, frequently below random performance.

  10. Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

    cs.CV 2026-04 unverdicted novelty 6.0

    DCP-PD improves macro F1 scores on CT report generation benchmarks and introduces a hierarchical location-aware evaluation protocol that reveals ongoing challenges in pathology spatial grounding.

  11. Anatomy Contextualized Adaption of CT Foundation Models

    cs.CV 2026-07 conditional novelty 5.0

    A lightweight inter-anatomy transformer on frozen CT foundation embeddings plus dual anatomy/scan contrastive losses beats global and fine-grained baselines on Merlin and CT-RATE zero-shot finding classification.

  12. CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

    cs.CV 2026-05 unverdicted novelty 4.0

    CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.