REVIEW 20 cited by
Vision Foundation Models for Computed Tomography
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.
Forward citations
Cited by 20 Pith papers
-
Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
CT-Lite combines Feature Attention Style Transfer (FAST) and Structured Factorized Projections (SFP) with contrastive learning to reach AUROC within 5-7% of uncompressed baselines on compressed CT volumes across three...
-
Semantically Calibrated Evidence Composition for CT Vision-Language Learning
SCOPE composes organ-level CT evidence under whole-volume context and calibrates it with diagnostic-summary supervision, reaching 85.0/72.2 macro AUC on CT-RATE and RadChestCT.
-
Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining
Cross-patient report-based pair mining plus burden-direction alignment improves CT vision-language pretraining, reaching 85.6 AUROC on CT-RATE zero-shot diagnosis.
-
Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction
Existing 3D CT foundation models generalize poorly to external head and neck cancer cohorts; CT-CLIP was the most robust, and gating fusion of imaging with clinical data gave the best prediction.
-
DALE-CT: Depth-Aware Foundation Models for Computed Tomography
DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.
-
Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
FlexiCT applies agglomerative pretraining across 2D, 3D, and vision-language stages on a large public CT collection to produce representations that match or exceed specialized models on segmentation, classification, r...
-
JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift
JANUS conditions Vision Transformer embeddings on macro-radiomic priors via anatomically guided gating, reaching macro-AUROC 0.88 on an internal test set of 5082 cases and 0.87 on an external set of 2000 cases while i...
-
Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
A student model trained on JPEG-compressed chest CT via attention-style distillation and a factorized projection head reaches AUROC within 3–5 percentage points of the uncompressed-teacher baseline on three datasets, ...
-
Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks
VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models...
-
M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
One self-supervised encoder trained on unpaired X-ray, ultrasound, endoscopy, and CT data gives competitive zero-shot retrieval and seems to generalize to unseen MRI tasks.
-
Disorder-induced stress-flow misalignment in soft glassy materials revealed using multi-directional shear
Soft glassy materials show a transient stress response orthogonal to a newly applied shear direction, which a mesoscopic elasto-plastic model attributes to local yield-stress disorder.
-
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
A new public dataset of 22,022 CT volumes labeled for 167 structures, and a nnU-Net model trained on it, outperform TotalSegmentator on most shared structures and expand coverage.
-
4KAgent: Agentic Any Image to 4K Super-Resolution
An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.
-
Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement
Bifurcation Connectedness Score (BCS) measures junction-level vessel connectivity that Dice misses, tracks geometric FFR-CT decision agreement in severe disease, and shows branch recovery and tree connectedness are se...
-
Anatomy Contextualized Adaption of CT Foundation Models
A lightweight inter-anatomy transformer on frozen CT foundation embeddings plus dual anatomy/scan contrastive losses beats global and fine-grained baselines on Merlin and CT-RATE zero-shot finding classification.
-
Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer
CT Foundation embeddings predict 2-year distant metastasis in head and neck cancer with AUC 0.791, matching radiomics-plus-deep-learning (0.794) while requiring no tumor contours.
-
Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
FlexiCT provides CT foundation models via agglomerative pretraining on 266227 volumes from 56 datasets that match or exceed task-specific models on five task families while organizing embeddings along tumor-stage gradients.
-
Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning
ASSFT combines active test-time sample selection via diversified knowledge divergence and anatomical segmentation difficulty with selective semi-supervised fine-tuning to adapt medical vision foundation models for vol...
-
Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing
Limited bedside impact of medical imaging AI stems from structural misalignment with clinical decision-making across six dimensions, not mainly from weak algorithms, regulation, or explainability.
-
Benchmarking Foundation Models for Renal Lesion Stratification in CT
Medical foundation models match a ResNet-50 but are outperformed by radiomics (AUC 0.88) on external validation for renal lesion stratification in CT.
Discussion (0). Sign in to comment.