Pith. sign in

REVIEW 20 cited by

Vision Foundation Models for Computed Tomography

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09001 v2 pith:ILANTPEV submitted 2025-01-15 eess.IV cs.CV

classification eess.IVcs.CV
keywords acrossct-fmmodelsfoundationimagingtaskscomputeddata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    CT-Lite combines Feature Attention Style Transfer (FAST) and Structured Factorized Projections (SFP) with contrastive learning to reach AUROC within 5-7% of uncompressed baselines on compressed CT volumes across three...

  2. Semantically Calibrated Evidence Composition for CT Vision-Language Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SCOPE composes organ-level CT evidence under whole-volume context and calibrates it with diagnostic-summary supervision, reaching 85.0/72.2 macro AUC on CT-RATE and RadChestCT.

  3. Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Cross-patient report-based pair mining plus burden-direction alignment improves CT vision-language pretraining, reaching 85.6 AUROC on CT-RATE zero-shot diagnosis.

  4. Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Existing 3D CT foundation models generalize poorly to external head and neck cancer cohorts; CT-CLIP was the most robust, and gating fusion of imaging with clinical data gave the best prediction.

  5. DALE-CT: Depth-Aware Foundation Models for Computed Tomography

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.

  6. Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    FlexiCT applies agglomerative pretraining across 2D, 3D, and vision-language stages on a large public CT collection to produce representations that match or exceed specialized models on segmentation, classification, r...

  7. JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    JANUS conditions Vision Transformer embeddings on macro-radiomic priors via anatomically guided gating, reaching macro-AUROC 0.88 on an internal test set of 5082 cases and 0.87 on an external set of 2000 cases while i...

  8. Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

    cs.CV 2026-05 conditional novelty 6.0 of 10

    A student model trained on JPEG-compressed chest CT via attention-style distillation and a factorized projection head reaches AUROC within 3–5 percentage points of the uncompressed-teacher baseline on three datasets, ...

  9. Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks

    cs.CV 2026-04 conditional novelty 6.0 of 10

    VoxelFM learns robust 3D CT visual features via DINO self-distillation that transfer effectively to seven clinical task categories using frozen backbones and lightweight heads, outperforming prior CT foundation models...

  10. M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

    cs.CV 2025-09 conditional novelty 6.0 of 10

    One self-supervised encoder trained on unpaired X-ray, ultrasound, endoscopy, and CT data gives competitive zero-shot retrieval and seems to generalize to unseen MRI tasks.

  11. Disorder-induced stress-flow misalignment in soft glassy materials revealed using multi-directional shear

    cond-mat.soft 2025-08 unverdicted novelty 6.0 of 10

    Soft glassy materials show a transient stress response orthogonal to a newly applied shear direction, which a mesoscopic elasto-plastic model attributes to local yield-stress disorder.

  12. CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A new public dataset of 22,022 CT volumes labeled for 167 structures, and a nnU-Net model trained on it, outperform TotalSegmentator on most shared structures and expand coverage.

  13. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  14. Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Bifurcation Connectedness Score (BCS) measures junction-level vessel connectivity that Dice misses, tracks geometric FFR-CT decision agreement in severe disease, and shows branch recovery and tree connectedness are se...

  15. Anatomy Contextualized Adaption of CT Foundation Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A lightweight inter-anatomy transformer on frozen CT foundation embeddings plus dual anatomy/scan contrastive losses beats global and fine-grained baselines on Merlin and CT-RATE zero-shot finding classification.

  16. Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

    cs.CV 2026-07 conditional novelty 5.0 of 10

    CT Foundation embeddings predict 2-year distant metastasis in head and neck cancer with AUC 0.791, matching radiomics-plus-deep-learning (0.794) while requiring no tumor contours.

  17. Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    FlexiCT provides CT foundation models via agglomerative pretraining on 266227 volumes from 56 datasets that match or exceed task-specific models on five task families while organizing embeddings along tumor-stage gradients.

  18. Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning

    eess.IV 2025-09 unverdicted novelty 5.0 of 10

    ASSFT combines active test-time sample selection via diversified knowledge divergence and anatomical segmentation difficulty with selective semi-supervised fine-tuning to adapt medical vision foundation models for vol...

  19. Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing

    physics.med-ph 2026-07 conditional novelty 4.0 of 10

    Limited bedside impact of medical imaging AI stems from structural misalignment with clinical decision-making across six dimensions, not mainly from weak algorithms, regulation, or explainability.

  20. Benchmarking Foundation Models for Renal Lesion Stratification in CT

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    Medical foundation models match a ResNet-50 but are outperformed by radiomics (AUC 0.88) on external validation for renal lesion stratification in CT.

Pith tools