Pith. sign in

REVIEW 4 cited by

Triad: Vision Foundation Model for 3D Magnetic Resonance Imaging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14064 v2 pith:WQC74WPU submitted 2025-02-19 cs.CV cs.AI

classification cs.CVcs.AI
keywords tasksacrossdatasetstriaddatafoundationimagingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting performance across a broad range of applications. However, existing vision foundation models that claim to be applicable to various clinical tasks are mostly pre-trained on 3D computed tomography (CT), which benefits from the availability of extensive 3D CT databases. Significant differences between CT and magnetic resonance imaging (MRI) in imaging principles, signal characteristics, and data distribution may hinder their practical performance and versatility in MRI-specific applications. Here, we propose Triad, a vision foundation model for 3D MRI. Triad adopts a widely used autoencoder architecture to learn robust representations from 131,170 3D MRI volumes and uses organ-independent imaging descriptions to constrain the semantic distribution of the visual modality. The above pre-training dataset is called Triad-131K, which is currently the largest 3D MRI pre-training dataset. We evaluate Triad across three tasks, namely, organ/tumor segmentation, organ/cancer classification, and medical image registration, in two data modalities (within-domain and out-of-domain) settings using 25 downstream datasets. By initializing models with Triad's pre-trained weights, nnUNet-Triad improves segmentation performance by 2.51% compared to nnUNet-Scratch across 17 datasets. Swin-B-Triad achieves a 3.97% improvement over Swin-B-Scratch in classification tasks across five datasets. SwinUNETR-Triad improves by 4.00% compared to SwinUNETR-Scratch in registration tasks across two datasets. Our study demonstrates that pre-training can improve performance when the data modalities and organs of upstream and downstream tasks are consistent.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

    cs.CV 2026-07 conditional novelty 6.5 of 10

    MST-based local boundary leakage and global topology divergence, fused by task complexity, rank SSL 3D medical encoders for segmentation without fine-tuning, beating prior TE metrics by 0.36 weighted Kendall τ at 56× speed.

  2. Pan-FM: A Pan-Organ Foundation Model with Saliency-Guided Masking for Missing Robustness

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Pan-FM learns balanced representations across seven organs by adaptively masking dominant organs during pre-training, yielding stronger disease prediction and missing-organ robustness than single-organ or naive multim...

  3. Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning

    eess.IV 2025-09 unverdicted novelty 5.0 of 10

    ASSFT combines active test-time sample selection via diversified knowledge divergence and anatomical segmentation difficulty with selective semi-supervised fine-tuning to adapt medical vision foundation models for vol...

  4. Capabilities of GPT-5 on Multimodal Medical Reasoning

    cs.CL 2025-08 reject novelty 4.0 of 10

    A benchmark study reports GPT-5 outperforming GPT-4o and pre-licensed human experts on most medical QA tasks, but not consistently on VQA-RAD.

Pith tools