Pith. sign in

REVIEW 3 major objections 5 minor 4 cited by

A Survey on Computational Pathology Foundation Models: Datasets, Adaptation Strategies, and Evaluation Tasks

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey maps 28 computational pathology foundation models through three lenses: pre-training datasets, adaptation strategies, and evaluation tasks.

desk verdict Useful survey with up-to-date coverage, but Table 2 has verifiable availability errors and the 'comprehensive' claim needs explicit inclusion criteria. read the letter →

arxiv 2501.15724 v2 pith:7UVT3THU submitted 2025-01-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords computationalpathologyfoundationmodelsself-supervisedlearningwhole-slideimagesvision-languageevaluationtaxonomyhistopathologydatasetsadaptationstrategies
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to give the field of computational pathology foundation models a usable map: which datasets are used to pre-train them, how the models adapt general self-supervised learning frameworks to pathology, and how the research community evaluates them. It claims that although the models differ in scale and modality, they fall into two paradigms—uni-modal models trained only on images and multi-modal models that align images with text—and that their evaluations can be summarized into six task categories. A sympathetic reader would care because the field lacks standardized benchmarks; a structured comparison is a step toward measuring which of these models actually generalizes to clinical use.

What carries the argument

The organizing device is a two-axis classification: uni-modal versus multi-modal pre-training, crossed with a six-category taxonomy of evaluation tasks. The survey constructs an architecture-and-adaptation matrix that records each model's self-supervised backbone, parameter count, input type, and pre-training strategy, and a task taxonomy that lists which models were evaluated on each task. This matrix-plus-taxonomy structure is what carries the review's claims of comprehensiveness and comparability.

What would settle it

A reader could check Table 2 against the full list of publicly released computational pathology foundation models as of the survey's final revision, or check Figure 3's six categories against every evaluation task reported in the cited model papers; finding a released model with a distinct self-supervised backbone that is absent from the table, or an evaluation task that fits none of the six categories, would falsify the survey's coverage claim.

Watch

Extended reading notes

Core claim

The central claim is organizational: 28 current computational pathology foundation models can be systematically reviewed through three lenses—pre-training datasets, adaptation strategies, and evaluation tasks—and doing so reveals that the field has converged on a small number of self-supervised learning frameworks. Uni-modal models mostly adapt DINO, DINOv2, or masked image modeling frameworks, while multi-modal models mostly adapt CLIP or CoCa; their evaluations cluster into classification, retrieval, generation, segmentation, prediction, and visual question answering. The paper further argues that no standardized benchmark exists and that this lack is the main obstacle to comparing models across institutions and tasks.

Load-bearing premise

The survey's map is only as complete as the authors' informal selection of 28 models and six task categories; no systematic search strategy or inclusion criteria is given, so the claim of comprehensiveness rests on an unstated judgment about which models and tasks count.

Editorial extensions

If this is right

  • Researchers can use the six-category taxonomy to position a new model's evaluation and see which task categories are under-tested.
  • The dominance of DINOv2 for uni-modal models and CLIP/CoCa for multi-modal models suggests that future models will likely build on these backbones rather than invent entirely new pre-training frameworks.
  • The lack of standardized benchmarks implies that results across papers remain not directly comparable until a common evaluation suite is adopted.
  • Coverage gaps, such as few models trained on multiplex immunofluorescence images, point to specific data types where foundation models are still missing.
  • The survey's future directions imply that clinical deployment will require work on fairness, explainability, security, and transparency before these models can be trusted in practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own tables suggest that private, large-scale slide collections dominate the largest models, so the public datasets it lists may understate the field's real data advantage; that reading goes beyond what the paper states explicitly.
  • A testable extension is to convert the evaluation-task taxonomy into a checklist and use it to score future computational pathology foundation model papers for evaluation coverage.
  • The six task categories may overlap—report generation and visual question answering both test cross-modal understanding—so a future taxonomy might merge them or add a separate cross-modal reasoning axis.
  • Because the survey excludes generative pathology assistants and non-histopathology multimodal models, its map is deliberately narrower than all medical foundation models; readers should not generalize beyond the histopathology scope the paper defines.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey reviews computational pathology foundation models (CPathFMs), covering 28 models, their pre-training datasets, adaptation strategies, and evaluation tasks. It organizes models into uni-modal and multi-modal categories, tabulates pre-training datasets in Table 1, summarizes architectures and adaptation strategies in Table 2, and introduces a six-category taxonomy of evaluation tasks in Figure 3. The paper also discusses data, adaptation, and evaluation challenges and proposes future directions such as trustworthy CPathFMs, MxIF imaging, and standardized benchmarking.

Significance. A well-curated survey of CPathFMs would be valuable to a broad community. The paper compiles a substantial amount of information in compact form, and the proposed taxonomy of evaluation tasks is a useful organizing device. However, because the survey's main deliverable is the structured map in Table 2, factual errors in the availability column—at least Phikon, Phikon-v2, and RudolfV are incorrectly marked as unavailable—directly undermine its reliability. The absence of a stated methodology also weakens the comprehensiveness claim. With corrections and a methodology section, the paper could serve as a useful reference.

major comments (3)
  1. [Table 2] The Availability column lists Phikon (Filiot et al., 2023), Phikon-v2 (Filiot et al., 2024), and RudolfV (Dippel et al., 2024) as unavailable, but all three models have public weights: Phikon and Phikon-v2 are distributed on HuggingFace, and RudolfV is released under a public model license. These are not peripheral entries; Phikon and Phikon-v2 are widely used public baselines. Because Table 2 is the paper's central structured comparison, these errors mislead practitioners and must be corrected, and every other row should be re-verified against primary sources.
  2. [Sections 4–5 and Introduction] The paper claims to provide a 'comprehensive' review of 28 existing and up-to-date models and to summarize evaluation tasks 'for the first time,' but it gives no inclusion/exclusion criteria, search strategy, or date cutoff. This makes the completeness claim unauditable: a reader cannot determine why these 28 models were chosen or whether the six evaluation categories in Figure 3 are exhaustive. A methodology subsection describing how models and tasks were selected, categorized, and cross-checked is needed.
  3. [Section 5 and Figure 3] The evaluation taxonomy mixes task type, granularity, and experimental setting in a single hierarchy. For example, 'OOD generalization' and 'imbalanced' are listed as subcategories under classification alongside tile-level and WSI-level granularity, while supervised/zero-shot/few-shot settings appear as parallel dimensions. This conflation makes the claimed 'six main perspectives' less well-defined and the model assignments harder to interpret. Clarify the dimensions of the taxonomy or present them as separate axes.
minor comments (5)
  1. [Table 2] The header contains 'A vailability' with an erroneous space; it should read 'Availability'.
  2. [Section 5] The sentence 'In addition to qualitative analysis, some CPathFMs have undergone qualitative analysis' appears to use 'qualitative' twice; one instance should likely be 'quantitative'.
  3. [Table 1] The Phikon-v2 row lists 'Proprietary×4' as a data source, which is unclear; specify the four proprietary sources or explain the notation in a footnote.
  4. [Figure 3] The figure is densely packed; consider increasing font size and using labels or patterns in addition to color to distinguish uni-modal and multi-modal models for accessibility.
  5. [References] Reference [Zhou et al., 2024a] contains the malformed author string 'Lifeng others Wang'; this should be 'Lifeng Wang et al.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a literature survey whose taxonomies and comparisons are editorial classifications, not derivations from fitted inputs.

full rationale

The paper's central claim is descriptive: 'This survey provides a comprehensive review of CPathFMs, with a focus on datasets, adaptation strategies, and evaluation tasks.' Its main outputs are Tables 1–2 and Figure 3, which organize external models, datasets, and evaluation tasks. There are no equations, no fitted parameters, and no prediction derived from a fitted input. The six-category evaluation taxonomy is an editorial classification, and a classification scheme is not a circular derivation because nothing is defined in terms of the survey's own output. The alleged factual errors in Table 2's availability column, if real, are accuracy/correctness problems about external artifacts, not circularity: the entries are empirical claims about model releases that can be checked against the cited sources. The absence of explicit inclusion criteria is a method-transparency limitation, not a self-referential or forced derivation. No self-citations by the present authors appear in the reference list, and no load-bearing argument relies on the authors' prior work. The survey is self-contained as a review and does not reduce to its own inputs, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities appear in this survey. The load-bearing commitments are the selection of 28 models and the six-category evaluation taxonomy, both of which are editorial choices rather than evidence-based derivations.

assumptions (2)
  • ad hoc to paper The 28 surveyed models are a representative and complete set of relevant computational pathology foundation models.
    Section 4 and Table 2 list the models, but the paper gives no inclusion or exclusion criteria, search strategy, or cutoff date.
  • ad hoc to paper The six evaluation task categories (classification, retrieval, generation, segmentation, prediction, VQA) partition all meaningful evaluation tasks.
    Section 5 and Figure 3 present the taxonomy as a given without deriving it from prior evaluation benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Computational Pathology Foundation Models: Datasets, Adaptation Strategies, and Evaluation Tasks." pith.science (2026). https://pith.science/paper/7UVT3THU

@misc{pith2026250115724,
  author       = {Pith},
  title        = {Pith review of: A Survey on Computational Pathology Foundation Models: Datasets, Adaptation Strategies, and Evaluation Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7UVT3THU}},
  note         = {Machine review of arXiv:2501.15724}
}
read the original abstract

Computational pathology foundation models (CPathFMs) have emerged as a powerful approach for analyzing histopathological data, leveraging self-supervised learning to extract robust feature representations from unlabeled whole-slide images. These models, categorized into uni-modal and multi-modal frameworks, have demonstrated promise in automating complex pathology tasks such as segmentation, classification, and biomarker discovery. However, the development of CPathFMs presents significant challenges, such as limited data accessibility, high variability across datasets, the necessity for domain-specific adaptation, and the lack of standardized evaluation benchmarks. This survey provides a comprehensive review of CPathFMs in computational pathology, focusing on datasets, adaptation strategies, and evaluation tasks. We analyze key techniques, such as contrastive learning and multi-modal integration, and highlight existing gaps in current research. Finally, we explore future directions from four perspectives for advancing CPathFMs. This survey serves as a valuable resource for researchers, clinicians, and AI practitioners, guiding the advancement of CPathFMs toward robust and clinically applicable AI-driven pathology solutions.

Figures

Figures reproduced from arXiv: 2501.15724 by the authors.

Figure 1
Figure 1. An illustrative example of data modalities and challenges in CPath. The figure illustrates different histopathology data types, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of the pre-training pipeline for CPathFMs. The process involves data curation, including image curation, text curation, [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy of evaluation tasks for pre-trained CPathFMs. Uni-modal and multi-modal CPathFMs are highlighted in [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer

    cs.CV 2026-08 conditional novelty 6.0 of 10

    SmartStu distills multiple teacher pathology models into compact breast-cancer encoders with an adversarial noise model and self-supervision, matching or improving external-cohort accuracy at over 30x smaller size.

  2. HISTAI: An Open-Source, Large-Scale Whole Slide Image Dataset for Computational Pathology

    eess.IV 2025-05 conditional novelty 6.0 of 10

    HISTAI is a new open-access collection of 60,110 whole-slide images across eight pathology subsets, each case linked to clinical notes, demographics, and ICD-10 codes.

  3. PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

    cs.CV 2025-12 conditional novelty 5.0 of 10

    Splitting slide captions into random sentence subcaptions and aligning them with region features via text-conditioned attention improves whole-slide classification, retrieval, captioning and VQA in computational pathology.

  4. Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

    cs.CV 2025-11 reject novelty 3.0 of 10

    JWTH achieves modest tissue-classification gains by adding attention pooling and stain augmentation to a DINOv3 backbone, but the biomarker claims in the abstract are unsupported by the experiments.

Reference graph

Works this paper leans on

48 extracted references · 23 canonical work pages · cited by 4 Pith papers

  1. [1]

    To- wards large-scale training of pathology foundation mod- els

    [Aben et al., 2024] Nanne Aben, Edwin D de Jong, Ioannis Gatopoulos, Nicolas K¨anzig, Mikhail Karasikov, et al. To- wards large-scale training of pathology foundation mod- els. arXiv:2404.15217,

  2. [6]

    Computational pathology at health system scale–self-supervised foundation models from billions of images

    [Campanella et al., 2024b] Gabriele Campanella, Chad Van- derbilt, and Thomas Fuchs. Computational pathology at health system scale–self-supervised foundation models from billions of images. In AAAI 2024 Spring Symposium on Clinical F oundation Models,

  3. [7]

    Emerging properties in self-supervised vi- sion transformers

    [Caron et al., 2021] Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vi- sion transformers. In ICCV,

  4. [8]

    A New Era in Computational Pathology: A Survey on Foundation and Vision-Language Models

    [Chanda et al., 2024] Dibaloke Chanda, Milan Aryal, Nasim Yahya Soltani, and Masoud Ganji. A new era in computational pathology: A survey on foundation and vision-language models. arXiv:2408.14496,

  5. [9]

    A simple framework for contrastive learning of visual representations

    [Chen et al., 2020] Ting Chen, Simon Kornblith, Moham- mad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. InICML. PMLR,

  6. [11]

    Towards a general-purpose foundation model for computational pathology

    [Chen et al., 2024] Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology. Nature Medicine ,

  7. [12]

    The genotype-tissue expression (gtex) pilot analysis: multitis- sue gene regulation in humans

    [Consortium et al., 2015] GTEx Consortium, Kristin G Ardlie, David S Deluca, Ayellet V Segr `e, Timothy J Sul- livan, Taylor R Young, Ellen T Gelfand, Casandra A Trowbridge, Julian B Maller, Taru Tukiainen, et al. The genotype-tissue expression (gtex) pilot analysis: multitis- sue gene regulation in humans. Science,

  8. [14]

    Multimodal whole slide foundation model for pathology

    [Ding et al., 2024] Tong Ding, Sophia J Wagner, Andrew H Song, Richard J Chen, Ming Y Lu, Andrew Zhang, Anurag J Vaidya, Guillaume Jaume, Muhammad Shaban, Ahrong Kim, et al. Multimodal whole slide foundation model for pathology. arXiv:2411.19666,

Show all 48 references
  1. [15]

    Rudolfv: a foundation model by pathologists for pathologists.arXiv:2401.04079,

    [Dippel et al., 2024] Jonas Dippel, Barbara Feulner, Tobias Winterhoff, Timo Milbich, et al. Rudolfv: a foundation model by pathologists for pathologists.arXiv:2401.04079,

  2. [16]

    Scaling self-supervised learning for histopathology with masked image modeling

    [Filiot et al., 2023] Alexandre Filiot, Ridouane Ghermi, An- toine Olivier, Paul Jacob, Lucas Fidon, Axel Camara, Al- ice Mac Kain, Charlie Saillard, and Jean-Baptiste Schi- ratti. Scaling self-supervised learning for histopathology with masked image modeling. medRxiv,

  3. [17]

    Phikon-v2, a large and public feature extractor for biomarker prediction

    [Filiot et al., 2024] Alexandre Filiot, Paul Jacob, Alice Mac Kain, and Charlie Saillard. Phikon-v2, a large and public feature extractor for biomarker prediction. arXiv:2409.09173,

  4. [18]

    Masked au- toencoders are scalable vision learners

    [He et al., 2022] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked au- toencoders are scalable vision learners. In CVPR,

  5. [20]

    Quilt-1m: One million image-text pairs for histopathology

    [Ikezogwo et al., 2024] Wisdom Ikezogwo, Saygin Sey- fioglu, Fatemeh Ghezloo, Dylan Geva, Fatwir Sheikh Mo- hammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro. Quilt-1m: One million image-text pairs for histopathology. NeurIPS,

  6. [21]

    Perceiver: General perception with iterative atten- tion

    [Jaegle et al., 2021] Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Car- reira. Perceiver: General perception with iterative atten- tion. In ICML. PMLR,

  7. [22]

    Transcriptomics-guided slide representation learning in computational pathology

    [Jaume et al., 2024] Guillaume Jaume, Lukas Oldenburg, Anurag Vaidya, Richard J Chen, Drew FK Williamson, Thomas Peeters, Andrew H Song, and Faisal Mahmood. Transcriptomics-guided slide representation learning in computational pathology. In CVPR,

  8. [23]

    Pluto: Pathology-universal transformer

    [Juyal et al., 2024] Dinkar Juyal, Harshith Padigela, Chintan Shah, Daniel Shenker, et al. Pluto: Pathology-universal transformer. arXiv:2405.07905,

  9. [24]

    Benchmarking self-supervised learning on diverse pathology datasets

    [Kang et al., 2023] Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, and S ´ergio Pereira. Benchmarking self-supervised learning on diverse pathology datasets. In CVPR,

  10. [25]

    Benchmarking pathology foun- dation models: Adaptation strategies and scenarios

    [Lee et al., 2024] Jeaung Lee, Jeewoo Lim, Keunho Byeon, and Jin Tae Kwak. Benchmarking pathology foun- dation models: Adaptation strategies and scenarios. arXiv:2410.16038,

  11. [26]

    A visual-language foundation model for computa- tional pathology

    [Lu et al., 2024] Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guil- laume Jaume, Igor Odintsov, Long Phi Le, Georg Gerber, et al. A visual-language foundation model for computa- tional pathology. Nature Medicine,

  12. [27]

    Towards a generalizable pathol- ogy foundation model via unified knowledge distillation

    [Ma et al., 2024] Jiabo Ma, Zhengrui Guo, Fengtao Zhou, Yihui Wang, et al. Towards a generalizable pathol- ogy foundation model via unified knowledge distillation. arXiv:2407.18449,

  13. [28]

    Hibou: A family of foundational vision transformers for pathology

    [Nechaev et al., 2024] Dmitry Nechaev, Alexey Pchelnikov, and Ekaterina Ivanova. Hibou: A family of foundational vision transformers for pathology. arXiv:2406.05074,

  14. [29]

    Benchmarking foundation models as feature extractors for weakly-supervised computational pathology

    [Neidlinger et al., 2024] Peter Neidlinger, Omar SM El Nah- has, Hannah Sophie Muti, Tim Lenz, Michael Hoffmeister, Hermann Brenner, et al. Benchmarking foundation models as feature extractors for weakly-supervised computational pathology. arXiv:2408.15823,

  15. [30]

    Pathology foundation models

    [Ochi et al., 2024] Mieko Ochi, Daisuke Komura, and Shumpei Ishikawa. Pathology foundation models. arXiv:2407.21317,

  16. [31]

    Dinov2: Learning robust visual features without supervision

    [Oquab et al., 2023] Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khali- dov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv:2304.07193,

  17. [32]

    Beit v2: Masked im- age modeling with vector-quantized visual tokenizers

    [Peng et al., 2022] Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked im- age modeling with vector-quantized visual tokenizers. arXiv:2208.06366,

  18. [33]

    Learning transferable visual models from nat- ural language supervision

    [Radford et al., 2021] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from nat- ural language supervision. In ICML. PMLR,

  19. [34]

    H-optimus-0,

    [Saillard et al., 2024] Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-L ´opez, Zelda Mariet, David Cahan ´e, Eric Durand, and Jean-Philippe Vert. H-optimus-0,

  20. [35]

    Prism: A multi-modal generative foundation model for slide-level histopathology

    [Shaikovski et al., 2024] George Shaikovski, Adam Casson, Kristen Severson, Eric Zimmermann, Yi Kan Wang, Jeremy D Kunz, Juan A Retamero, et al. Prism: A multi-modal generative foundation model for slide-level histopathology. arXiv:2405.10254,

  21. [36]

    Pathasst: A generative foun- dation ai assistant towards artificial general intelligence of pathology

    [Sun et al., 2024] Yuxuan Sun, Chenglu Zhu, Sunyi Zheng, Kai Zhang, Lin Sun, Zhongyi Shui, Yunlong Zhang, Honglin Li, and Lin Yang. Pathasst: A generative foun- dation ai assistant towards artificial general intelligence of pathology. In AAAI,

  22. [37]

    Virchow: A million-slide digital pathology founda- tion model

    [V orontsovet al., 2023] Eugene V orontsov, Alican Bozkurt, Adam Casson, George Shaikovski, Michal Zelechowski, et al. Virchow: A million-slide digital pathology founda- tion model. arXiv:2309.07778,

  23. [38]

    Image as a foreign lan- guage: Beit pretraining for all vision and vision-language tasks

    [Wang et al., 2022a] Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, et al. Image as a foreign lan- guage: Beit pretraining for all vision and vision-language tasks. arXiv:2208.10442,

  24. [39]

    A pathol- ogy foundation model for cancer diagnosis and prognosis prediction

    [Wang et al., 2024] Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, Jietian Jin, Jiayu Zhang, Ruijiang Li, Hongping Tang, Kanran Wang, Yu Li, et al. A pathol- ogy foundation model for cancer diagnosis and prognosis prediction. Nature,

  25. [40]

    The cancer genome atlas pan-cancer analysis project

    [Weinstein et al., 2013] John N Weinstein, Eric A Collisson, Gordon B Mills, Kenna R Shaw, Brad A Ozenberger, Kyle Ellrott, Ilya Shmulevich, Chris Sander, and Joshua M Stu- art. The cancer genome atlas pan-cancer analysis project. Nature genetics,

  26. [42]

    A whole-slide foundation model for digital pathology from real-world data

    [Xu et al., 2024] Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier Gonz ´alez, Yu Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature,

  27. [43]

    A foundation model for gen- eralizable cancer diagnosis and survival prediction from histopathological images

    [Yang et al., 2024] Zhaochang Yang, Ting Wei, Ying Liang, Xin Yuan, Ruitian Gao, et al. A foundation model for gen- eralizable cancer diagnosis and survival prediction from histopathological images. bioRxiv,

  28. [44]

    Coca: Contrastive captioners are image-text foundation models

    [Yu et al., 2022] Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv:2205.01917,

  29. [45]

    Biomed- clip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

    [Zhang et al., 2023] Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, et al. Biomed- clip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv:2303.00915,

  30. [46]

    ibot: Image bert pre-training with online tokenizer

    [Zhou et al., 2021] Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. arXiv:2111.07832,

  31. [47]

    A knowledge-enhanced pathology vision-language founda- tion model for cancer diagnosis

    [Zhou et al., 2024a] Xiao Zhou, Luoyi Sun, Dexuan He, Wenbin Guan, Ruifen Wang, and Lifeng others Wang. A knowledge-enhanced pathology vision-language founda- tion model for cancer diagnosis. arXiv:2412.13126,

  32. [48]

    Virchow2: Scaling self-supervised mixed magnification models in pathology

    [Zimmermann et al., 2024] Eric Zimmermann, Eugene V orontsov, Julian Viret, Adam Casson, Michal Zele- chowski, et al. Virchow2: Scaling self-supervised mixed magnification models in pathology. arXiv:2408.00738, 2024

  33. [2013]

    A vision–language foundation model for pre- cision oncology

    [Xiang et al., 2025] Jinxi Xiang, Xiyue Wang, Xiaoming Zhang, Yinghua Xi, Feyisope Eweje, Yijiang Chen, Yuchen Li, Colin Bergstrom, Matthew Gopaulchan, Ted Kim, et al. A vision–language foundation model for pre- cision oncology. Nature,

  34. [2015]

    Longnet: Scaling transformers to 1,000,000,000 tokens

    [Ding et al., 2023] Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, Nanning Zheng, and Furu Wei. Longnet: Scaling transformers to 1,000,000,000 tokens. arXiv:2307.02486,

  35. [2020]

    An empirical study of training self-supervised vision transformers

    [Chen et al., 2021] Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV,

  36. [2021]

    A clinical benchmark of public self-supervised pathology foundation models

    [Campanella et al., 2024a] Gabriele Campanella, Shengjia Chen, Ruchika Verma, Jennifer Zeng, et al. A clinical benchmark of public self-supervised pathology foundation models. arXiv:2407.06508,

  37. [2022]

    A visual–language foundation model for pathology image analysis using medical twitter

    [Huang et al., 2023] Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual–language foundation model for pathology image analysis using medical twitter. Nature medicine,

  38. [2023]

    Beit: Bert pre-training of image transformers

    [Bao et al., 2021] Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv:2106.08254,

  39. [2024]

    A novel pathology foundation model by mayo clinic, charit\’e, and aignostics

    [Alber et al., 2025] Maximilian Alber, Stephan Tietz, Jonas Dippel, Timo Milbich, Timoth ´ee Lesort, Panos Korfiatis, et al. A novel pathology foundation model by mayo clinic, charit\’e, and aignostics. arXiv:2501.05409,

  40. [2025]

    Robust and data-efficient gener- alization of self-supervised machine learning for diagnos- tic imaging

    [Azizi et al., 2023] Shekoofeh Azizi, Laura Culp, Jan Frey- berg, Basil Mustafa, et al. Robust and data-efficient gener- alization of self-supervised machine learning for diagnos- tic imaging. Nature Biomedical Engineering,

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.