Pith. sign in

REVIEW 3 major objections 5 minor 3 cited by

This paper claims one foundation-model family can lead in prediction accuracy, robustness, and compute efficiency for clinical pathology.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 11:45 UTC pith:LVBXUNN3

load-bearing objection Atlas 2 is a large-scale, credible empirical advance in pathology foundation models, but the missing train/test overlap audit and lack of released weights make the SOTA margins conditional. the 3 major comments →

arxiv 2601.05148 v2 pith:LVBXUNN3 submitted 2026-01-08 cs.CV cs.AIcs.LG

Atlas 2 -- Foundation models for clinical deployment

classification cs.CV cs.AIcs.LG
keywords pathology foundation modelwhole-slide imagesself-supervised learningVision Transformerknowledge distillationrobustnessclinical deploymentbenchmark evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The report claims that a single pathology vision model family can overcome the usual tradeoffs between predictive performance, robustness to hospital-specific variation, and computational cost. Atlas 2, a 2-billion-parameter ViT trained on the largest pathology pretraining set assembled so far (5.5 million whole-slide images from three medical centers, sampled at several magnifications), is reported to be the best or second-best model in a comparison of fifteen competitors across eighty public benchmarks. The distilled versions, Atlas 2-B and Atlas 2-S, are reported to keep near-competitive accuracy while running 3.4x and 9x faster than the teacher. If these results hold, clinical AI systems could deploy one family of features instead of selecting between accuracy, robustness, and speed.

Core claim

The central claim is that Atlas 2 achieves state-of-the-art performance and robustness simultaneously: on the benchmark subset used for head-to-head comparison, it is best in 22 of 27 tasks and second-best in 3, with a 44.8% average on the HEST gene-expression benchmark, 82.9% on the eva morphology benchmark, and an 85.7% average robustness score—9.7 percentage points above the closest competitor (Virchow2). The authors further claim that the distilled variants Atlas 2-B (86M parameters) and Atlas 2-S (22M parameters) are the best in their compute class, with Atlas 2-B matching or exceeding several larger models and Atlas 2-S approaching Virchow2-class accuracy at 4.3–7.4x greater inference

What carries the argument

Atlas 2 is a 2-billion-parameter Vision Transformer with patch size 8, trained with a self-supervised recipe built on DINOv2/DINOv3 principles and the authors' earlier pathology-model pipeline, over 5.5 million whole-slide images at four resolutions (0.25–2.0 µm/pixel). The key mechanism for efficiency is knowledge distillation from Atlas 2 into ViT-B and ViT-S encoders, and for evaluation the models act as frozen feature extractors whose CLS+mean-patch embeddings feed lightweight heads (linear probes, ABMIL, PCA-ridge regression) in five public benchmark suites plus internal linear-probing tasks.

Load-bearing premise

The load-bearing premise is that the 5.5-million-slide pretraining corpus (from three medical institutions) does not overlap with the public benchmark evaluation data; if any patient or tile leakage exists, the reported state-of-the-art margins, especially on TCGA-based tasks, would be inflated.

What would settle it

Cross-reference slide identifiers, patient metadata, and tile coordinates between the pretraining archives and the public benchmark datasets (e.g., TCGA, CAMELYON, PLISM); finding even a small fraction of shared tiles or patients would invalidate the claimed margins, while confirming zero overlap would support them.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the claim holds, pathology departments can use one model family for accurate, robust features without needing separate models for accuracy and for scanner/stain invariance.
  • Atlas 2-B and 2-S make foundation-model-level features practical on standard hospital GPUs, since they run 3.4x and 9x faster than Atlas 2 while staying close in accuracy.
  • The reported robustness margin of 9.7 p.p. over the closest contender implies that pretraining on very large multi-center archives reduces sensitivity to technical variation better than previously seen.
  • Consistent top results across gene-expression, morphology, and molecular tasks suggest the same frozen embeddings can serve both research and clinical prediction workloads.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A controlled ablation separating corpus scale from multi-magnification sampling would clarify which ingredient drives the robustness gain; the paper does not isolate these factors.
  • The comparison against Pluto-4G and H-Optimus-1 relies on published numbers rather than a shared evaluation harness, so some margins could shift if preprocessing or splits differed; re-running those models under the same protocol would be a direct test.
  • If the no-leakage assumption holds, the results suggest diminishing returns from simply scaling pretraining data; a smaller corpus at the same recipe might produce similar gains, which would lower the barrier for future independent reproduction.
  • A natural next validation would be evaluating Atlas 2 on slides from an institution not represented in the pretraining corpus, which would test whether the robustness advantage transfers beyond the three training centers.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Atlas 2, a 2-billion-parameter ViT-based pathology foundation model trained on 5.5 million whole-slide images from Charité, LMU, and Mayo Clinic, together with two distilled lightweight variants, Atlas 2-B (ViT-B/8) and Atlas 2-S (ViT-S/8). The authors evaluate frozen embeddings across eighty public benchmarks organized into five frameworks — HEST, eva, PathoROB, Plismbench, and Patho-Bench — plus additional MSI and TCGA Uniform tasks. They report that Atlas 2 achieves state-of-the-art performance on most tasks and, in particular, strong robustness to center/staining shifts, while the distilled models are competitive with much larger models at 3–9x higher inference throughput. The central claim is that Atlas 2 simultaneously improves prediction performance, robustness, and resource efficiency relative to existing pathology foundation models.

Significance. If the claims are correct, Atlas 2 would be a notable step toward clinically deployable pathology foundation models: it combines large-scale pretraining, multi-magnification training, and distilled efficient variants, with strong results across diverse public benchmarks. The use of pinned public evaluation framework versions (HEST v1.2.0, eva v0.4.2, PathoROB, Plismbench, Patho-Bench) and the breadth of comparison models are strengths that make the results independently checkable. However, the absence of any stated overlap audit between the 5.5M-WSI pretraining corpus and the evaluation data is a serious unaddressed risk for the headline SOTA claim, as is the lack of uncertainty quantification for most comparisons. The paper's main value — a large, multi-centric pathology foundation model with efficient distilled versions — is real, but the current evidence does not yet fully secure the claimed margins.

major comments (3)
  1. [§3.1, §3.3, Tables 1–3]
  2. [§4.1, Tables 1–2]
  3. [§4.1, Table 1]
minor comments (5)
  1. [§1, Abstract] Typo: 'Universtätsmedizin' should be 'Universitätsmedizin' in the abstract and Section 1. Also 'eva 2' in Section 4.1 should probably be 'eva' (or a clearer label).
  2. [Table 2 caption] Typo: 'chancel level' should be 'chance level'.
  3. [§4.1, Figure 1] The processing speed of Pluto-4G is approximated from 'the fastest speed of same sized models'; this approximation should be stated more prominently in the figure and in any speed-efficiency claim, since it directly affects the Pareto-front visual in Figure 1(C).
  4. [Table 1] The labels 'Morphology-Average-Pluto-4G' and 'Morphology-Average' are confusing, especially because Pluto-4G has missing entries in many rows. Clarify which tasks are included in each average and why the Pluto-4G average differs.
  5. [§3.3, PathoROB] The statement that PathoROB averages for Virchow2 and Prov-GigaPath differ from official values due to bicubic resizing should be accompanied by a note on whether this resizing change affects all models equally or could alter relative rankings in other frameworks.

Circularity Check

0 steps flagged

No circularity: Atlas 2's SOTA claims are empirical benchmark results against external frameworks; no derivation reduces to its inputs.

full rationale

The paper's central claims are empirical evaluation outcomes, not analytic derivations or first-principles predictions. Atlas 2 is trained on unlabeled WSIs and then evaluated on public benchmark frameworks (HEST, eva, Patho-Bench, Plismbench, PathoROB) using frozen encoders and task-specific heads; the reported numbers are direct measurements rather than quantities derived from the training objective or from the evaluation benchmarks by construction. No parameter is fitted to a subset of benchmark data and then reported as a prediction of a closely related benchmark quantity. The robustness claim relies on PathoROB, whose authors overlap with the present paper, but PathoROB is a fixed public benchmark with externally defined datasets, code, and protocol; using it is not a self-citation chain that forces the result. The potential overlap between the Mayo Clinic pretraining corpus and TCGA-based evaluation sets is a data-contamination/validation risk, not a circularity in the derivation: it concerns whether the empirical comparison is fair, not whether the claim is equivalent to its inputs by definition. Therefore no self-definitional, fitted-input-called-prediction, or load-bearing self-citation circularity is present.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The paper contributes empirical benchmarks and a trained model; it does not introduce new theoretical entities. The main unstated inputs are the pretraining corpus composition, the training recipe, and the assumption that benchmark tasks generalize to clinical use.

axioms (5)
  • domain assumption Pretraining on a large, diverse multi-centric WSI corpus transfers to downstream clinical benchmarks.
    The entire evaluation depends on this; no theoretical or causal proof is given (Sections 3.1 and 4).
  • domain assumption Evaluation benchmarks are representative of clinical deployment and are free of data leakage from pretraining.
    The paper never states that the 5.5M pretraining WSIs are disjoint from TCGA/CAMELYON/HEST test data (Section 3.1 vs Section 3.3).
  • ad hoc to paper The 53-task Patho-Bench subset is representative of the full 95-task benchmark.
    Selection is described as determined by implementation difficulties and dataset access (Section 3.3, Patho-Bench paragraph), not by a predefined protocol.
  • domain assumption Competing models' available weights or published results are evaluated under equivalent conditions.
    Table 1 footnotes: Pluto-4G results are taken from [60] and H0-mini from [22]/[62]; PathoROB uses bicubic resizing, which changes exact values for some models (Section 3.3).
  • domain assumption The DINOv2/DINOv3/RudolfV/Atlas training recipes are valid for histopathology representation learning.
    Section 3.2 states adapted RudolfV/Atlas and DINOv2/DINOv3 components, but no derivation or proof is provided; this is accepted background in self-supervised learning.

pith-pipeline@v1.3.0-alltime-deepseek · 31884 in / 15562 out tokens · 164040 ms · 2026-08-03T11:45:50.867102+00:00 · methodology

0 comments
read the original abstract

Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this report, we present Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models which bridge these shortcomings by showing state-of-the-art prediction performance, robustness, and resource efficiency in a comprehensive evaluation across eighty public benchmarks. Our models were trained on the largest pathology foundation model dataset to date comprising 5.5 million histopathology whole slide images, collected from three medical institutions Charit\'e - Universit\"atsmedizin Berlin, LMU Munich, and Mayo Clinic.

Figures

Figures reproduced from arXiv: 2601.05148 by Alessandro Benetti, Alexandra Carpen-Amarie, Andrew Norgan, Aniruddh Jammoria, Beatriz Perez Cancer, David Horst, Elias Eulig, Frederick Klauschen, Gabriel Dernbach, Jake Matras, J\'er\^ome L\"uscher, Jiasen Wu, Jonas Dippel, Klaus-Robert M\"uller, Lukas Muttenthaler, Lukas Ruff, Matt Redlon, Maximilian Alber, Moritz Kr\"ugener, Neelay Shah, Panos Korfiatis, Patrick Duffy, Philipp Jurmeister, Sayed Abid Hashimi, Simon Schallenberg, Stephan Tietz, Timo Milbich.

Figure 1
Figure 1. Figure 1: (A) The results show that Atlas 2 is the best performing model. Additionally Atlas 2-B and 2-S have similar prediction performance as contenders but are up to a magnitude more efficient. The performance average is based on morphology prediction tasks from eva [27] and molecular prediction tasks from HEST [34] in [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (A) - (E) show the average performance per evaluation framework [ [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy

    cs.CV 2026-06 conditional novelty 6.0

    An Atlas-foundation-model system for multi-cancer H&E tissue and cell profiling matches pathologist H&E accuracy against IHC-informed consensus and generalizes across 1,500+ cases.

  2. Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy

    cs.CV 2026-06 unverdicted novelty 5.0

    Atlas H&E-TME is a new AI system for cell-level tissue profiling on H&E slides that matches pathologist performance when validated against an IHC-informed consensus and a large multi-cancer H&E annotation set.

  3. OpenTME: An Open Dataset of AI-powered H&E Tumor Microenvironment Profiles from TCGA

    cs.CV 2026-04 unverdicted novelty 4.0

    OpenTME provides pre-computed TME profiles with over 4,500 quantitative readouts per slide from 3,634 TCGA H&E images using an AI pipeline based on pathology foundation models.

Reference graph

Works this paper leans on

83 extracted references · 11 linked inside Pith · cited by 2 Pith papers

  1. [1]

    de Jong, Ioannis Gatopoulos, Nicolas Känzig, Mikhail Karasikov, Axel Lagré, Roman Moser, Joost van Doorn, and Fei Tang

    Nanne Aben, Edwin D. de Jong, Ioannis Gatopoulos, Nicolas Känzig, Mikhail Karasikov, Axel Lagré, Roman Moser, Joost van Doorn, and Fei Tang. Towards Large-Scale Training of Pathology Foundation Models, March 2024. arXiv:2404.15217

  2. [2]

    Atlas: A novel pathology foundation model by mayo clinic, charité, and aignostics, 2025

    Maximilian Alber, Stephan Tietz, Jonas Dippel, Timo Milbich, Timothée Lesort, Panos Korfiatis, Moritz Krügener, Beatriz Perez Cancer, Neelay Shah, Alexander Möllers, Philipp Seegerer, Alexandra Carpen-Amarie, Kai Standvoss, Gabriel Dernbach, Edwin de Jong, Simon Schal- lenberg, Andreas Kunft, Helmut Hoffer von Ankershoffen, Gavin Schaeferle, Patrick Duffy...

  3. [3]

    BACH: Grand challenge on breast cancer histology images.Medical image analysis, 56:122– 139, 2019

    Guilherme Aresta, Teresa Araújo, Scotty Kwok, Sai Saketh Chennamsetty, Mohammed Safwan, Varghese Alex, Bahram Marami, Marcel Prastawa, Monica Chan, Michael Donovan, et al. BACH: Grand challenge on breast cancer histology images.Medical image analysis, 56:122– 139, 2019. 9

  4. [4]

    Automated gleason grading of prostate cancer tissue microarrays via deep learning, 03 2018

    Eirini Arvaniti, Kim Fricker, Michael Moret, Niels Rupp, Thomas Hermanns, Christian Fankhauser, Norbert Wey, Peter Wild, Jan Rueschoff, and Manfred Claassen. Automated gleason grading of prostate cancer tissue microarrays via deep learning, 03 2018

  5. [5]

    Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer.JAMA, 318(22):2199–2210, 2017

    Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer.JAMA, 318(22):2199–2210, 2017

  6. [6]

    Knowledge distillation: A good teacher is patient and consistent

    Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. Knowledge distillation: A good teacher is patient and consistent. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  7. [7]

    Morphological and molecular breast cancer profiling through explainable machine learning

    Alexander Binder, Michael Bockmayr, Miriam Hägele, Stephan Wienert, Daniel Heim, Katha- rina Hellweg, Masaru Ishii, Albrecht Stenzinger, Andreas Hocke, Carsten Denkert, et al. Morphological and molecular breast cancer profiling through explainable machine learning. Nature Machine Intelligence, 3(4):355–366, 2021

  8. [8]

    H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025

    Bioptimus. H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025

  9. [9]

    Artificial intelligence for diagnosis and gleason grading of prostate cancer: the PANDA challenge.Nature medicine, 28(1):154–163, 2022

    Wouter Bulten, Kimmo Kartasalo, Po-Hsuan Cameron Chen, Peter Ström, Hans Pinckaers, Kunal Nagpal, Yuannan Cai, David F Steiner, Hester Van Boven, Robert Vink, et al. Artificial intelligence for diagnosis and gleason grading of prostate cancer: the PANDA challenge.Nature medicine, 28(1):154–163, 2022

  10. [10]

    Péter Bándi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, Quanzheng Li, Farhad Ghazvinian Zanjani, Svitlana Zinger, Keisuke Fukuta, Daisuke Komura, Vlado Ovtcharov, Shenghua Cheng, Shaoqun Zeng, Jeppe Thagaard, Anders B. Dahl, Huangjing Lin, Hao...

  11. [11]

    Pathology foundation models are scanner sensitive: Benchmark and mitigation with contrastive scangen loss

    Gianluca Carloni, Biagio Brattoli, Seongho Keum, Jongchan Park, Taebum Lee, Chang Ho Ahn, and Sergio Pereira. Pathology foundation models are scanner sensitive: Benchmark and mitigation with contrastive scangen loss. InInternational Workshop on Foundation Models for General Medical AI, pages 44–53. Springer, 2025

  12. [12]

    Impact of tissue staining and scanner variation on the performance of pathology foundation models: a study of sarcomas and their mimics.bioRxiv, pages 2025–08, 2025

    Binghao Chai, Jianan Chen, Paul Cool, Fatine Oumlil, Anna Tollitt, David F Steiner, Tapabrata Chakraborti, and Adrienne M Flanagan. Impact of tissue staining and scanner variation on the performance of pathology foundation models: a study of sarcomas and their mimics.bioRxiv, pages 2025–08, 2025

  13. [13]

    Chen, Tong Ding, Ming Y

    Richard J. Chen, Tong Ding, Ming Y . Lu, Drew F. K. Williamson, Guillaume Jaume, Andrew H. Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, Mane Williams, Lukas Oldenburg, Luca L. Weishaupt, Judy J. Wang, Anurag Vaidya, Long Phi Le, Georg Gerber, Sharifa Sahai, Walt Williams, and Faisal Mahmood. Towards a general-purpose foundation model for ...

  14. [14]

    Current pathology foundation models are unrobust to medical center differences.arXiv preprint arXiv:2501.18055, 2025

    Edwin D de Jong, Eric Marcus, and Jonas Teuwen. Current pathology foundation models are unrobust to medical center differences.arXiv preprint arXiv:2501.18055, 2025

  15. [15]

    Human-interpretable image features derived from densely mapped cancer pathology slides predict diverse molecular phenotypes.Nature communications, 12(1):1613, 2021

    James A Diao, Jason K Wang, Wan Fung Chui, Victoria Mountain, Sai Chowdary Gullapally, Ramprakash Srinivasan, Richard N Mitchell, Benjamin Glass, Sara Hoffman, Sudha K Rao, et al. Human-interpretable image features derived from densely mapped cancer pathology slides predict diverse molecular phenotypes.Nature communications, 12(1):1613, 2021

  16. [16]

    A multimodal whole-slide foundation model for pathology.Nature Medicine, pages 1–13, 2025

    Tong Ding, Sophia J Wagner, Andrew H Song, Richard J Chen, Ming Y Lu, Andrew Zhang, Anurag J Vaidya, Guillaume Jaume, Muhammad Shaban, Ahrong Kim, et al. A multimodal whole-slide foundation model for pathology.Nature Medicine, pages 1–13, 2025. 10

  17. [17]

    RudolfV: A Foundation Model by Pathologists for Pathologists, June 2024

    Jonas Dippel, Barbara Feulner, Tobias Winterhoff, Timo Milbich, Stephan Tietz, Simon Schal- lenberg, Gabriel Dernbach, Andreas Kunft, Simon Heinke, Marie-Lisa Eich, Julika Ribbat-Idel, Rosemarie Krupar, Philipp Anders, Niklas Prenißl, Philipp Jurmeister, David Horst, Lukas Ruff, Klaus-Robert Müller, Frederick Klauschen, and Maximilian Alber. RudolfV: A Fo...

  18. [18]

    Ai-based anomaly detection for clinical-grade histopathological diagnostics.NEJM AI, 1(11):AIoa2400468, 2024

    Jonas Dippel, Niklas Prenißl, Julius Hense, Philipp Liznerski, Tobias Winterhoff, Simon Schallenberg, Marius Kloft, Oliver Buchstab, David Horst, Maximilian Alber, Lukas Ruff, Klaus-Robert Müller, and Frederick Klauschen. Ai-based anomaly detection for clinical-grade histopathological diagnostics.NEJM AI, 1(11):AIoa2400468, 2024

  19. [19]

    An image is worth 16x16 words: Transformers for image recognition at scale.ICLR, 2021

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale.ICLR, 2021

  20. [20]

    David Jacob Drexlin, Jonas Dippel, Julius Hense, Niklas Prenißl, Grégoire Montavon, Frederick Klauschen, and Klaus-Robert Müller. Medi: Metadata-guided diffusion models for mitigating biases in tumor classification.International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 379–388, 2025

  21. [21]

    https://github.com/ kaiko-ai/eva/tree/0.4.2, 2025

    eva: Evaluation framework for oncology foundation models (FMs). https://github.com/ kaiko-ai/eva/tree/0.4.2, 2025. v0.4.2

  22. [22]

    Dis- tilling foundation models for robust and efficient models in digital pathology

    Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Rémy Dubois, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, and Antoine Olivier. Dis- tilling foundation models for robust and efficient models in digital pathology. In James C. Gee, Daniel C. Alexander, Jaesung Hong, Juan Eugenio Iglesias, Carole H. Sudre, Arch...

  23. [23]

    Distilling foundation models for robust and efficient models in digital pathology, 2025

    Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, and Antoine Olivier. Distilling foundation models for robust and efficient models in digital pathology, 2025

  24. [24]

    Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023

    Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023

  25. [25]

    Phikon-v2, A large and public feature extractor for biomarker prediction, September 2024

    Alexandre Filiot, Paul Jacob, Alice Mac Kain, and Charlie Saillard. Phikon-v2, A large and public feature extractor for biomarker prediction, September 2024. arXiv:2409.09173

  26. [26]

    Lempitsky

    Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor S. Lempitsky. Domain-adversarial training of neural networks.J. Mach. Learn. Res., 17:59:1–59:35, 2016

  27. [27]

    eva: Evaluation framework for pathology foundation models

    Ioannis Gatopoulos, Nicolas Känzig, Roman Moser, and Sebastian Otálora. eva: Evaluation framework for pathology foundation models. InMedical Imaging with Deep Learning, 2024

  28. [28]

    Simon Graham, Quoc Dang Vu, Shan E Ahmed Raza, Jin Tae Kwak, and Nasir M. Rajpoot. XY network for nuclear segmentation in multi-tissue histology images.CoRR, abs/1812.06499, 2018

  29. [29]

    Evaluating computational pathology foundation models for prostate cancer grading under distribution shifts.arXiv preprint arXiv:2410.06723, 2024

    Fredrik K Gustafsson and Mattias Rantalainen. Evaluating computational pathology foundation models for prostate cancer grading under distribution shifts.arXiv preprint arXiv:2410.06723, 2024

  30. [30]

    https://github

    Hest-library: Bringing spatial transcriptomics and histopathology together. https://github. com/mahmoodlab/HEST/tree/v1.2.0, 2024. v1.2.0

  31. [31]

    Distilling the knowledge in a neural network

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In NeurIPS 2014 - Deep Learning Workshop, 2015. 11

  32. [32]

    Attention-based deep multiple instance learning

    Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. InInternational Conference on Machine Learning, pages 2127–2136, 2018

  33. [33]

    Tomczak, and Max Welling

    Maximilian Ilse, Jakub M. Tomczak, and Max Welling. Attention-based deep multiple instance learning, 2018

  34. [34]

    Song, Ming Y

    Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y . Lu, Cristina Almagro Pérez, Sophia J Wagner, Anurag Jayant Vaidya, Richard J. Chen, Drew FK Williamson, Ahrong Kim, and Faisal Mahmood. HEST-1k: A dataset for spatial transcriptomics and histology image analysis. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Bench...

  35. [35]

    Hipp, Darren Fahy, Benjamin Glass, Eric Walk, John Abel, Harsha Vardhan pokkalla, Andrew H

    Dinkar Juyal, Harshith Padigela, Chintan Shah, Daniel Shenker, Natalia Harguindeguy, Yi Liu, Blake Martin, Yibo Zhang, Michael Nercessian, Miles Markey, Isaac Finberg, Kelsey Luu, Daniel Borders, Syed Ashar Javed, Emma L Krause, Raymond Biju, Aashish Sood, Allen Ma, Jackson Nyman, John Shamshoian, Guillaume Chhor, Darpan Sanghavi, Marc Thibault, Limin Yu,...

  36. [36]

    Champkit: A framework for rapid evaluation of deep neural networks for patch-based histopathology classification.Computer methods and programs in biomedicine, 239:107631, 2023

    Jakub R Kaczmarzyk, Rajarsi Gupta, Tahsin M Kurc, Shahira Abousamra, Joel H Saltz, and Peter K Koo. Champkit: A framework for rapid evaluation of deep neural networks for patch-based histopathology classification.Computer methods and programs in biomedicine, 239:107631, 2023

  37. [37]

    Benchmarking self-supervised learning on diverse pathology datasets

    Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, and Sérgio Pereira. Benchmarking self-supervised learning on diverse pathology datasets. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023

  38. [38]

    Daniel Kaplan, Ratna Sagari Grandhi, Connor Lane, Benjamin Warner, Tanishq Mathew Abraham, and Paul S. Scotti. How to train a state-of-the-art pathology foundation model with $1.6k, 2025

  39. [39]

    Training state-of-the-art pathology foundation models with orders of magnitude less data

    Mikhail Karasikov, Joost van Doorn, Nicolas Känzig, Melis Erdal Cesur, Hugo Mark Horlings, Robert Berke, Fei Tang, and Sebastian Otálora. Training state-of-the-art pathology foundation models with orders of magnitude less data. InMedical Image Computing and Computer Assisted Intervention – MICCAI 2025, volume LNCS 15967, pages 573–583. Springer Nature Swi...

  40. [40]

    100,000 histological images of human colorectal cancer and healthy tissue (v0.1) [Data set]

    Jakob Nikolas Kather, Niels Halama, and Alexander Marx. 100,000 histological images of human colorectal cancer and healthy tissue (v0.1) [Data set]. Zenodo. https://doi.org/10. 5281/zenodo.1214456, May 2018

  41. [41]

    Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study.PLoS medicine, 16(1):e1002730, 2019

    Jakob Nikolas Kather, Johannes Krisam, Pornpimol Charoentong, Tom Luedde, Esther Herpel, Cleo-Aron Weis, Timo Gaiser, Alexander Marx, Nektarios A Valous, Dyke Ferber, et al. Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study.PLoS medicine, 16(1):e1002730, 2019

  42. [42]

    Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer.Nature medicine, 25(7):1054–1056, 2019

    Jakob Nikolas Kather, Alexander T Pearson, Niels Halama, Dirk Jäger, Jeremias Krause, Sven H Loosen, Alexander Marx, Peter Boor, Frank Tacke, Ulf Peter Neumann, et al. Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer.Nature medicine, 25(7):1054–1056, 2019

  43. [43]

    Explainable AI reveals Clever Hans effects in unsupervised learning models.Nature Machine Intelligence, 7:412—-422, 2025

    Jacob Kauffmann, Jonas Dippel, Lukas Ruff, Wojciech Samek, Klaus-Robert Müller, and Grégoire Montavon. Explainable AI reveals Clever Hans effects in unsupervised learning models.Nature Machine Intelligence, 7:412—-422, 2025

  44. [44]

    Patient-level proteomic network prediction by ex- plainable artificial intelligence.NPJ Precision Oncology, 6(1):35, 2022

    Philipp Keyl, Michael Bockmayr, Daniel Heim, Gabriel Dernbach, Grégoire Montavon, Klaus- Robert Müller, and Frederick Klauschen. Patient-level proteomic network prediction by ex- plainable artificial intelligence.NPJ Precision Oncology, 6(1):35, 2022. 12

  45. [45]

    Toward explainable artificial intelligence for precision pathology.Annual Review of Pathology: Mechanisms of Disease, 19(1):541–570, 2024

    Frederick Klauschen, Jonas Dippel, Philipp Keyl, Philipp Jurmeister, Michael Bockmayr, Andreas Mock, Oliver Buchstab, Maximilian Alber, Lukas Ruff, Grégoire Montavon, et al. Toward explainable artificial intelligence for precision pathology.Annual Review of Pathology: Mechanisms of Disease, 19(1):541–570, 2024

  46. [46]

    Do histopathological foundation models eliminate batch effects? A comparative study

    Jonah Kömen, Hannah Marienwald, Jonas Dippel, and Julius Hense. Do histopathological foundation models eliminate batch effects? A comparative study. 2024. Presented at: NeurIPS 2024 Workshop on Advancements In Medical Foundation Models: Explainability, Robustness, Security, and Beyond

  47. [47]

    Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9), 2022

    Daisuke Komura, Akihiro Kawabe, Keisuke Fukuta, Kyohei Sano, Toshikazu Umezaki, Hi- rotomo Koda, Ryohei Suzuki, Ken Tominaga, Mieko Ochi, Hiroki Konishi, et al. Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9), 2022

  48. [48]

    de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller

    Jonah Kömen, Edwin D. de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller. Towards robust foundation models for digital pathology, 2025

  49. [49]

    Unmasking Clever Hans predictors and assessing what machines really learn.Nature Communications, 10(1):1096, 2019

    Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. Unmasking Clever Hans predictors and assessing what machines really learn.Nature Communications, 10(1):1096, 2019

  50. [50]

    Unsupervised foundation model-agnostic slide-level representation learning, 2024

    Tim Lenz, Peter Neidlinger, Marta Ligero, Georg Wölflein, Marko van Treeck, and Jakob Niko- las Kather. Unsupervised foundation model-agnostic slide-level representation learning, 2024. arXiv:2411.13623

  51. [51]

    Unveiling institution-specific bias in pathology foundation models: Detriments, causes, and potential solutions.arXiv preprint arXiv:2502.16889, 2025

    Weiping Lin, Shen Liu, Runchen Zhu, and Liansheng Wang. Unveiling institution-specific bias in pathology foundation models: Detriments, causes, and potential solutions.arXiv preprint arXiv:2502.16889, 2025

  52. [52]

    A visual-language foundation model for computational pathology.Nature Medicine, 30:863–874, 2024

    Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guil- laume Jaume, Igor Odintsov, Long Phi Le, Georg Gerber, et al. A visual-language foundation model for computational pathology.Nature Medicine, 30:863–874, 2024

  53. [53]

    Marron, David Borland, John Woosley, Xiaojun Guan, Charles Schmitt, and Nancy Thomas

    Marc Macenko, Marc Niethammer, J. Marron, David Borland, John Woosley, Xiaojun Guan, Charles Schmitt, and Nancy Thomas. A method for normalizing histology slides for quantitative analysis. InIEEE International Symposium on Biomedical Imaging: From Nano to Macro (ISBI)., 2009

  54. [54]

    Mind the gap: Continuous magnification sampling for pathology foundation models, 2026

    Alexander Möllers, Julius Hense, Florian Schulz, Timo Milbich, Maximilian Alber, and Lukas Ruff. Mind the gap: Continuous magnification sampling for pathology foundation models, 2026

  55. [55]

    Hibou: A Family of Foundational Vision Transformers for Pathology, June 2024

    Dmitry Nechaev, Alexey Pchelnikov, and Ekaterina Ivanova. Hibou: A Family of Foundational Vision Transformers for Pathology, June 2024. arXiv:2406.05074

  56. [56]

    Hibou: A family of foundational vision transformers for pathology, 2024

    Dmitry Nechaev, Alexey Pchelnikov, and Ekaterina Ivanova. Hibou: A family of foundational vision transformers for pathology, 2024

  57. [57]

    fmMAP: A framework reducing site-bias batch effect from foundation models in pathology

    Hai Cao Truong Nguyen and David Joon Ho. fmMAP: A framework reducing site-bias batch effect from foundation models in pathology. InMICCAI Workshop on Computational Pathology with Multimodal Data (COMPAYL), 2025

  58. [58]

    Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11:330, 2024

    Masaki Ochi, Daisuke Komura, Takumi Onoyama, et al. Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11:330, 2024

  59. [59]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick L...

  60. [60]

    Pluto-4: Frontier pathology foundation models, 2025

    Harshith Padigela, Shima Nofallah, Atchuth Naveen Chilaparasetti, Ryun Han, Andrew Walker, Judy Shen, Chintan Shah, Blake Martin, Aashish Sood, Elliot Miller, Ben Glass, Andy Beck, Harsha Pokkalla, and Syed Ashar Javed. Pluto-4: Frontier pathology foundation models, 2025

  61. [61]

    https://github.com/mahmoodlab/Patho-Bench/tree/ 660e77044640e3d7d2f1150cc6721e97454993bf, 2025

    Patho-Bench: Standardized benchmark for computational pathology foun- dation models. https://github.com/mahmoodlab/Patho-Bench/tree/ 660e77044640e3d7d2f1150cc6721e97454993bf, 2025

  62. [62]

    https://github.com/bifold-pathomics/PathoROB/tree/ ac1abe4df8c4d5b03aab13d6d9aabbc7205061e6, 2025

    Pathorob: Benchmark for pathology foundation models robustness to non-biological med- ical center differences. https://github.com/bifold-pathomics/PathoROB/tree/ ac1abe4df8c4d5b03aab13d6d9aabbc7205061e6, 2025

  63. [63]

    https://github

    Plismbench: A robustness benchmark of pathology foundation models. https://github. com/owkin/plism-benchmark/tree/b4f32b7c233e0f5ba8340da89ef65803cfc2ce49, 2025

  64. [64]

    Bridging local inductive bias and long-range dependencies with pixel-mamba for end-to-end whole slide image analysis

    Zhongwei Qiu, Hanqing Chao, Tiancheng Lin, Wanxing Chang, Zijiang Yang, Wenpei Jiao, Yixuan Shen, Yunshuo Zhang, Yelin Yang, Wenbin Liu, Hui Jiang, Yun Bian, Ke Yan, Dakai Jin, and Le Lu. Bridging local inductive bias and long-range dependencies with pixel-mamba for end-to-end whole slide image analysis. InProceedings of the IEEE/CVF International Confere...

  65. [65]

    Patricia Raciti, Jillian Sue, Juan A Retamero, Rodrigo Ceballos, Ran Godrich, Jeremy D Kunz, Adam Casson, Dilip Thiagarajan, Zahra Ebrahimzadeh, Julian Viret, et al. Clinical validation of artificial intelligence–augmented pathology diagnosis demonstrates significant gains in diagnostic accuracy in prostate cancer detection.Archives of Pathology & Laborat...

  66. [66]

    Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001

    Erik Reinhard, Michael Adhikhmin, Bruce Gooch, and Peter Shirley. Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001

  67. [67]

    H-optimus-0

    Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-López, Zelda Mariet, David Cahané, Eric Durand, and Jean-Philippe Vert. H-optimus-0. https://github.com/bioptimus/ releases/tree/main/models/h-optimus/v0, 2024

  68. [68]

    Kunz, Juan A

    George Shaikovski, Adam Casson, Kristen Severson, Eric Zimmermann, Yi Kan Wang, Jeremy D. Kunz, Juan A. Retamero, Gerard Oakley, David Klimstra, Christopher Kanan, Matthew Hanna, Michal Zelechowski, Julian Viret, Neil Tenenholtz, James Hall, Nicolo Fusi, Razik Yousfi, Peter Hamilton, William A. Moye, Eugene V orontsov, Siqi Liu, and Thomas J. Fuchs. PRISM...

  69. [69]

    Oriane Siméoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julie...

  70. [70]

    Spanhol, Luiz S

    Fabio A. Spanhol, Luiz S. Oliveira, Caroline Petitjean, and Laurent Heutte. A dataset for breast cancer histopathological image classification.IEEE Transactions on Biomedical Engineering, 63(7):1455–1462, 2016

  71. [71]

    Cpath-omni: A unified multimodal foundation model for patch and whole slide image analysis in computational pathology

    Yuxuan Sun, Yixuan Si, Chenglu Zhu, Xuan Gong, Kai Zhang, Pingyi Chen, Ye Zhang, Zhongyi Shui, Tao Lin, and Lin Yang. Cpath-omni: A unified multimodal foundation model for patch and whole slide image analysis in computational pathology. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 10360–10371, 2025

  72. [72]

    Yuri Tolkach, Lisa Marie Wolgast, Alexander Damanakis, Alexey Pryalukhin, Simon Schallen- berg, Wolfgang Hulla, Marie-Lisa Eich, Wolfgang Schroeder, Anirban Mukhopadhyay, Moritz Fuchs, et al. Artificial intelligence for tumour tissue detection and histological regression grading in oesophageal adenocarcinomas: a retrospective algorithm development and val...

  73. [73]

    Song, Tong Ding, Sophia J

    Anurag Vaidya, Andrew Zhang, Guillaume Jaume, Andrew H. Song, Tong Ding, Sophia J. Wagner, Ming Y . Lu, Paul Doucet, Harry Robertson, Cristina Almagro-Perez, Richard J. Chen, Dina ElHarouni, Georges Ayoub, Connor Bossi, Keith L. Ligon, Georg Gerber, Long Phi Le, and Faisal Mahmood. Molecular-driven foundation model for oncologic pathology, 2025

  74. [74]

    Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling

    Bastiaan S. Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling. Rotation equivariant cnns for digital pathology. In Alejandro F. Frangi, Julia A. Schnabel, Christos Davatzikos, Carlos Alberola-López, and Gabor Fichtinger, editors,Medical Image Computing and Computer Assisted Intervention – MICCAI 2018, pages 210–218, Cham, 2018. Springer In...

  75. [75]

    Ahmed Raza, Nasir Rajpoot, Xiyi Wu, Huai Chen, Yijie Huang, Lisheng Wang, Hyun Jung, G

    Ruchika Verma, Neeraj Kumar, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E. Ahmed Raza, Nasir Rajpoot, Xiyi Wu, Huai Chen, Yijie Huang, Lisheng Wang, Hyun Jung, G. Thomas Brown, Yanling Liu, Shuolin Liu, Seyed Alireza Fatemi Jahromi, Ali Asghar Khani, Ehsan Montahaei, Mahdieh Soleymani Baghshah, Hami...

  76. [76]

    Bernhard, Ran A

    Eugene V orontsov, George Shaikovski, Adam Casson, Julian Viret, Eric Zimmermann, Neil Tenenholtz, Yi Kan Wang, Jan H. Bernhard, Ran A. Godrich, Juan A. Retamero, Jinru Shia, Mithat Gonen, Martin R. Weiser, David S. Klimstra, Razik Yousfi, Nicolo Fusi, Thomas J. Fuchs, Kristen Severson, and Siqi Liu. Prism2: Unlocking multi-modal general pathology ai with...

  77. [77]

    A pathology foundation model for cancer diagnosis and prognosis prediction

    Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, Jietian Jin, Jiayu Zhang, Ruijiang Li, Hongping Tang, Kanran Wang, Yu Li, Fang Wang, Yulong Peng, Junyou Zhu, Jing Zhang, and Christopher. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature, 634(8035):970–978, October 2024

  78. [78]

    A petri dish for histopathology image analysis

    Jerry Wei, Arief Suriawinata, Bing Ren, Xiaoying Liu, Mikhail Lisovsky, Louis Vaickus, Charles Brown, Michael Baker, Naofumi Tomita, Lorenzo Torresani, et al. A petri dish for histopathology image analysis. InArtificial Intelligence in Medicine: 19th International Conference on Artificial Intelligence in Medicine, AIME 2021, Virtual Event, June 15–18, 202...

  79. [79]

    Wright, Ari Robicsek, Brian Piening, Carlo Bifulco, Sheng Wang, and Hoifung Poon

    Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier González, Yu Gu, Yanbo Xu, Mu Wei, Wenhui Wang, Shuming Ma, Furu Wei, Jianwei Yang, Chunyuan Li, Jianfeng Gao, Jaylen Rosemon, Tucker Bower, Soohee Lee, Roshanthi Weerasinghe, Bill J. Wright, Ari Robicsek, Brian Piening, Carlo Bifulco, Shen...

  80. [80]

    A multi- modal knowledge-enhanced whole-slide pathology foundation model.Nature Communications, 16:11406, 2025

    Yingxue Xu, Yihui Wang, Fengtao Zhou, Jiabo Ma, Cheng Jin, Shu Yang, Jinbang Li, Zhengyu Zhang, Chenglong Zhao, Huajun Zhou, Zhenhui Li, Huangjing Lin, Xin Wang, Jiguang Wang, Anjia Han, Ronald Cheong Kin Chan, Li Liang, Xiuming Zhang, and Hao Chen. A multi- modal knowledge-enhanced whole-slide pathology foundation model.Nature Communications, 16:11406, 2025

Showing first 80 references.