Pith. sign in

REVIEW 4 major objections 5 minor 69 references

Robustifying pathology foundation models via fine-tuning

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A single fine-tuning step applied to ten pathology foundation models raises average robustness by 23% and cross-benchmark performance by 43%, with no observed trade-off.

desk verdict Big empirical claim, but the fine-tuning recipe is missing and the fine-tuning data are undisclosed, so the central result cannot be checked. read the letter →

arxiv 2607.22861 v1 pith:F2GGFDPO submitted 2026-07-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords pathologyfoundationmodelsrobustnessfine-tuningscannervariabilitystaindomainshiftfeaturespaceinvariancehistopathology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a single fine-tuning step can make pathology foundation models robust to scanner and staining variability without sacrificing their generic quality. Applied uniformly to ten existing encoders, the step raised the average PathoROB robustness index from 0.72 to 0.87 and improved combined HEST, THUNDER, and Patho-Bench performance by 43%, with every model moving up on robustness and overall rank. If true, it means a cheap, label-free procedure can decouple pathology representations from acquisition confounders, so a single encoder can be deployed across laboratories with different scanners and stain protocols. The paper also gives a mechanistic picture: fine-tuning reorients the dominant directions of feature space from 'where the slide was digitized' to 'what the morphology is.'

What carries the argument

The load-bearing object is the fine-tuned encoder's reorganized feature geometry, produced by a fine-tuning step applied uniformly to all ten models. The paper supports its mechanism with two observations: on the PLISM dataset, scanner shift appears as a near-linear offset in feature space, so a feature-level correction seems possible in principle; but cross-scanner retrieval on SCORPION shows that invariance builds up across transformer depth, with fine-tuning reaching a given retrieval quality roughly eight blocks earlier and a higher asymptote (mAP about 0.99 versus 0.91 at the final block). This locates acquisition invariance deep in the transformer and identifies fine-tuning as re-purposing the dominant feature-space directions from acquisition site to biological class.

What would settle it

Check the fine-tuning data manifest against PathoROB (including its TCGA and Tolkach cohorts), HEST, THUNDER, and Patho-Bench: if any of those slides or tiles appear in fine-tuning, the robustness and performance gains are partly in-distribution rather than evidence of generalization to unseen acquisition sources. If the data are clean, rerunning the recipe with those benchmark cohorts explicitly held out should reproduce the average 23% PathoROB gain and the 43% cross-benchmark improvement.

Watch

Extended reading notes

Core claim

The central claim is that acquisition robustness is not a property that must be bought with pretraining scale or traded off against downstream utility: fine-tuning the encoder itself is sufficient. Across ten pathology foundation models spanning different architectures and pretraining recipes, every model improved its PathoROB robustness index after fine-tuning (one-sided Wilcoxon signed-rank $p<10^{-4}$), and every model improved its overall rank on the combined HEST, THUNDER, and Patho-Bench leaderboards ($p<10^{-4}$); the best fine-tuned encoder, UNI2-h, moved from total rank 21 to 5. On the CAMELYON subset of PathoROB, the leading axis of Phikon-v2's feature space switched from clustering by medical center (Adjusted Rand Index 0.46 against center, 0.00 against metastasis) to clustering by metastasis status (ARI 0.67 against metastasis, 0.01 against center), for five centers that were not seen during fine-tuning. The paper interprets the joint up-and-to-the-right shift as evidence that scanner- and stain-related directions are nuisance dimensions: removing them frees capacity for biologically relevant structure.

Load-bearing premise

The claim that robustness generalizes to unseen acquisition sources depends on the fine-tuning data being disjoint from every evaluation benchmark; the paper demonstrates this for five CAMELYON centers but never states whether the TCGA or Tolkach parts of PathoROB, or the data behind HEST, THUNDER, and Patho-Bench, were excluded.

Editorial extensions

If this is right

  • A laboratory can adopt a robustified encoder as a drop-in feature extractor and expect it to keep working when the scanner or stain protocol changes, without retraining the encoder.
  • Fine-tuning acts as an equalizer: the largest robustness gains land on the least robust base models, so older or smaller encoders can be brought closer to top-tier performance without larger pretraining corpora.
  • Because every fine-tuned model improves or matches its base model on aggregate benchmarks, robustification can be applied before downstream heads are trained, with no observed performance tax.
  • The released robust versions of Phikon-v2 and Midnight-12k (Phaet and Mascaret) make the effect immediately available for other pipelines; Mascaret ranks first on PathoROB and second on average downstream performance among publicly available models.
  • Since scanner invariance emerges roughly eight transformer blocks earlier after fine-tuning, even models that read intermediate features inherit part of the robustness gain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If scanner shift really is a near-linear offset, a testable extension is to combine the fine-tuning recipe with an explicit linear correction estimated on paired multi-scanner slides; the paper's depth analysis suggests the linear correction alone would be partial, because the invariance is assembled deep in the transformer.
  • Because the fine-tuning data and recipe hyperparameters are not disclosed in the main text, external replication on completely unseen scanners will determine whether the no-trade-off claim is a property of the method or of the specific evaluation setup.
  • The PathoROB index measures feature-space dominance of biology over confounders, not end-task robustness; a natural next check is whether fine-tuned encoders reduce site-specific errors in biomarker or survival tasks under external validation, which the paper explicitly leaves for future work.
  • If the gains hold under strict data separation, the recipe becomes a general post-processing step that could be applied to any new pathology encoder, including vision-language models, rather than a one-off training scheme.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript claims a novel fine-tuning recipe that, applied to ten pathology foundation models, jointly improves robustness to acquisition factors and downstream task performance with no observed trade-off. The evaluation uses PathoROB for robustness, and HEST, THUNDER, and Patho-Bench for performance, reporting consistent gains across all ten pairs, with Wilcoxon signed-rank tests, and releases two fine-tuned models (Phaet and Mascaret). The paper also presents analyses on PLISM and SCORPION to argue that acquisition shifts are near-linear in feature space and that fine-tuning instills invariance across transformer depth.

Significance. If the empirical claims hold, the paper would be practically valuable: it offers a model-agnostic way to improve robustness across very different pathology FMs, with gains also on downstream tasks, and it releases models that practitioners could adopt directly. The benchmark coverage is broad and external (PathoROB, HEST, THUNDER, Patho-Bench), the set of base FMs includes models from several independent groups, and the internal numbers are reproducible in the sense that the tables are detailed and the paired Wilcoxon tests support the direction of the robustness effect. However, the central contribution, the fine-tuning recipe itself, is never described, and the training data are not disclosed, so the generalization claim cannot currently be evaluated or reproduced.

major comments (4)
  1. [Section 3 (Experimental setup)] The manuscript never specifies the fine-tuning recipe that the abstract and introduction present as the central contribution. There is no loss function, no fine-tuning dataset, no hyperparameters, no optimization details, and no compute budget anywhere in Sections 3–5. As a result, a reader cannot reproduce the method, cannot determine what makes it 'novel', and cannot assess whether the recipe is distinct from existing robustness methods discussed in Section 2. This is load-bearing: the entire paper is an evaluation of a method that is never stated. A complete method section, including training data provenance, objective, and hyperparameters, is required.
  2. [Section 4.1 and Figure 2 caption] The only statement that evaluation data were unseen during fine-tuning concerns the five CAMELYON centers shown in Figure 2. No statement is made for the TCGA and Tolkach components of PathoROB (Section 3.2), for HEST, THUNDER, or Patho-Bench (Section 3.3), or for PLISM and SCORPION (Section 5). If any of these cohorts or slides were used for fine-tuning, then the claimed robustness gains on 'unseen acquisition sources' are partly in-distribution and the central generalization claim collapses. The paper must provide an explicit list of the fine-tuning data and a per-benchmark statement of disjointness.
  3. [Section 4.1 and Conclusion] The claim 'no observed trade-off' is contradicted by the paper's own tables. Table 2 shows that H0-mini's THUNDER rank sum worsens from 61 to 65 and AquaViT's is unchanged at 60. Table 3 shows that GenBio-PathFM's average HEST Pearson correlation decreases from 0.4197 to 0.4178. These are regressions or ties on individual benchmarks, even if aggregate cross-benchmark ranks improve. The text should be qualified to say that overall performance improves on aggregate, while individual benchmark regressions do occur, rather than claiming no trade-off at all.
  4. [Appendix A (Tables 5–7)] The extended leaderboards mix official published values for base models with in-house computed values for fine-tuned models and for some mixed-precision base models. If the evaluation environments differ (e.g., full precision vs. mixed precision, different preprocessing or hardware), the ranks in Table 1 are not strictly comparable across rows. Please state explicitly which rows were computed in-house, which were taken from official leaderboards, and whether official leaderboard values were produced under the same precision and preprocessing protocol.
minor comments (5)
  1. [Section 3.3 vs. Table 3] Section 3.3 lists nine cancer types including 'hepatocellular carcinoma (liver cancer, HCC)', but Table 3's columns are labeled IDC, PRAD, PAAD, SKCM, COAD, READ, CCRCC, LUNG, LYMPH IDC. The LUNG column appears in place of HCC. Please align the list and the table headers.
  2. [Section 5, Figure 3] Figure 3 shows 225 points per scanner but the text does not explain how this subset of the 16,278 PLISM tiles was sampled. A brief sampling description would help interpret the PCA projection.
  3. [Section 4.2, Table 2] The THUNDER rank sums in Table 2 (e.g., UNI2-h base 26, fine-tuned 20) differ from the extended leaderboard in Table 5 (UNI2-h base 35, fine-tuned 31) because Table 5 includes more models and uses a different ranking pool. The relationship between the two tables should be explained in the text.
  4. [Section 3.1 and Appendix A] The paper refers to 'wearewaiv.github.io/histoboard/models' for detailed model information. This is not a stable scientific citation; please provide a persistent repository, versioned release, or arXiv reference for the model details.
  5. [General] Several hyperlinks are written as raw URLs (e.g., huggingface.co/wearewaiv/models). Please format them properly and ensure they are accessible at the time of publication.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the fine-tuning gains are measured on external benchmarks and the self-citations are base models, not evaluators.

full rationale

The paper's central claim is empirical: applying a fine-tuning recipe to ten pathology foundation models improves PathoROB robustness and cross-benchmark performance. The evaluation benchmarks (PathoROB, HEST, THUNDER, Patho-Bench) are external artifacts, and the effect is demonstrated on foundation models developed by other groups (UNI2-h, Virchow2, Prov-GigaPath, GenBio-PathFM, H-Optimus-0), so the measurement is not tautological and does not reduce to the paper's own definitions. The self-citations present in the reference list (Phikon, Phikon-v2, H0-mini) are base models used as inputs to the fine-tuning procedure, not as the source of the robustness index or the downstream metrics; they do not carry the load of the central claim. The main substantive weakness is that the fine-tuning data, loss function, and training protocol are never specified, which makes the independence of the fine-tuning data from the evaluation benchmarks unverifiable; however, this is a reproducibility and possible-leakage concern, not a circularity of the kind where a prediction is equivalent to its inputs by construction or where a fitted parameter is renamed as a prediction. No equation, definition, or quoted passage in the manuscript exhibits such a reduction, so the appropriate finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The fine-tuning recipe is entirely undisclosed, so the ledger cannot enumerate its free parameters in detail; instead we list the structural unknowns that the central claim depends on: the training objective, the data, and the hyperparameters. Axioms include the validity of the PathoROB proxy, the disjointness of training and evaluation data (for which only CAMELYON is attested), and the comparability of external base-model leaderboard numbers with in-house fine-tuned numbers.

free parameters (3)
  • Fine-tuning objective and loss weights = Not disclosed
    The 'novel fine-tuning recipe' is never defined, so the loss function, any invariance term(s), and their relative weights are free parameters beyond the reader's reach.
  • Fine-tuning dataset composition = Not disclosed
    The data used to fine-tune, its size, composition, and exclusion rules are never stated; whether any evaluation cohort (e.g., TCGA) appears in the training set is unknown.
  • Fine-tuning hyperparameters = Not disclosed
    Learning rate, epochs, batch size, augmentation, and optimization details are absent, making the recipe impossible to reproduce or compare.
assumptions (3)
  • domain assumption PathoROB robustness index is a valid proxy for downstream robustness to acquisition shift.
    The conclusion explicitly notes 'We measure robustness through the PathoROB index... it does not directly measure downstream robustness under domain shift.' The paper's central robustness claim depends on this proxy.
  • domain assumption Fine-tuning data is disjoint from all evaluation benchmarks (PathoROB, HEST, THUNDER, Patho-Bench).
    Only the CAMELYON centers in Figure 2 are explicitly declared unseen. No statement covers the TCGA and Tolkach components of PathoROB, or the other benchmarks. Without this axiom, the 'generalization to unseen acquisition sources' claim is unsupported.
  • domain assumption Official leaderboard results for base models and in-house results for fine-tuned models are directly comparable.
    Section 3.3 says evaluations use 'official implementations and default hyperparameters,' and Supplementary A quotes official published values for base models while fine-tuned values are computed in-house; any protocol difference (e.g., mixed precision, feature concatenation rules) could bias the paired comparisons.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robustifying pathology foundation models via fine-tuning." pith.science (2026). https://pith.science/paper/F2GGFDPO

@misc{pith2026260722861,
  author       = {Pith},
  title        = {Pith review of: Robustifying pathology foundation models via fine-tuning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F2GGFDPO}},
  note         = {Machine review of arXiv:2607.22861}
}
read the original abstract

Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness of pathology FMs to acquisition factors. Applied to ten different FMs, our fine-tuning strategy consistently improves robustness for every model as well as downstream performance, with no observed trade-off. On average, it raises the PathoROB robustness index by 23% (from 0.72 to 0.87) and increases the overall cross-benchmark performance by 43% on Patho-Bench, HEST and THUNDER combined, with individual gains reaching up to 72% in robustness (Phikon-v2) and 76% in performance (Midnight-12k). We publicly release the fine-tuned versions of Phikon-v2 (Phaet) and Midnight-12k (Mascaret) at https://huggingface.co/wearewaiv/models.

Figures

Figures reproduced from arXiv: 2607.22861 by the authors.

Figure 1
Figure 1. Fine-tuning improves robustness and performance jointly. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Fine-tuning reorganizes Phikon-v2’s feature space around biology rather than [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Acquisition factors seem linearly encoded in feature space. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Fine-tuning makes scanner invariance emerge earlier and stronger with depth. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 27 canonical work pages

  1. [1]

    Atlas 2 – foundation models for clinical deployment.arXiv preprint arXiv:2601.05148, 2026

    Maximilian Alber et al. Atlas 2 – foundation models for clinical deployment.arXiv preprint arXiv:2601.05148, 2026

  2. [2]

    Péter Bándi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, et al. From detection of individual metastases to classification of lymph node status at the patient level: The camelyon17 challenge.IEEE Transactions on Medical Imaging, 38(2): 550–560, 2019. doi: 10.1109/TMI.2018.2867350

  3. [3]

    Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen A. W. M. Van Der Laak, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer.JAMA, 318(22):2199–2210, 2017. doi: 10.1001/jama.2017.14585

  4. [4]

    H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025

    Bioptimus. H-optimus-1.https://huggingface.co/bioptimus/H-optimus-1, 2025

  5. [5]

    Pathology foundation models are scanner sensitive: Benchmark and mitigation with contrastive scangen loss

    Gianluca Carloni, Biagio Brattoli, Seongho Keum, Jongchan Park, Taebum Lee, Chang Ho Ahn, and Sergio Pereira. Pathology foundation models are scanner sensitive: Benchmark and mitigation with contrastive scangen loss. InMICCAI Workshop on Foundation Models for General Medical AI (MedAGI), pages 44–53. Springer, 2025. arXiv:2507.22092

  6. [6]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InIEEE/CVF International Conference on Computer Vision (ICCV), pages 9650–9660, 2021

  7. [7]

    Steiner, Tapabrata Chakraborti, and Adrienne M

    Binghao Chai, Jianan Chen, Paul Cool, Fatine Oumlil, Anna Tollitt, David F. Steiner, Tapabrata Chakraborti, and Adrienne M. Flanagan. Impact of tissue staining and scanner variation on the performance of pathology foundation models: a study of sarcomas and their mimics.The Journal of Pathology: Clinical Research, 12(2):e70080, 2026. doi: 10.1002/2056-4538.70080

  8. [8]

    Chen, Tong Ding, Ming Y

    Richard J. Chen, Tong Ding, Ming Y . Lu, Drew F. K. Williamson, Guillaume Jaume, Andrew H. Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, Mane Williams, Lukas Oldenburg, Luca L. Weishaupt, Judy J. Wang, Anurag Vaidya, Long Phi Le, Georg Gerber, Sharifa Sahai, Walt Williams, and Faisal Mahmood. Towards a general-purpose foundation model for ...

Show all 69 references
  1. [9]

    Improved baselines with momentum contrastive learning.arXiv preprint arXiv:2003.04297, 2020

    Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning.arXiv preprint arXiv:2003.04297, 2020

  2. [10]

    Ozan Ciga, Tony Xu, and Anne L. Martel. Self supervised contrastive learning for digital histopathology. Machine Learning with Applications, 7:100198, 2022. doi: 10.1016/j.mlwa.2021.100198

  3. [11]

    Nicholson, Jean-Yves Blay, Françoise Galateau-Sallé, Gilles Wainrib, and Thomas Clozel

    Pierre Courtiol, Charles Maussion, Matahi Moarii, Elodie Pronier, Samuel Pilcer, Meriem Sefta, Pierre Manceron, Sylvain Toldo, Mikhail Zaslavskiy, Nolwenn Le Stang, Nicolas Girard, Olivier Elemento, Andrew G. Nicholson, Jean-Yves Blay, Françoise Galateau-Sallé, Gilles Wainrib,...

  4. [12]

    de Jong, Eric Marcus, and Jonas Teuwen

    Edwin D. de Jong, Eric Marcus, and Jonas Teuwen. Current pathology foundation models are unrobust to medical center differences.arXiv preprint arXiv:2501.18055, 2025

  5. [13]

    Self-supervision closes the gap between weak and strong supervision in histology.arXiv preprint arXiv:2012.03583, 2020

    Olivier Dehaene, Axel Camara, Olivier Moindrot, Axel de Lavergne, and Pierre Courtiol. Self-supervision closes the gap between weak and strong supervision in histology.arXiv preprint arXiv:2012.03583, 2020. ML4H Workshop, NeurIPS 2020

  6. [14]

    Taher Dehkharghanian, Azam Asilian Bidgoli, Abtin Riasatian, Pooria Mazaheri, Clinton J. V . Camp- bell, Liron Pantanowitz, H. R. Tizhoosh, and Shahryar Rahnamayan. Biased data, biased ai: deep networks predict the acquisition site of tcga images.Diagnostic Pathology, 18(1):67...

  7. [15]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255,

  8. [16]

    Wagner, Andrew H

    Tong Ding, Sophia J. Wagner, Andrew H. Song, et al. A multimodal whole-slide foundation model for pathology.Nature Medicine, 2025. arXiv:2411.19666

  9. [17]

    Medi: Metadata-guided diffusion models for mitigating biases in tumor classification

    David Jacob Drexlin, Jonas Dippel, Julius Hense, Niklas Prenißl, Grégoire Montavon, Frederick Klauschen, and Klaus-Robert Müller. Medi: Metadata-guided diffusion models for mitigating biases in tumor classification. InMedical Image Computing and Computer Assisted Intervention ...

  10. [18]

    Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023

    Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling.medRxiv, 2023. doi: 10.1101/2023.07.21.23292757

  11. [19]

    Phikon-v2, a large and public feature extractor for biomarker prediction.arXiv preprint arXiv:2409.09173, 2024

    Alexandre Filiot, Paul Jacob, Alice Mac Kain, and Charlie Saillard. Phikon-v2, a large and public feature extractor for biomarker prediction.arXiv preprint arXiv:2409.09173, 2024

  12. [20]

    Distilling foundation models for robust and efficient models in digital pathology

    Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Rémy Dubois, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, and Antoine Olivier. Distilling foundation models for robust and efficient models in digital pathology. InMedical Image Co...

  13. [21]

    Domain-adversarial training of neural networks.Journal of Machine Learning Research, 17(59):1–35, 2016

    Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks.Journal of Machine Learning Research, 17(59):1–35, 2016

  14. [22]

    Schüffler

    Christian Grashei, Christian Brechenmacher, Rao Muhammad Umer, Jingsong Liu, Carsten Marr, Ewa Szczurek, and Peter J. Schüffler. Pathryoshka: Compressing pathology foundation models via multi-teacher knowledge distillation with nested embeddings.arXiv preprint arXiv:2511.23204, 2025

  15. [23]

    Gustafsson and Mattias Rantalainen

    Fredrik K. Gustafsson and Mattias Rantalainen. Evaluating computational pathology foundation models for prostate cancer grading under distribution shifts.arXiv preprint arXiv:2410.06723, 2024

  16. [24]

    Audun L. Henriksen, Ole-Johan Skrede, Lisa van der Schee, Enric Domingo, Karolina Cyll, Wanja Kildal, Joakim Kalsnes, Manohar Pradhan, Hanne Askautrud, Tarjei Sveinsgjerd Hveem, Knut Liestøl, David N. Church, David J. Kerr, and Andreas Kleppe. Enabling clinical use of foundati...

  17. [25]

    Howard, James Dolezal, Sara Kochanny, Jefree Schulte, Heather Chen, Lara Heij, Dezheng Huo, Rita Nanda, Olufunmilayo I

    Frederick M. Howard, James Dolezal, Sara Kochanny, Jefree Schulte, Heather Chen, Lara Heij, Dezheng Huo, Rita Nanda, Olufunmilayo I. Olopade, Jakob N. Kather, Nicole Cipriani, Robert L. Grossman, and Alexander T. Pearson. The impact of site-specific digital histology signature...

  18. [26]

    Knowledge-guided adaptation of pathology foundation models effectively improves cross-domain generalization and demographic fairness.Nature Communications, 16:11485, 2025

    Yanyan Huang et al. Knowledge-guided adaptation of pathology foundation models effectively improves cross-domain generalization and demographic fairness.Nature Communications, 16:11485, 2025. doi: 10.1038/s41467-025-66300-y

  19. [27]

    Tomczak, and Max Welling

    Maximilian Ilse, Jakub M. Tomczak, and Max Welling. Attention-based deep multiple instance learning. In International Conference on Machine Learning (ICML), volume 80 ofProceedings of Machine Learning Research, pages 2132–2141, 2018

  20. [28]

    Domain generalization in computational pathology: Survey and guidelines.ACM Computing Surveys, 2025

    Mostafa Jahanifar, Manahil Raza, Kesi Xu, Trinh Vuong, Robert Jewsbury, Adam Shephard, Neda Zamanitajeddin, Jin Tae Kwak, Shan E Ahmed Raza, Fayyaz Minhas, and Nasir Rajpoot. Domain generalization in computational pathology: Survey and guidelines.ACM Computing Surveys, 2025. a...

  21. [29]

    Song, Ming Y

    Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y . Lu, Cristina Almagro-Pérez, Sophia J. Wagner, Anurag J. Vaidya, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, and Faisal Mahmood. Hest-1k: A dataset for spatial transcriptomics and histology image analysis. InAdvance...

  22. [30]

    Saarthak Kapse, Mehmet Aygün, Elijah Cole, Emma Lundberg, Le Song, and Eric P. Xing. Genbio-pathfm: A state-of-the-art foundation model for histopathology.bioRxiv, 2026. doi: 10.64898/2026.03.17.712534. https://huggingface.co/genbio-ai/genbio-pathfm

  23. [31]

    Training state-of-the-art pathology foundation models with orders of magnitude less data

    Mikhail Karasikov, Joost van Doorn, Nicolas Känzig, Melis Erdal Cesur, Hugo Mark Horlings, Robert Berke, Fei Tang, and Sebastian Otálora. Training state-of-the-art pathology foundation models with orders of magnitude less data. InMedical Image Computing and Computer Assisted I...

  24. [32]

    Pearson, Niels Halama, Dirk Jäger, Jeremias Krause, Sven H

    Jakob Nikolas Kather, Alexander T. Pearson, Niels Halama, Dirk Jäger, Jeremias Krause, Sven H. Loosen, Alexander Marx, Peter Boor, Frank Tacke, Ulf Peter Neumann, Heike I. Grabsch, Takaki Yoshikawa, Hermann Brenner, Jenny Chang-Claude, Michael Hoffmeister, Christian Trautwein,...

  25. [33]

    Viergever, and Josien P

    Stefan Klein, Marius Staring, Keelin Murphy, Max A. Viergever, and Josien P. W. Pluim. elastix: a toolbox for intensity-based medical image registration.IEEE Transactions on Medical Imaging, 29(1):196–205,

  26. [34]

    de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller

    Jonah Kömen, Edwin D. de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller. Towards robust foundation models for digital pathology.Nature Communications, 17...

  27. [35]

    Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9):110424, 2022

    Daisuke Komura, Akihiro Kawabe, Keisuke Fukuta, Kyohei Sano, Toshikazu Umezaki, Hirotomo Koda, Ryohei Suzuki, Ken Tominaga, Mieko Ochi, Hiroki Konishi, et al. Universal encoding of pan-cancer histology by deep texture representations.Cell Reports, 38(9):110424, 2022. doi: 10.1...

  28. [36]

    Beyond diagnostic performance: Revealing and quantifying ethical risks in pathology foundation models.arXiv preprint arXiv:2502.16889, 2025

    Weiping Lin, Shen Liu, Runchen Zhu, and Liansheng Wang. Beyond diagnostic performance: Revealing and quantifying ethical risks in pathology foundation models.arXiv preprint arXiv:2502.16889, 2025

  29. [37]

    A generalizable pathology foundation model using a unified knowledge distillation pretraining framework.Nature Biomedical Engineering, 10(3):545–564, 2026

    Jiabo Ma, Zhengrui Guo, Fengtao Zhou, Yihui Wang, Yingxue Xu, Jinbang Li, Fang Yan, Yu Cai, Zhengjie Zhu, Cheng Jin, Yi Lin, Xinrui Jiang, Chenglong Zhao, Danyi Li, Anjia Han, Zhenhui Li, Ronald Cheong Kin Chan, Jiguang Wang, Peng Fei, Kwang-Ting Cheng, Shaoting Zhang, Li Lian...

  30. [38]

    Marc Macenko, Marc Niethammer, J. S. Marron, David Borland, John T. Woosley, Xiaojun Guan, Charles Schmitt, and Nancy E. Thomas. A method for normalizing histology slides for quantitative analysis. In IEEE International Symposium on Biomedical Imaging: From Nano to Macro (ISBI...

  31. [39]

    Thunder: Tile-level histopathology image understanding benchmark.Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025

    Pierre Marza, Leo Fillioux, Sofiène Boutaj, Kunal Mahatha, Christian Desrosiers, Pablo Piantanida, Jose Dolz, Stergios Christodoulidis, and Maria Vakalopoulou. Thunder: Tile-level histopathology image understanding benchmark.Advances in Neural Information Processing Systems (N...

  32. [40]

    fmmap: A framework reducing site-bias batch effect from foundation models in pathology

    Hai Cao Truong Nguyen and David Joon Ho. fmmap: A framework reducing site-bias batch effect from foundation models in pathology. InMICCAI Workshop on Computational Pathology with Multimodal Data (COMPAYL), 2025

  33. [41]

    doi: 10.1109/ISBI.2009.5193250. 13

  34. [42]

    Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11(1):330, 2024

    Masaki Ochi, Daisuke Komura, Takumi Onoyama, Koki Shinbo, Haruya Endo, Hiroto Odaka, Miwako Kakiuchi, Hiroto Katoh, Tetsuo Ushiku, and Shumpei Ishikawa. Registered multi-device/staining histology image dataset for domain-agnostic machine learning models.Scientific Data, 11(1):...

  35. [43]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...

  36. [44]

    Nguyen et al

    Tan H. Nguyen et al. Contrimix: Scalable stain color augmentation for domain generalization without domain labels.arXiv preprint arXiv:2306.04527, 2023

  37. [45]

    Scorpion: Addressing scanner-induced variability in histopathology

    Jeongun Ryu, Heon Song, Seungeun Lee, Soo Ick Cho, Jiwon Shin, Kyunghyun Paeng, and Sérgio Pereira. Scorpion: Addressing scanner-induced variability in histopathology. InUncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE), MICCAI 2025 Workshop, Lec...

  38. [46]

    Self supervised learning improves dmmr/msi detection from histology slides across multiple cancers

    Charlie Saillard, Olivier Dehaene, Tanguy Marchand, Olivier Moindrot, Aurélien Kamoun, Benoît Schmauch, and Simon Jegou. Self supervised learning improves dmmr/msi detection from histology slides across multiple cancers. InMICCAI Workshop on Computational Pathology (COMPAY), v...

  39. [47]

    Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001

    Erik Reinhard, Michael Ashikhmin, Bruce Gooch, and Peter Shirley. Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41, 2001. doi: 10.1109/38.946629

  40. [48]

    H-optimus-0

    Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-López, Zelda Mariet, David Cahané, Eric Durand, and Jean-Philippe Vert. H-optimus-0. https://github.com/bioptimus/releases/tree/main/models/ h-optimus/v0, 2024

  41. [49]

    A deep learning model to predict rna-seq expression of tumours from whole slide images.Nature Communications, 11(1):3877, 2020

    Benoît Schmauch, Alberto Romagnoni, Elodie Pronier, Charlie Saillard, Pascale Maillé, Julien Calderaro, Aurélien Kamoun, Meriem Sefta, Sylvain Toldo, Mikhail Zaslavskiy, Thomas Clozel, Matahi Moarii, Pierre Courtiol, and Gilles Wainrib. A deep learning model to predict rna-seq...

  42. [50]

    Validation of msintuit as an ai-based pre-screening tool for msi detection from colorectal cancer histology slides.Nature Communications, 14(1):6695, 2023

    Charlie Saillard, Rémy Dubois, Oussama Tchita, Nicolas Loiseau, Théophile Garcia, Aurélie Adriansen, Séverine Carpentier, Joël Reyre, Diana Enea, Katharina V on Loga, Aurélien Kamoun, Stéphane Rossat, Céline Wiscart, Meriem Sefta, Michaël Auffret, Lionel Guillou, Arnaud Fouill...

  43. [51]

    Transmil: Transformer based correlated multiple instance learning for whole slide image classification

    Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, and Yongbing Zhang. Transmil: Transformer based correlated multiple instance learning for whole slide image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2021. arXiv:2106.00908

  44. [52]

    Randstainna: Learning stain-agnostic features by bridging stain augmentation and normalization

    Yiqing Shen, Yulin Luo, Dinggang Shen, and Jing Ke. Randstainna: Learning stain-agnostic features by bridging stain augmentation and normalization. InMedical Image Computing and Computer Assisted Intervention (MICCAI). Springer, 2022. arXiv:2206.12694

  45. [53]

    Schönpflug, Nikki van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor H

    Lydia A. Schönpflug, Nikki van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor H. Koelzer, and Maxime W. Lafarge. A protocol for evaluating robustness to h&e staining variation in computational pathology models.arXiv preprint arXiv:2603.12886, 2026

  46. [54]

    Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology.Medical Image Analysis, 58:101544, 2019

    David Tellez, Geert Litjens, Péter Bándi, Wouter Bulten, John-Melle Bokhorst, Francesco Ciompi, and Jeroen van der Laak. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology.Medical Image Analys...

  47. [55]

    Gustafsson, Kajsa Ledesma Eriksson, and Mattias Rantalainen

    Erik Thiringer, Fredrik K. Gustafsson, Kajsa Ledesma Eriksson, and Mattias Rantalainen. Scanner-induced domain shifts undermine the robustness of pathology foundation models.arXiv preprint arXiv:2601.04163, 2026

  48. [56]

    V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, et al

    Oriane Siméoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. 14

  49. [57]

    Yuri Tolkach, Lisa Marie Wolgast, Alexander Damanakis, Alexey Pryalukhin, Simon Schallenberg, Wolfgang Hulla, Marie-Lisa Eich, Wolfgang Schroeder, Anirban Mukhopadhyay, Moritz Fuchs, et al. Artificial intelligence for tumour tissue detection and histological regression grading...

  50. [58]

    Structure-preserving color normalization and sparse stain separation for histological images.IEEE Transactions on Medical Imaging, 35(8):1962–1971, 2016

    Abhishek Vahadane, Tingying Peng, Amit Sethi, Shadi Albarqouni, Lichao Wang, Maximilian Baust, Katja Steiger, Anna Melissa Schlitter, Irene Esposito, and Nassir Navab. Structure-preserving color normalization and sparse stain separation for histological images.IEEE Transaction...

  51. [59]

    Tizhoosh

    Hamid R. Tizhoosh. Beyond the failures: Rethinking foundation models in pathology.arXiv preprint arXiv:2510.23807, 2025

  52. [60]

    Georg Wölflein, Dyke Ferber, Asier Rabasco Meneghetti, Omar S. M. El Nahhas, Daniel Truhn, Zunamys I. Carrero, David J. Harrison, Ognjen Arandjelovi ´c, and Jakob Nikolas Kather. A good feature extractor is all you need for weakly supervised pathology slide classification. InC...

  53. [61]

    Wright, Ari Robicsek, Brian Piening, Carlo Bifulco, Sheng Wang, and Hoifung Poon

    Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier González, Yu Gu, Yanbo Xu, Mu Wei, Wenhui Wang, Shuming Ma, Furu Wei, Jianwei Yang, Chunyuan Li, Jianfeng Gao, Jaylen Rosemon, Tucker Bower, Soohee Lee, Roshant...

  54. [62]

    A pathology foundation model for cancer diagnosis and prognosis prediction.Nature, 634(8035):970–978, 2024

    Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, et al. A pathology foundation model for cancer diagnosis and prognosis prediction.Nature, 634(8035):970–978, 2024. doi: 10.1038/s41586-024-07894-z

  55. [63]

    Accelerating data processing and benchmarking of ai models for pathology.arXiv preprint arXiv:2502.06750, 2025

    Andrew Zhang, Guillaume Jaume, Anurag Vaidya, Tong Ding, and Faisal Mahmood. Accelerating data processing and benchmarking of ai models for pathology.arXiv preprint arXiv:2502.06750, 2025

  56. [64]

    ibot: Image bert pre-training with online tokenizer

    Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. InInternational Conference on Learning Representations (ICLR),

  57. [65]

    Virchow2: Scaling self-supervised mixed magnification models in pathology.arXiv preprint arXiv:2408.00738, 2024

    Eric Zimmermann, Eugene V orontsov, Julian Viret, Adam Casson, Michal Zelechowski, George Shaikovski, Neil Tenenholtz, James Hall, David Klimstra, Razik Yousfi, Thomas Fuchs, Nicolo Fusi, Siqi Liu, and Kristen Severson. Virchow2: Scaling self-supervised mixed magnification mod...

  58. [66]

    Barlow twins: Self-supervised learning via redundancy reduction

    Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. InInternational Conference on Machine Learning (ICML), volume 139 ofProceedings of Machine Learning Research, pages 12310–12320, 2021

  59. [2009]

    doi: 10.1109/CVPR.2009.5206848

  60. [2010]

    doi: 10.1109/TMI.2009.2035616

  61. [2024]

    doi: 10.1038/s41586-024-07441-w

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.